Why Better AI Starts With Better Data with Gemma Newlove – VistaTalks Ep 204

Keywords: AI, Data, Localization, Data Annotation, Gemma Newlove, VistaTalks, Vistatec, Multilingual AI Data, Data collection, Data Quality, Synthetic Data, Multilingual AI
Run Time:
Release Date: September 16, 2026

Listen to the audio or watch the video below.

Gemma Newlove, Director of Sales at Vistatec, joins Host Simon Hodgkins for a timely conversation about the rapidly expanding world of AI data, multilingual model development, and the critical role localization expertise plays in building AI systems that work reliably across global markets. Drawing on a decade in localization and an increasing focus on AI data, Gemma explores data collection, annotation, multimodal AI, synthetic data, representation, quality, governance, and the enduring importance of human judgment.

Gemma Newlove, Director of Sales at Vistatec, joins Host Simon Hodgkins for a timely conversation about the rapidly expanding world of AI data, multilingual model development, and the critical role localization expertise plays in building AI systems that work reliably across global markets. Drawing on a decade in localization and an increasing focus on AI data, Gemma explores data collection, annotation, multimodal AI, synthetic data, representation, quality, governance, and the enduring importance of human judgment.

From Localized Content to AI Training Data

For decades, localization has largely focused on transforming source content into experiences designed for different languages and markets. Generative AI has changed that relationship. As Gemma explains, localized content is no longer necessarily the finished product. Increasingly, it can become the raw material that AI models learn from. That shift is creating new opportunities for organizations with deep expertise in language, culture, quality, and global delivery.

The challenge for AI companies is also evolving. The question is becoming less about whether a model can perform a task and more about whether it can do so safely, consistently, and appropriately in real-world markets. Organizations may have access to similar underlying models, but they do not necessarily have the same quality of information about how customers actually communicate in Brazilian Portuguese, Arabic, Spanish, or countless other languages, dialects, and cultural contexts.

That makes data a significant competitive differentiator.

Why Better AI Starts With Better Data Collection

One of Gemma's strongest messages in the episode is that data quality begins long before annotation.

“Annotation can't save bad collection,” she explains.

Organizations can label poor or inappropriate data with exceptional accuracy and still end up with the wrong dataset. The key is to begin with a clear specification: What does the model need to learn? Where does it currently fail? Which edge cases need to be represented? Only then should teams decide who to recruit, which environments and devices to include, and under what conditions to collect the data.

Speech provides a particularly clear example. Audio recorded exclusively in quiet environments with high-quality microphones may produce technically excellent recordings, but it may fail to represent users speaking inside cars, in busy environments, with background noise, or while switching languages.

As Gemma puts it, the environment itself can effectively become part of what the model learns. The same problem applies to scripted speech. People naturally pause, restart sentences, use filler words, change direction, and code-switch between languages. Removing those behaviors from a dataset may clean the data while also making the resulting AI less capable of handling real conversations.

Multimodal AI Makes Global Data More Complex

AI systems increasingly need to understand far more than written text. Audio, images, video, speech, gestures, faces, and other multimodal inputs are becoming integral to how people interact with technology. That creates considerable complexity when collecting data internationally. Hardware varies by market. A dataset collected entirely using premium devices may unintentionally overrepresent wealthier users. Video and in-person collection can require local facilities, recruitment, and on-the-ground teams. Voice and facial information may also fall under different biometric and privacy regulations depending on geography.

Language introduces another layer.

A label such as “Spanish” or “Arabic” is not enough to describe a representative dataset. Teams need to consider dialect, register, accent, demographic differences, and the fact that multilingual speakers often switch between languages in normal conversation.

Visual information brings cultural questions too. Concepts such as kitchens, weddings, clothing, food, homes, celebrations, and gestures can look dramatically different worldwide. The underlying challenge, however, is familiar to the localization industry: how do you maintain one quality framework while operating consistently across many languages, cultures, markets, and in-country teams?

Who Is the AI Actually For?

For Gemma, one of the most important questions in developing a dataset is deceptively simple: Who is this system supposed to serve? Organizations need to define that population and then sample against it. Without a clear definition, datasets can gradually become dominated by whoever was easiest to recruit rather than the people who will ultimately use the product. Representation can involve accent, dialect, age, gender, urban versus rural populations, income, disability, atypical speech, skin tone, and the context surrounding visual information.

The danger is building an AI system that performs extremely well “on average” but poorly at the edges. Those edges can include the emerging or fast-growing markets an organization hopes to reach. As Gemma points out, aggregate performance metrics can hide these shortcomings. Companies need to examine performance by individual segments if they want to understand whether an AI product genuinely works for different users.

This is also where decades of localization and transcreation expertise become relevant. Localization professionals have long understood that something can be technically correct and still feel culturally wrong. AI data presents the same challenge in a different form.

Synthetic Data is Powerful, but Not a Replacement for Reality

Synthetic data has become an increasingly prominent part of AI development, and Gemma sees considerable value in it. It can help organizations create volume, model unusual edge cases, work in privacy-sensitive environments, and develop resources for low-resource languages where sufficient real-world data may not exist. But she also cautions against making synthetic data the foundation of an entire system.

Synthetic data is generated by another model, which means it may reproduce that model's biases and blind spots. Used repeatedly, it can gradually pull information toward an artificial middle ground. The problem can become especially visible in multilingual applications, where synthetic non-English material may be linguistically fluent while still carrying an underlying English-centric perspective.

Gemma advocates a hybrid approach summarized in one of the episode's most memorable principles:

“Human seeded, synthetically expanded, human validated.”

Automation can help create scale, but judgment should remain human. She is equally clear about evaluation: synthetic information should not be the ultimate test of whether a system works. Evaluation needs to involve real people and real markets. Otherwise, as Gemma puts it, the model is effectively “grading its own homework.”

What Does High-Quality AI Data Actually Mean?

Quality does not simply mean eliminating errors. According to Gemma, quality means asking whether the data performs the job it was specified to do. That involves several considerations: whether annotators follow instructions correctly, whether independent reviewers reach consistent conclusions, whether the dataset represents intended users, whether the information stays current, and whether every individual asset can be traced back to its origin and permissions. This is another area where localization offers valuable foundations.

Established approaches to language quality already include error categorization, severity weighting, sampling, reviewer calibration, and structured arbitration. Gemma argues that annotation quality assurance is fundamentally a related discipline. Poor AI data can be particularly dangerous because it often fails silently. Traditional software may produce an obvious error when something breaks. A model trained on poor labels can instead deliver an incorrect answer with considerable confidence, making the original data problem much harder to diagnose.

Ambiguous annotation guidelines create another challenge. When two annotators interpret the same edge case differently, simply averaging that disagreement does not solve the problem. Instead, Gemma recommends arbitration by a senior reviewer. Importantly, repeated disagreements can reveal something bigger than an isolated annotation issue: they can expose ambiguity in the guidelines themselves.

That allows teams to fix the underlying problem before the model learns inconsistent behavior.

Governance Can Help Organizations Move Faster

Privacy, security, consent, and governance become increasingly important as AI data programs scale. But Gemma challenges the assumption that governance necessarily slows innovation. In her experience, organizations can move faster when they establish the framework in advance rather than attempting to resolve privacy, legal, consent, and security questions independently for every project.

Consent needs to be captured when data is collected, with clear definitions around what the information can be used for, whether model training is permitted, how long it can be retained, and whether redistribution is allowed. Other requirements can include personally identifiable information detection and redaction, secure working environments, regional data controls, vetted contributors, role-based access, audit trails, and detailed data lineage.

The goal is to be able to determine where an asset came from, what permissions apply to it, and who has interacted with it throughout its lifecycle. For organizations accustomed to serving highly regulated industries, these requirements are not entirely new. They represent an extension of quality, security, and governance disciplines that have existed for years.

Why Localization Expertise Matters in the AI Data Era

Although AI data can look like an entirely new discipline from the outside, Gemma sees it as part of a much longer technological evolution. The localization industry has already worked through machine translation, neural machine translation, language quality evaluation, speech technologies, cultural adaptation, and increasingly sophisticated automation.

AI data brings these capabilities together in a new way.

Language expertise, local-market knowledge, quality frameworks, global contributor networks, cultural judgment, governance, and scalable international program management are suddenly essential capabilities for companies building global AI products. For Gemma, this is less a reinvention than an evolution. The raw material has changed, and the workflows have expanded, but many of the fundamental challenges remain remarkably familiar.

The Future of Global AI Is Human-Informed

The central takeaway from Gemma Newlove's conversation on VistaTalks is that more data is not automatically better data. Effective AI requires datasets designed intentionally around the people, markets, environments, and situations in which a system will actually operate. Real-world data matters. Synthetic data has an important supporting role. Representation needs to be designed rather than assumed. Governance must begin at collection. Quality requires clear specifications and meaningful arbitration.

And throughout all of it, human judgment remains indispensable.

For organizations building AI products for international audiences, Vistatec’s decades of experience managing language, cultural complexity, quality, and global operations may prove more valuable than ever.


Why Better AI Starts With Better Data with Gemma Newlove – VistaTalks Ep 204

Gemma Newlove, Director of Sales at Vistatec, joins Host Simon Hodgkins for a timely conversation about the rapidly expanding world of AI data, multilingual model development, and the critical role localization expertise plays in building AI systems that work reliably across global markets. Drawing on a decade in localization and an increasing focus on AI data, Gemma explores data collection, annotation, multimodal AI, synthetic data, representation, quality, governance, and the enduring importance of human judgment.