Synthetic test data powered by GenAI

As your Swiss specialist for test automation, test data management, and artificial intelligence, we love exploring the latest technologies and possibilities. One of these new possibilities is the generation of synthetic test data using the power of generative artificial intelligence (Generative AI). Our Infometis communities proactively invest a lot of time in future-oriented trends and technologies, which is why we analyzed this topic in more detail as a team during a joint hackathon.

Internal exchange regarding synthetic test data

‍

The importance of synthetic test data for modern test data management

Challenges in test data management

When we talk about testing, whether automated or manual, we inevitably have to talk about test data management. Efficient testing is only possible with the corresponding test data. The rapid provision of adequate test data is therefore an essential success factor for all testing activities.

Especially in complex enterprise environments, providing test data or finding suitable data constellations is an enormous challenge for companies and often a difficult-to-penetrate jungle of dependencies, exceptions and legacy issues.

When traversing this jungle, one often encounters snares:

  • A specific constellation of data for testing a new function or for regression testing is not available
  • Insufficient data for load and performance testing
  • Data consumed after test runs
  • Lack of integration into the CI/CD pipeline
  • Data protection requirements for LLM models that need to preserve data structures
  • This results in a high total cost of ownership.

‍

Traditional solutions and their limitations

Current solutions to these challenges include:

  • Cloning production data to test environments (🫣)
  • Transfer production data, masked, pseudonymized or anonymized, into test environments

These approaches entail the following risks and disadvantages:

Data protection risks and cybercrime: With increasing regulation in data protection, purpose limitation and access restrictions, especially for PII (Personally Identifiable Information), are major concerns. Using production data for testing purposes, even masked or anonymized, poses a significant data protection risk and increases the attack surface for cybercriminals.

Data unreliability: The efficient execution of regression tests, whether functional or non-functional, relies on reliable test data. Data based on production systems can contain unintended configurations that negatively impact test results. The resulting error analysis may deviate from the actual test objective. Furthermore, a higher degree of anonymization reduces the suitability for broader test campaigns.

Structure: To identify edge cases and simulate realistic scenarios during load and performance testing, it is crucial that the structure of the data, including time series and transactions, closely mirrors the production data. Anonymizing data or generating mock data alters the structure or fails to capture its full complexity. This drawback is also critical when training proprietary AI models.

High manual effort: The effort required to provide anonymized datasets is usually very high. This ties up resources and reduces the availability of the test systems. Furthermore, if data is provided irregularly, the test environments quickly become outdated, and data quality deteriorates.

More than anonymization: Synthetic test data

Synthetic test data, the artificial simulation of real-world data, offers the significant advantage of completely decoupling it from production data while simultaneously preserving complex data relationships. This counteracts the disadvantages of traditional test data in terms of data protection and structural representation.

Tools like iSynth help to define and maintain synthetic test data as code and to deploy it consistently in test environments in any variance and quantity

Platforms like Tonic.ai or k2view offer a lot of great tools to automatically identify, transform, and generate synthetic data from PII in production data sources.

Automation frameworks such as Tosca, TAMI , or GenRocket benefit from an organization's ability to create and maintain automated regression tests, including the necessary data constellations.

However, what remains is the effort required for analyzing and maintaining the structure of the real data, deriving the rules and algorithms, and implementing them.

The breakthrough of artificial intelligence , and in particular generative AI, could support us in precisely this area.

The solution? Use AI to generate synthetic test data

AI tools and platforms for generating synthetic test data utilize the capabilities of generative AI to learn the structure and relationships of production data and to generate synthetic test data with the resulting model.

Infometis Synthetic Data - Mostly AI Docs
Source: What is synthetic data? – MOSTLY AI Docs

This delegates the time-consuming task of analyzing and documenting structures and dependencies to AI, capturing them in a data-specific AI model. Using this completely anonymous model, the AI ​​can then, in a second step, generate the desired amount of synthetic data on demand within the desired environment.

Market analysis: Tools for synthetic test data with generative AI

At Infometis, we continuously monitor the market, analyze new developments, actively engage with innovative and market-leading technologies, and maintain close contact with manufacturers. This ensures that we can always offer our customers the best solutions.

For generating synthetic test data with AI, we took a closer look at the following two platforms. They are explicitly specialized in generating synthetic test data using generative AI.

The goal was to develop a sense of where we currently stand and where the journey could lead.

Mostly AI

Infometis Synthetic Test Data GenAI - Mostly AI

‍

‍

At the heart of Mostly AI are the generators. These can be fed with data via connectors or file uploads. In the next step, a data-based AI model can be trained, which in turn forms the basis for generating synthetic data.

Using a wizard (GPT), you can interact with a generator and, for example, analyze data tables in more detail. It is also possible to generate on-demand test data or analyze a data file independently of the generators via a prompt

Mostly AI impresses with its low barrier to entry. With just a few clicks, an AI model and synthetic data can be generated from a CSV file containing data. Users can also gain comprehensive insights into the data's statistics if desired.

‍

Infometis Synthetic Test Data GenAI Mostly AI - Cloud
Mostly AI Cloud

‍

Furthermore, Mostly offers a Python CLI (Command Line Interface) and can therefore be integrated into CI/CD workflows and other automations. This allows, for example, the on-demand generation and provision of data for automated tests. All generators set up in the cloud can also be represented as Python code.

‍

Infometis Synthetic Test Data GenAI - Python CLI
Python CLI

‍

Gretel

Infometis Synthetic Test Data GenAI - Gretel Blueprint Model
Gretel Blueprint Model

‍

Gretel is organized into projects. Within a project, I can train multiple models with different data. Gretel provides so-called blueprints for different use cases to facilitate this.

‍

One project, multiple models

‍

The Navigator allows me to generate on-demand test data via a prompt , similar to the assistant in Mostly AI. I can also add columns to existing tables according to specific rules

With Gretel, it's very quick to achieve initial results. Blueprints allow different use cases to be applied to the source data. I can either train a model that generates synthetic test data or one that classifies my data according to PII (Personally Identifiable Information).

For each model, Gretel generates a quality report with a wealth of information. Gretel also offers a Python CLI. More precisely, two: a low-level and a high-level Gretel SDK.

Gretel's workflow feature is worth mentioning. It allows you to configure your own automations, such as creating 500 test records every Friday at 8:00 AM.

‍

Gretel Scheduled Workflow Feature
Scheduled Workflow Feature

Deployment & Integration

We tested both tools in the cloud using the credits available in the free model. In most business scenarios, it's crucial that production data remains within the company context.

Both tools therefore offer the option of deploying them in a Kubernetes cluster under the company's control. This allows the AI ​​model to be trained with production data within the company's own context.

⚠️ The trained AI model in both tools no longer has any connection to the original data. It can therefore be used to generate synthetic data in the cloud.

Gretel is still limited to the three major cloud providers, while Mostly AI also offers a Helm chart for deployment in a self-hosted environment.

Possible uses:

  • Mostly AI: Ideal for companies that prefer self-hosting to keep sensitive data local.
  • Gretel.ai: Effective for cloud-based, automated data delivery in regulated industries.

Next Steps

Test data has always been one of the major challenges for efficient testing. Increased data protection requirements, growing cybercrime, and increasing industry regulations further increase the complexity.

Test data management, and in particular the provision of synthetic test data, will become even more important in the near future than it already is today. The technology is also advancing rapidly.

The growing capabilities of AI solutions for generating synthetic data come at just the right time. However, this doesn't solve the major challenges of providing adequate test data overnight. In our view, building AI expertise and assessing technological progress are central to sustainable and secure strategies for providing test data.

As your Swiss partner for software quality and automation, we are always on the pulse of the latest market trends and continuously incorporate our findings into our consulting and support.

Do you want to update your test data management? We look forward to an initial discussion.

‍

Training on this topic

Show all
No items found.

We are ready for your next step!

Would you like to utilize our expertise and implement technological innovations?

This website
uses cookies

Cookies are used for user guidance and web analytics and help to improve this website. You can view our cookie policy here or adjust your cookie settings here . By continuing to use this website, you agree to our cookie policy.

All accept
Accept selection
Optimal. Functional cookies to optimize the website, social media cookies, cookies for advertising purposes and the provision of relevant offers on this website and third-party websites, as well as analytical cookies to track website visits.
Limited functionality. Several functional cookies are used for the proper display of the website, e.g., to save your personal settings. No personal data is stored.
Back to overview

Speak to an expert

Do you have a question or are you looking for more information? Provide your contact information and we will call you back.