
Who will recognize the following situations?.
A new test iteration is due. To run all the tests, I need certain test data configurations, such as customers in different statuses (prospect, active, former, etc.). This starts the (sometimes lengthy) search for suitable test data. Only after I've found the right test data can the test iteration begin. The effort required to find suitable test data thus reduces the time available for testing.
After a new deployment, the automated regression tests run. The result shows that 20% of all test cases failed. I just think, "What the f***?" and start analyzing. It turns out that the tests couldn't be executed due to incorrectly configured master data or missing data (keywords: false negatives, false positives). I look at the clock and realize that I've now spent an hour figuring out that the test data is (once again) to blame for the failed tests (e.g., 20 tests * 3 min analysis = 60 min).
Anyone who recognizes themselves in these situations will only think: "Test data management will rule the world".

Let's take a step back and analyze the possible causes of test data problems. I can classify the causes of errors and problems I've encountered in my testing life into the following categories. (Note: This is not an exhaustive list, and the order does not reflect priority/importance):

testing (especially test automation) using production data is, in my view, doomed to failure. Production data typically has a limited lifespan. Furthermore, according to the GDPR, testing with production data is not permitted. Therefore, testing should be based on synthetic and/or anonymized data.
The complexity of the required test data primarily stems from the fact that the test data should be as close to production as possible. For example, testing a reporting functionality or a history function requires many different data sets. Replicating this with synthetic test data can be very complex. In contrast, verifying a change functionality, such as an address, is less complex to replicate with synthetic test data. The effort required to define high-quality, production-like synthetic test data should not be underestimated!
The amount of test data is usually considered in relation to the type of test. For example, performance tests (non-functional tests) require a large amount of test data, but only a few test data configurations. Regression tests, on the other hand, require a large number of different test data configurations, albeit in small quantities.
The limited lifespan of test data is, in my view, one of the major and most important challenges for stable test automation. It's not enough to create the test data once and reuse it with every test execution. The test data is consumed during test execution and/or is subject to aging. For example, if a customer's status in a test case changes from "prospect" to "active," that test case can no longer be executed with the same customer.
The fact that the test data is environment-dependent is self-explanatory; that is, it is not sufficient to create the test data once in one environment, e.g., the acceptance environment. The test data must be created in all environments on which tests are executed.
Time travel of test data may be necessary when functionality needs to be tested based on a specific date. In the banking sector, this would be, for example, month-end or year-end processing.
Specifications for test data, whether legal, industry-specific, or company-internal, play a central role. The test data must comply with these specifications. This is particularly important when the GDPR applies. Under the GDPR, personal data may only be used for specific purposes. Testing is not included in this purpose and means that testing with production customer data is not permitted (unless the data subject's consent is obtained). Failure to comply can be costly, with fines of up to €20 million plus potential claims for damages.
In addition to these points, good test data is useless if it is not provided in the right scope, in the right environment, at the right time, and in the right quality .
Provided the points listed above are/have been observed, nothing stands in the way of successful testing. If this is not the case, risks can arise that must be managed. These risks can be classically divided into the following categories:
A product risk exists, for example, if the test data does not meet the technical requirements and/or does not cover all business-critical combinations. In this case, undiscovered errors could end up in production.
A project risk arises, for example, if the test data is not defined or provided in a timely manner. This leads to delays in the project.
A particular business risk lies in non-compliance with legal and industry-specific regulations. As mentioned above, this can be costly and pose a significant risk to the company.
Now that we understand the problems, challenges, and risks of lacking test data management, my next blog post will focus on developing a suitable test data strategy and concept. So stay tuned.
March 27, 2020, Frederic Hesse
We take a holistic approach to software quality and, drawing on our many years of experience and continuous professional development, provide support where we can offer the greatest added value for you and your project. In doing so, we deliberately break free from the constraints of traditional project roles and focus on skills that benefit your software quality across the board.
> Learn more
Would you like to utilize our expertise and implement technological innovations?


Do you have a question or are you looking for more information? Provide your contact information and we will call you back.