AI in test automation: A customer project shows what is possible

6.8.2026

Five challenges that AI has only made visible

How many hours in your test automation project are spent on maintenance instead of new coverage? And how much of the knowledge that holds your project together resides in the minds of only two or three people?

We asked ourselves these questions in a customer project in the insurance sector: several hundred automated business transactions across multiple clients, tested from the user interface to the database, in a regulated environment where every test run influences release approval. When we implemented Claude Code, five problems came to light that had been holding us back for a long time, but no one had openly addressed them before.

Initial situation

We are responsible for the end-to-end test automation of this inventory management platform: test suites that check the entire chain, from the user interface through the business logic and interfaces to the database. The results must be documented in an audit-proof manner.

A project like this grows. And with it grow things that are rarely talked about because they creep up on you. It was precisely these issues that actually cost us time. Not writing new tests.

The first realization was therefore uncomfortable: An AI assistant doesn't automatically make a project faster. It amplifies what's already there – existing order and existing chaos alike.

Pain Point 1: Excessive dependence on individual people

The problem: In every organically grown test project, there are rules that aren't written down anywhere. Why certain test suites shouldn't run in parallel. Why an automatic retry at a specific point does more harm than good. Why an error message from test management almost never means what it says. This knowledge was held by only two or three people in our company.

Every team knows the consequences. Onboarding new colleagues takes a long time. Mistakes are repeated because no one has documented them. And the project depends on the availability of individual people, which is a real risk in a consulting environment.

How we solved it: We had to explain to the assistant how the project works. And to explain it to him, we first had to clearly define it ourselves. This necessity led to the creation of a set of rules, which is stored in the repository, versioned, and visible to everyone.

The result: The biggest gain wasn't the assistant. It was the documentation we wrote because of him. These rules are now the first thing new team members read. They answer precisely the questions you'd otherwise ask in the third week, after you've already made the mistake.

Pain Point 2: Green tests that don't check anything

The problem. This is the most uncomfortable topic in test automation, and it's rarely addressed openly. A broken test turns red and stands out. A bad test turns green.

If a test is applied to an empty result set, it will run. The test remains in the suite, consuming runtime and creating unjustified confidence. For release decisions, this is worse than no test at all, because nobody checks.

This is precisely where the real risk of AI support in testing lies. A language model optimizes for plausibility, not correctness. It delivers a convincing-sounding answer, even if it cannot know the correct one. In our case, this concerned internal business keys that cannot be derived from what is displayed on the surface, even though it appears to be.

How we solved it: We formulated a rule that doesn't affect code, but rather behavior. A value is either inherited from a comparable, already defined test case, or it isn't set but marked as an open question and verified on the first run. A value that triggers a green light in a test without being defined is considered a defect.

The result: The approach shifts from always providing an answer to highlighting where the answer is missing. An honest red test is worth more than a deceptive green one. This applies regardless of AI, by the way. The assistant simply forced us to finally write it down.

Pain Point 3: Maintenance costs instead of further development

The problem: In many automation projects, the balance eventually tips. The team spends more time keeping existing tests alive than building new coverage. Two causes dominate: Unstable access to user interface elements that break with every deployment, and duplicates because existing components aren't found and are therefore built a second time.

Recording tools exacerbate both problems. They preferentially suggest precisely those access methods that work at the moment of recording and will break down most quickly afterwards.

How we solved it: We defined a binding hierarchy specifying which access method should be used in which situation, and explicitly stated which are prohibited. In addition, we implemented a verification requirement. Before a new component is created, a four-step check is performed to determine if it already exists.

The result: The rules have a dual effect. The assistant adheres to them, and the team has, for the first time, a written standard for reviews. Previously, maintainability was a matter of experience. Now it's verifiable.

Pain Point 4: Troubleshooting that goes in the wrong direction

The problem: A test run fails. The report states that an object was not found. So, a search is made for the missing object. In fact, an access token had expired, and the interface is deliberately providing a misleading response for security reasons.

In addition, we encountered a self-inflicted problem. In older code, the technical error trail was discarded before the error was reported. The reporting then showed an error, but without any information about its origin. Every analysis started from scratch.

Both of these things together regularly cost hours. And an assistant without this knowledge very efficiently searches in completely the wrong direction.

How we solved it: The technical fault trace remains intact; that's a strict rule these days. And the known misleading error patterns are included in the rulebook, each along with the specific diagnostic step that clarifies the true cause in one minute.

The result: recurring hours became minutes. The real progress, however, is that this knowledge is no longer lost when the person who possessed it is on vacation.

Pain Point 5: Compliance and Confidentiality

The problem: In a regulated environment, the question of what can be copied into a chat window is not an academic one. Someone investigating a login error might, in case of doubt, copy the entire configuration into a tool because speed is of the essence. The risk lies in human behavior. This cannot be technically reconfigured away.

How we solved it: Clear, explicitly stated guidelines. No access data or configuration content in prompts, logs, tickets, or commit messages. No circumvention of established version control validation mechanisms without explicit instruction. No modifications to shared core components without assessing the impact on the entire test suite.

The result: Anyone introducing AI support in a regulated environment needs these points in writing, not as a verbal agreement. They are also the first things that will be asked about in an audit.

The central finding

An AI assistant is only as good as the context you give it. And that context doesn't happen by itself.

The measurable time savings for us occurred where a clear rule-based foundation existed. This included creating new test cases based on existing patterns, translating documentation into maintainable code, and recurring structural work. Where this foundation was lacking, rework was required, sometimes more than with manual creation. That's the honest truth, and it's why we discuss structure first and then tools.

It's noteworthy that none of the five points above are an AI problem. Undocumented knowledge, deceptively positive test results, maintenance backlogs, misleading error patterns, and unclear confidentiality rules already existed. The assistant simply made them visible and urgent.

What this means for your project

If you're considering bringing AI support into your test automation, these are, in our experience, the right starting questions to ask:

- Where in your project could something go wrong without anyone noticing?

- What knowledge resides with individuals and nowhere else?

- How much of your capacity is spent on maintenance instead of new covers?

- How can you tell from a review whether a test actually checks what its name claims?

- What is an employee allowed to enter into an AI tool, and where is this written down?

Anyone who can answer these questions has already done most of the work. Finding the right tool is then the easier step.

- Do you want to know how this can be applied to your situation? Talk to us.

‍

In a follow-up article, we will examine this topic from a compliance perspective. What evidence does an auditor expect when AI support is used in the testing process?.

We are ready for your next step!

Would you like to utilize our expertise and implement technological innovations?

This website
uses cookies

Cookies are used for user guidance and web analytics and help to improve this website. You can view our cookie policy here or adjust your cookie settings here . By continuing to use this website, you agree to our cookie policy.

All accept
Accept selection
Optimal. Functional cookies to optimize the website, social media cookies, cookies for advertising purposes and the provision of relevant offers on this website and third-party websites, as well as analytical cookies to track website visits.
Limited functionality. Several functional cookies are used for the proper display of the website, e.g., to save your personal settings. No personal data is stored.
Back to overview

Speak to an expert

Do you have a question or are you looking for more information? Provide your contact information and we will call you back.