← All Posts

What is automated testing? The determinism problem nobody explains

What is automated testing? The determinism problem nobody explains

Every vendor explaining automated testing skips the one concept that actually determines whether it works: a test can only be automated if “correct” has a fixed, predictable answer. A surprising amount of software doesn’t work that way. Once you see that line clearly, most of the confusing parts of automation, what to automate, why suites go flaky, what AI really changes, start to make sense.

The one concept that makes or breaks automated testing: determinism

A function is deterministic if the same input always produces the same output. 1 + 1 always equals 2, with no exceptions.

That matters because of how automated tests work. Every test compares an actual result to an expected one. That comparison only means anything if the expected result is fixed. If the “right” answer moves around, a pass tells you nothing and a fail tells you even less.

Most UI testing is deterministic by design. The same login with the same credentials should always reach the same dashboard, which is exactly why it automates well.

Some things genuinely aren’t deterministic: a personalized recommendation feed, an AI-generated summary, a “does this look right” visual judgment. Pretending they are is where automated suites quietly become unreliable.

What a manual tester’s job actually breaks down into

When a person tests a feature by hand, they are doing four separate things:

  1. Navigating: getting the app into the right state.
  2. Acting: entering data, clicking, submitting.
  3. Observing: noticing what happened.
  4. Judging: deciding whether what happened is correct.

Automation replaces the first three cleanly because they are mechanical. It only replaces the fourth where the correct answer was fixed in advance.

This is the actual boundary of automated testing. It isn’t a list of test types. It is a line drawn by whether judgment was predefined or genuinely exercised in the moment.

Two numbers every automation decision actually comes down to

  • Build cost: what it takes to write the first version of a test. It is usually a one-time cost, and it is the only number most teams budget for.
  • Maintenance cost: what it takes to keep that test working every time the application changes underneath it. It is recurring, it compounds, and it is the number that actually determines ROI.

A suite that costs little to build but a lot to maintain can end up more expensive over a year than one that took longer to build but barely needs touching.

Most public vendor claims you’ll see, such as 85% less maintenance, 95% less maintenance, or 15x faster creation, are claims about the second number, not the first. It is worth knowing what is actually being measured when you see one.

A worked example: one login test, from Day 1 to Day 41

Day 1: the test is written. It has an action (enter credentials, submit), data (a test account stored in a test data management platform), and an assertion (a dashboard element is present). That takes roughly 15 to 30 minutes.

Days 2 to 40: it runs automatically on every deploy and passes silently. This is automation working exactly as intended, and it is invisible because of it.

Day 41: the login page gets redesigned. The button label changes and the dashboard element gets renamed. The test fails, correctly, because something genuinely changed.

The actual cost of maintenance isn’t “the test broke.” It is the 10 to 20 minutes it takes a person to confirm it is a stale selector and not a real bug, then fix it. That repeats every time this page changes, for as long as the test exists.

Where this breaks: when determinism isn’t there to rely on

  • Personalized or dynamic content: a recommendation engine or a live data feed, where “correct” varies by design, not by bug.
  • Subjective visual or UX judgment: “does this look right” doesn’t reduce to a fixed assertion without losing the thing that made the judgment valuable.
  • One-off checks: a migration validation run exactly once has no repeated executions to earn back its build cost.
  • Workflows still in active flux: automating against a page being redesigned weekly means re-paying the Day 41 cost every week.

How AI changes each number, specifically

Build cost: natural language test creation removes the manual authoring step from Day 1. More importantly, it removes the scripting-skill requirement that is usually the actual bottleneck.

Maintenance cost: on Day 41, a self-healing test recognizes the change and updates itself, escalating to a person only when it can’t confidently resolve the change alone. That removes the recurring 10 to 20 minute interruption, not the underlying event. This is what AI test maintenance refers to.

Neither eliminates determinism as a requirement. AI-generated tests still need a fixed expected outcome to check against. What changes is who defines it and who maintains it, not whether it is needed.

FAQs

No. Automated testing is the broader practice of having a machine navigate, act, observe, and check results against a fixed expectation. AI testing is a subset where AI writes or maintains those tests. The expected outcome still has to be fixed either way; AI changes who defines it and who keeps it up to date.
Partly. You can automate the parts that do have a fixed answer, such as checking that a recommendation feed loads, returns the right number of items, or stays within allowed categories. What you can't automate is the judgment of whether the specific output is good, because there is no single correct answer to compare against.
A stable, high-traffic flow with a clear expected result, such as login or checkout. It runs on every deploy, so it earns back its build cost quickly, and because the page rarely changes it keeps maintenance cost low.
━━━━

See how Klarent can achieve 90%+ test coverage in just two weeks.

Book a free 30 minute call
Dyuwan Shukla, Product Manager at Klarent

Written by

Dyuwan Shukla

Product Manager at Klarent

View full bio

More blogs