End-to-end test data management with icaria TDM.
Based on icaria Technology's Masterclass series.
Understanding what TDM is, why it matters, and the real challenges QA teams face today.
Test Data Management (TDM) is the set of mechanisms and processes that ensure test data is realistic, secure, representative, consistent, and readily available. It goes far beyond simply copying production databases — it means designing a coordinated system that makes testing repeatable and safe.
TDM is more than having masked copies of production data. It covers how test data is generated, how it is validated, and how it is provisioned in a way that is both functionally and technically consistent — and irreversible. Masking protects; TDM orchestrates, enabling teams to test with secure and consistent data as many times as needed, without bottlenecks and with reliable results.
Poor test data comes at a high cost and has a direct impact on software quality. The numbers speak for themselves:
QA teams deal with multiple obstacles on a daily basis:
A paradigm shift: treating test data with the same discipline you would apply to a software product.
Modern TDM treats test data with the same discipline applied to a software product: with clearly defined consumers, ownership, metrics, a roadmap, and a polished consumption experience.
The goal is to move test data from something someone copies to a stable, observable service with measurable value.
If your team depends on an SQL expert to deliver or create data, you do not have a product — what you have is a recurring favour with high risk.
The product is ready-to-use test datasets that each of these consumers can request on demand. The catalogue includes:
Compliance by design means that privacy and regulatory requirements are embedded directly into the test data pipeline — not as manual checks at the end, but as automatic, measurable, and auditable steps.
The result: datasets that are useful for QA, free of real data exposure, repeatable, and backed by evidence that things are being done right.
Consistency is not the same as copying production — it is not about cloning everything. It is controlled equivalence: a test that passes in QA should carry the same meaning in UAT (User Acceptance Testing) and in production.
A modern TDM system relies on four workflows that together guarantee reliable and visible data.
This is what enables a QA team or a CI/CD pipeline to obtain a dataset that is ready for testing.
Instead of spinning up entire environments or asking the DBA (Database Administrator) for a favour, you request a pre-built, masked, consistent dataset that the TDM system delivers.
Enables you to re-run tests with the same data or roll back if a new release breaks something.
Every time a dataset is generated or delivered, the system must leave a trail: logs, metrics, full audit history.
Protects data from the source, ensuring privacy is built into every step of the cycle — not bolted on afterwards.
Masking rules are applied automatically whenever data is extracted or generated.
Means being able to see, measure, and audit what happens with your test data.
Every dataset carries logs, metrics, usage statistics, and full traceability.
If I need 200 active customers with valid orders, the TDM system delivers them directly — no environment downtime, no third-party requests. On-demand provisioning turns test data into a service, not a manual task.
The habits we keep without realising it — and that prevent us from building a serious test data system. Overcoming them is the first step toward TDM maturity.
It sounds efficient, but it creates legal and compliance risks. Realism should not be confused with reusing real data. In many cases the necessary infrastructure or environment security for such a copy simply does not exist.
Alternatives:
This leaves you dependent on one or two people who know the data model — and you are stuck when they are unavailable.
This costs you days every time the environment is reset. If every data refresh wipes out your tests, the data is not really yours. With TDM, a data refresh stops being a disaster and becomes a push-button operation.
Nobody knows who owns the test data. Is it QA? The DBA? Security? Infrastructure? The result is that everyone touches it but nobody governs it. Priority conflicts arise because QA demands more data, Security blocks it, and Infrastructure complains about the volume.
Tests that sometimes pass and sometimes fail with the same code because the data has changed, does not meet preconditions, or is inconsistent. If a test passes or fails depending on how the environment was set up, it is not a reliable test.
Solution: Deliver consistent data with referential integrity and traceability.
The team relies on one or two people who know how to query the database and prepare data with their own scripts. That is tacit knowledge rather than documented process. You cannot afford to wait days to get the data you need for testing.
Solution: An archetype-based search system with self-service capabilities that empowers QA teams to request data with full guarantees.
When building an environment or preparing data takes days or weeks, you are effectively losing the sprint waiting for everything to be ready. If it takes days to get data, your bottleneck is not testing — it is the data.
Production data is used in test environments with no controls. Common reasons:
The smallest set of components you need to demonstrate real value in test data management. Just as in development we build an MVP to validate an idea, in TDM we can build a minimum viable architecture to validate the model.
This architecture is not meant to be a throwaway pilot. It is the first step of a system that grows based on evidence. It is a reproducible and measurable system in miniature — one that already demonstrates the benefits of TDM.
Test data is not just managed — it is designed. And designing it well turns testing into reliable knowledge. The TDM mechanisms work together to make this possible.
Data masking is the transformation of sensitive data through controlled rules that preserve the technical value of the data so it remains usable. It is the cornerstone of compliance by design.
It involves transforming or replacing sensitive values (names, national IDs, phone numbers, health-related information) with alternatives that retain the original format and logic without allowing the identification of real individuals.
Performs a one-time replacement of data values and, once the process is complete, destroys any evidence of the anonymisation. It always uses non-deterministic algorithms so that values cannot be reversed to identify individuals.
This approach is more common for ephemeral environments or when data needs to be delivered to third parties — the data is anonymised today and any link to the original information is permanently severed.
Retains a protected mapping. The masking process runs and stores — in a fully secured manner — the fact that ID number 1 was changed to 5, that bank account X was changed to Y, and so on (individual data points that cannot be linked back to a specific person).
Scan the database — every table, every column, the actual content — to identify where sensitive data resides. This is done through inspectors that apply:
The output is a heat map (very high, high, medium) indicating the likelihood that each column contains a specific type of sensitive information.
Synchronise the table metamodel in the application and assign specific masking algorithms to each column:
Select the target environment (test, development, integration, pre-production) and run the process. Steps execute in a specific order, with parallelisation where possible:
The process can be automated via an orchestrator (e.g. over the weekend the production database is copied and the next step is a bulk masking run). icaria TDM provides a full REST API layer for these executions.
A masking process can target databases of different technologies simultaneously. If the environment comprises a Postgres database, an Oracle instance, and a DB2 system, the same process can address all of them at the same time — seamlessly.
All icaria TDM processes can run simultaneously across multiple applications and environments, completely agnostic to the underlying database technology.
How do we ensure that data masking does not degrade the quality or usefulness of the test dataset?
records anonymised across our customer base. icaria TDM can parallelise as long as there are available CPU and memory resources.
Subsetting extracts a coherent subset of data from production or a test environment and delivers it to another environment while preserving full referential integrity and all table relationships.
Rather than copying entire databases, subsetting selects a representative set by extracting only the entities you need.
reduction in database volume through intelligent subsetting.
A production incident occurs involving a specific customer, and you need to reproduce it.
With traditional bulk masking: Perhaps the data exists because a copy was made last week, but it is already stale. The developer confirms the bug is reproducible... and now the data is inconsistent and retesting becomes impossible.
With subsetting: The moment the issue is detected, you take the data from the source environment (production) and deliver it to another environment with masking applied. You get the data you need, when you need it. It is repeatable and takes seconds.
Repositories allow you to save a snapshot of a customer (or as many as needed) in multiple versions, independent of what happens in the source environment.
Even with "unlimited" data (a full copy of production), the question remains: how do you find exactly the data you need in this ocean of records? Out of millions of customers, how do you locate the ones that actually satisfy your test requirements?
"My test needs a newly created customer." But what conditions define "newly created"?
Managing data through queries comes with multiple challenges:
Archetypes (or data profiles) are standardised definitions that everyone agrees on. They are categorisations of data profiles built from a set of conditions that define the data — where everyone shares the same, standardised understanding.
Using archetypes, a self-provisioning data pool is created: data matching the archetype definition is automatically located across the available test environments.
If the existing archetypes do not cover what you need, you can create a search request from scratch:
Requests are processed in batches to properly manage database access. Organisations automate execution in scheduled time windows that do not impact their environments.
Self-service completes the process and allows you to request dataset provisioning directly into the environment where it is needed — without raising tickets or waiting on other teams.
QA and DevOps gain real autonomy, and the testing workflow accelerates exponentially. It breaks the dependency on specific individuals or teams and turns test data into an on-demand service.
autonomy. With the appropriate training, teams can manage the entire cycle themselves.
From the self-service user's perspective:
Log in to icaria and request the data delivery
Choose the data structure (e.g. customers)
Select the target environment (test, development, repository)
Log the reason for audit purposes (can include a ticket number or test case ID)
Select the seed or search criteria to use
The background queue processes pending requests
This is the fifth mechanism and it closes the loop. The concept shifts: it is no longer the tester who searches for or decides which data to use — instead, the data is already designed as part of the test case itself. It is the starting point for building and executing automated tests.
Each test carries an associated dataset that meets the exact conditions it requires (users, customers, orders, statuses, dates), and the system provisions it automatically before execution.
Typically, managing this data is not included in the plan — it is neither part of the test definition nor the automation setup. That is why we start well with automation (Selenium scripts are straightforward to write) but things fall apart as tests grow in complexity.
From a data perspective, the design includes:
Locate the data in production that meets the required conditions
Move the data to an internal, protected database so no one can modify it
On the test's request, deliver the data to the appropriate environment
Once the data has been consumed, restore it to its original state for reuse
Run the test and verify the outcome from the data perspective
Validation rules define the expected outcomes from a data perspective:
Pre-delivery rules ensure data quality before it reaches the test environment.
If the test requires the user to be under 18: icaria TDM reads the date, calculates how old the person would be at the time of delivery, and adjusts it to meet the condition. The delivered date of birth becomes today minus 17 years, ensuring the person is always 17 at the time of delivery.
To deliver consistent data, a data domain is defined: what needs to be delivered (e.g. a complete partner record with at least one quotation in draft status across both applications).
The subsetting structure has a header entity (e.g. partner) with all its required related entities to complete the domain across every application involved.
TDM maturity is a journey, not a leap. icaria TDM has identified five levels, from reactive to continuous improvement.
This model does not expect every organisation to reach the highest level. Its purpose is to help you identify where you stand, what maturity level each process has reached, and what the next logical step looks like.
| Level | Characteristics | Indicators |
|---|---|---|
| 1 — Initial | Test data is managed reactively, manually, ad-hoc. Direct copies of production with no controls. | No catalogue, no masking, reliance on "local heroes" |
| 2 — Repeatable | Some basic masking processes exist. Sensitive data has been identified in the main systems. | First sensitive data map, basic masking rules |
| 3 — Defined | A Minimum Viable Architecture has been implemented. Automated provisioning is in place for at least one data domain. | Automated pipeline, functional subsetting, defined archetypes |
| 4 — Managed | TDM integrated into CI/CD. Self-service operational. Usage metrics and lead time visible. | Data as a service, full traceability, multiple domains |
| 5 — Optimised | Continuous improvement driven by metrics. Test data designed as part of the test case. Compliance by design across all environments. | 100% coverage, measured ROI, zero real data in non-production |
The goal is not to do everything at once, but to demonstrate value quickly, learn, and then scale. The two levers that drive this are:
By incorporating icaria TDM into the testing ecosystem, you complete the continuous testing pipeline.
Testing teams commonly use tools such as:
Yet data management tooling is rarely part of the conversation. Adding icaria TDM completes the ecosystem.
The orchestrator (e.g. Jenkins) asks the test case manager which test to run and in which environment
The orchestrator calls icaria TDM to provision the data into the target environment
Once icaria TDM finishes, the orchestrator triggers the execution tool to run the test
After the test, the orchestrator asks icaria TDM to verify results from the data perspective
All information is consolidated into a report and the results are stored with data from every tool in the chain
When a test case is triggered, the system automatically:
This eliminates the typical bottlenecks in data preparation and synchronisation phases, enabling parallel test execution without conflicts or dependencies. It is the final step toward continuous testing, where data and tests advance at the same pace as development.
The measurable impact of TDM on speed, coverage, costs, risk, and team satisfaction.
Improved communication within and across teams:
We couldn't live without icaria (TDM)
Data management specialists. We help ensure data is governed, secure, and usable.
icaria Technology is a company specialising in data management. It helps organisations design and operationalise their data — from data governance and the protection of production environments through the automated enforcement of privacy rights (such as the right to erasure), to the reproducible provisioning of test data.
icaria TDM is recognised by Gartner as a Sample Vendor in their 2025 report on three steps to optimise test data management.
TDM adoption, according to Gartner, is still at a very early stage in many organisations. Most companies — even large ones — still rely on manual processes. But TDM technology has matured: what the tools deliver matches what users expect.
icaria TDM is the Test Data Management platform built for mission-critical applications and OSS/BSS environments. It delivers realistic, secure, accurate, and consistent data exactly when and as many times as it is needed — ensuring that every test result is effective and reliable.
With icaria TDM, testers and developers dramatically reduce the hours spent producing and managing data, allowing them to focus on higher-value tasks. Leading companies in banking, insurance, and telecommunications already trust icaria TDM to transform their testing processes.
Security and compliance without compromise. Applies advanced anonymisation and pseudonymisation techniques to ensure data is representative without exposing sensitive information. GDPR-compliant and consistent across multiple environments.
Data ready for manual testing. Testers gain autonomy by obtaining optimised subsets on demand, finding the exact data for each test case through the data finder, and extracting datasets while preserving integrity and relationships.
Automated data for continuous testing. Automates data provisioning in CI/CD pipelines, eliminating bottlenecks. Every test run gets the right, up-to-date data — no manual intervention required.
Compatible with leading platforms — Oracle, SAP, Salesforce, IBM, Hadoop, and many other data sources — icaria TDM integrates into any technology stack, delivering a robust and scalable solution. QA teams test more, test better, and test faster.
Installed on the customer's own infrastructure (cloud or on-premise). icaria TDM is privacy-aware by design: the data belongs to the customer and stays within their infrastructure.
Getting a TDM project up and running requires:
The project takes a few months, but the return on investment is visible very quickly — even during the implementation itself, data quality issues are uncovered and smaller, leaner data deliveries already provide tangible benefits.
Large organisations with complex structures, demanding regulatory compliance requirements, and very large data volumes trust icaria TDM: banks, telecommunications companies, insurers. The platform scales from smaller teams working with hundreds of tables to organisations with thousands of tables, multiple technologies, and diverse teams.
Discover how icaria TDM can help your organisation cut costs, increase test coverage, and achieve regulatory compliance — from day one.
Request a Demoicaria Technology — Data Management Specialists
© 2026 icaria Technology. All rights reserved.