icaria TDM
Strategic & Practical Guide

The Complete Guide
to Test Data
Management

End-to-end test data management with icaria TDM.
Based on icaria Technology's Masterclass series.

Table of Contents

01

Introduction to Test Data Management

Understanding what TDM is, why it matters, and the real challenges QA teams face today.

1.1 What is Test Data Management?

Test Data Management (TDM) is the set of mechanisms and processes that ensure test data is realistic, secure, representative, consistent, and readily available. It goes far beyond simply copying production databases — it means designing a coordinated system that makes testing repeatable and safe.

Key Concept

TDM is more than having masked copies of production data. It covers how test data is generated, how it is validated, and how it is provisioned in a way that is both functionally and technically consistent — and irreversible. Masking protects; TDM orchestrates, enabling teams to test with secure and consistent data as many times as needed, without bottlenecks and with reliable results.

1.2 The current problem

Poor test data comes at a high cost and has a direct impact on software quality. The numbers speak for themselves:

80% of testers say creating test data is extremely difficult
90% of companies believe data privacy regulations impact their processes
100x more expensive to fix a bug in production than at design time
~50% of a tester's time is lost waiting for data
<20% of test cases are automated due to data issues
Many test failures are caused by the data, not the code

1.3 The challenges QA teams face

QA teams deal with multiple obstacles on a daily basis:

02

Modern TDM: test data as a product

A paradigm shift: treating test data with the same discipline you would apply to a software product.

2.1 The paradigm shift

Modern TDM treats test data with the same discipline applied to a software product: with clearly defined consumers, ownership, metrics, a roadmap, and a polished consumption experience.

The goal is to move test data from something someone copies to a stable, observable service with measurable value.

Think About This

If your team depends on an SQL expert to deliver or create data, you do not have a product — what you have is a recurring favour with high risk.

2.2 Defining the product

2.2.1 Who are the consumers?

Functional QA
Needs datasets that satisfy business test cases
Automation (CI/CD)
Needs reproducible datasets
Performance
Needs representative volumes of clean data
Security
Needs PII masked and traceable

2.2.2 What is the offering?

The product is ready-to-use test datasets that each of these consumers can request on demand. The catalogue includes:

2.2.3 What do we measure?

Lead time
From request to a ready dataset
Integrity and consistency
Are we delivering end-to-end consistent data?
Privacy
Personal and sensitive data properly de-identified

2.3 Compliance by design

Compliance by design means that privacy and regulatory requirements are embedded directly into the test data pipeline — not as manual checks at the end, but as automatic, measurable, and auditable steps.

The result: datasets that are useful for QA, free of real data exposure, repeatable, and backed by evidence that things are being done right.

Key Insight

Consistency is not the same as copying production — it is not about cloning everything. It is controlled equivalence: a test that passes in QA should carry the same meaning in UAT (User Acceptance Testing) and in production.

03

The four core TDM workflows

A modern TDM system relies on four workflows that together guarantee reliable and visible data.

1. On-demand provisioning

This is what enables a QA team or a CI/CD pipeline to obtain a dataset that is ready for testing.

Instead of spinning up entire environments or asking the DBA (Database Administrator) for a favour, you request a pre-built, masked, consistent dataset that the TDM system delivers.

2. Traceability

Enables you to re-run tests with the same data or roll back if a new release breaks something.

Every time a dataset is generated or delivered, the system must leave a trail: logs, metrics, full audit history.

3. Masking by design

Protects data from the source, ensuring privacy is built into every step of the cycle — not bolted on afterwards.

Masking rules are applied automatically whenever data is extracted or generated.

4. Dataset observability

Means being able to see, measure, and audit what happens with your test data.

Every dataset carries logs, metrics, usage statistics, and full traceability.

Practical Example

If I need 200 active customers with valid orders, the TDM system delivers them directly — no environment downtime, no third-party requests. On-demand provisioning turns test data into a service, not a manual task.

04

Anti-patterns and warning signs

The habits we keep without realising it — and that prevent us from building a serious test data system. Overcoming them is the first step toward TDM maturity.

4.1 Common anti-patterns

Anti-pattern #1 — Copy Production and Call It a Day

It sounds efficient, but it creates legal and compliance risks. Realism should not be confused with reusing real data. In many cases the necessary infrastructure or environment security for such a copy simply does not exist.

Alternatives:

  • Use representative subsets rather than a full copy
  • Apply masking before moving data
  • Combine de-identified real data with synthetic data for new scenarios
Anti-pattern #2 — The Local Heroes

This leaves you dependent on one or two people who know the data model — and you are stuck when they are unavailable.

  • Knowledge accumulates but is never documented
  • There is no versioning or catalogue
  • Scripts are modified ad-hoc, with new criteria added on the fly
  • Nothing guarantees that two requests for the same dataset return the same results
  • Every iteration requires manual coordination
  • When data spans multiple technologies, it becomes unmanageable
Anti-pattern #3 — The Refresh That Breaks Everything

This costs you days every time the environment is reset. If every data refresh wipes out your tests, the data is not really yours. With TDM, a data refresh stops being a disaster and becomes a push-button operation.

Anti-pattern #4 — The Ownership Grey Area

Nobody knows who owns the test data. Is it QA? The DBA? Security? Infrastructure? The result is that everyone touches it but nobody governs it. Priority conflicts arise because QA demands more data, Security blocks it, and Infrastructure complains about the volume.


4.2 Warning signs in your organisation

Warning #1 — Data-Driven Instability

Tests that sometimes pass and sometimes fail with the same code because the data has changed, does not meet preconditions, or is inconsistent. If a test passes or fails depending on how the environment was set up, it is not a reliable test.

Solution: Deliver consistent data with referential integrity and traceability.

Warning #2 — Single-Person Dependency

The team relies on one or two people who know how to query the database and prepare data with their own scripts. That is tacit knowledge rather than documented process. You cannot afford to wait days to get the data you need for testing.

Solution: An archetype-based search system with self-service capabilities that empowers QA teams to request data with full guarantees.

Warning #3 — Environments Take Days to Set Up

When building an environment or preparing data takes days or weeks, you are effectively losing the sprint waiting for everything to be ready. If it takes days to get data, your bottleneck is not testing — it is the data.

Warning #4 — Real Data in QA

Production data is used in test environments with no controls. Common reasons:

  • Realism is confused with reusing real data
  • No masking rules exist
  • Pressure to move fast leads to direct copies
05

Minimum viable TDM architecture

The smallest set of components you need to demonstrate real value in test data management. Just as in development we build an MVP to validate an idea, in TDM we can build a minimum viable architecture to validate the model.

5.1 The four pillars

1. The data domain
Choose a key dataset that represents a real business case. Examples:
  • Customers and orders
  • Customers, policies, and receipts
  • Patients and treatments
The idea is to avoid trying to cover every system from day one.
2. The orchestration
Build a simple pipeline that manages the data: extract, mask where needed, and deliver. The orchestration automates the entire flow.
3. The masking policies
Define the policies that govern which data gets masked and how. Policies drive security and compliance.
4. The provisioning
The mechanism that delivers data to the target environment. Provisioning puts data where and when it is needed.

5.2 Benefits of the minimum viable architecture

Important

This architecture is not meant to be a throwaway pilot. It is the first step of a system that grows based on evidence. It is a reproducible and measurable system in miniature — one that already demonstrates the benefits of TDM.

06

TDM mechanisms

Test data is not just managed — it is designed. And designing it well turns testing into reliable knowledge. The TDM mechanisms work together to make this possible.

6.1 Data masking

6.1.1 Definition and purpose

Data masking is the transformation of sensitive data through controlled rules that preserve the technical value of the data so it remains usable. It is the cornerstone of compliance by design.

It involves transforming or replacing sensitive values (names, national IDs, phone numbers, health-related information) with alternatives that retain the original format and logic without allowing the identification of real individuals.

Regulatory compliance
GDPR, DORA, and other regulations
Stronger security
No real sensitive data in non-production environments
Risk reduction
Eliminates the risk of re-identifying individuals

6.1.2 Types of masking

A) Full anonymisation (100%)

Performs a one-time replacement of data values and, once the process is complete, destroys any evidence of the anonymisation. It always uses non-deterministic algorithms so that values cannot be reversed to identify individuals.

This approach is more common for ephemeral environments or when data needs to be delivered to third parties — the data is anonymised today and any link to the original information is permanently severed.

B) Pseudonymisation

Retains a protected mapping. The masking process runs and stores — in a fully secured manner — the fact that ID number 1 was changed to 5, that bank account X was changed to Y, and so on (individual data points that cannot be linked back to a specific person).

Advantages of pseudonymisation
  • After an initial anonymisation cycle, if teams identify the data they need, the next run will produce the same masked values — so the same data keeps working
  • Simplifies the process of finding and selecting data for tests
  • When multiple applications share related data but cannot be processed at the same time, you can mask one today and the other in a week or a month — and still get consistent results

6.1.3 The data masking process

1

Sensitive data discovery

Scan the database — every table, every column, the actual content — to identify where sensitive data resides. This is done through inspectors that apply:

  • Static analysis: Column name, type, and length. If we are looking for a name and the column is numeric, it is ruled out
  • Dynamic analysis: Reads the content of each column and applies various methods to determine the data type (validation algorithms, format checks, length analysis)

The output is a heat map (very high, high, medium) indicating the likelihood that each column contains a specific type of sensitive information.

2

Masking policy configuration

Synchronise the table metamodel in the application and assign specific masking algorithms to each column:

  • This column needs to be anonymised with the national ID algorithm
  • This one with the personal name algorithm
  • This one with the bank account algorithm
3

Masking execution

Select the target environment (test, development, integration, pre-production) and run the process. Steps execute in a specific order, with parallelisation where possible:

  • Encrypted, serialised data download
  • Privilege validation
  • Temporary disabling of indexes and foreign keys for performance
  • Application of masking algorithms
  • Database update with masked values
  • Re-enabling of indexes and constraints
4

Automation

The process can be automated via an orchestrator (e.g. over the weekend the production database is copied and the next step is a bulk masking run). icaria TDM provides a full REST API layer for these executions.

6.1.4 Multi-system consistency

A masking process can target databases of different technologies simultaneously. If the environment comprises a Postgres database, an Oracle instance, and a DB2 system, the same process can address all of them at the same time — seamlessly.

Important

All icaria TDM processes can run simultaneously across multiple applications and environments, completely agnostic to the underlying database technology.

6.1.5 Data quality assurance

How do we ensure that data masking does not degrade the quality or usefulness of the test dataset?

6.1.6 Processing capacity

+1 Billion

records anonymised across our customer base. icaria TDM can parallelise as long as there are available CPU and memory resources.


6.2 Subsetting

6.2.1 Definition and purpose

Subsetting extracts a coherent subset of data from production or a test environment and delivers it to another environment while preserving full referential integrity and all table relationships.

Rather than copying entire databases, subsetting selects a representative set by extracting only the entities you need.

99%

reduction in database volume through intelligent subsetting.

6.2.2 Problems it solves

Cloning costs
A 70 TB environment means significant storage costs, cloning times of 2+ weeks, and hidden expenses in licences and staff time.
Over-provisioning
The environment ends up with excess data that gets in the way. Sometimes less is more: more infrastructure, heavier systems, and harder to find what you need.
Unnecessary risk
Copying the entire database creates re-identification risk. Subsetting complies with the GDPR data minimisation principle and other privacy regulations.
Lack of agility
When it takes two weeks to set up an environment, agility is practically non-existent.

6.2.3 Use case: production incident

Scenario

A production incident occurs involving a specific customer, and you need to reproduce it.

With traditional bulk masking: Perhaps the data exists because a copy was made last week, but it is already stale. The developer confirms the bug is reproducible... and now the data is inconsistent and retesting becomes impossible.

With subsetting: The moment the issue is detected, you take the data from the source environment (production) and deliver it to another environment with masking applied. You get the data you need, when you need it. It is repeatable and takes seconds.

6.2.4 Key characteristics

6.2.5 Internal repositories

Repositories allow you to save a snapshot of a customer (or as many as needed) in multiple versions, independent of what happens in the source environment.


6.3 Data finder and archetypes

6.3.1 The problem of finding the right data

Even with "unlimited" data (a full copy of production), the question remains: how do you find exactly the data you need in this ocean of records? Out of millions of customers, how do you locate the ones that actually satisfy your test requirements?

Example

"My test needs a newly created customer." But what conditions define "newly created"?

  • When all the forms have been filled in?
  • When a waiting period has passed for the customer to become active?
  • When the customer has uploaded documentation to online banking?
  • For development it might mean one thing; for testing it could mean something entirely different

6.3.2 The problem with queries

Managing data through queries comes with multiple challenges:

6.3.3 The solution: data archetypes

Archetypes (or data profiles) are standardised definitions that everyone agrees on. They are categorisations of data profiles built from a set of conditions that define the data — where everyone shares the same, standardised understanding.

6.3.4 Self-provisioning data pool

Using archetypes, a self-provisioning data pool is created: data matching the archetype definition is automatically located across the available test environments.

6.3.5 Custom search requests

If the existing archetypes do not cover what you need, you can create a search request from scratch:

  1. Start from an archetype, a template, or a blank slate
  2. Select the data structure (e.g. customers)
  3. Choose the environment to search (production, test, etc.)
  4. Specify how many items to locate (1, 5, 1,000...)
  5. Select the conditions that match your requirement

6.3.6 Batch request execution

Requests are processed in batches to properly manage database access. Organisations automate execution in scheduled time windows that do not impact their environments.


6.4 Self-service

6.4.1 Concept

Self-service completes the process and allows you to request dataset provisioning directly into the environment where it is needed — without raising tickets or waiting on other teams.

QA and DevOps gain real autonomy, and the testing workflow accelerates exponentially. It breaks the dependency on specific individuals or teams and turns test data into an on-demand service.

6.4.2 Level of autonomy

100%

autonomy. With the appropriate training, teams can manage the entire cycle themselves.

6.4.3 Self-service process for subsetting

From the self-service user's perspective:

1

Request delivery

Log in to icaria and request the data delivery

2

Select the structure

Choose the data structure (e.g. customers)

3

Choose the destination

Select the target environment (test, development, repository)

4

Provide a reason

Log the reason for audit purposes (can include a ticket number or test case ID)

5

Set the search criteria

Select the seed or search criteria to use

6

Submit and wait

The background queue processes pending requests


6.5 Data design in the test case

6.5.1 The core concept

This is the fifth mechanism and it closes the loop. The concept shifts: it is no longer the tester who searches for or decides which data to use — instead, the data is already designed as part of the test case itself. It is the starting point for building and executing automated tests.

Each test carries an associated dataset that meets the exact conditions it requires (users, customers, orders, statuses, dates), and the system provisions it automatically before execution.

6.5.2 The 3 ingredients of a test

1. The source code
The code under test
2. The actions
The script with the steps to execute
3. The data
The data required for the test to run reliably
A Common Pitfall

Typically, managing this data is not included in the plan — it is neither part of the test definition nor the automation setup. That is why we start well with automation (Selenium scripts are straightforward to write) but things fall apart as tests grow in complexity.

6.5.3 Complete test case design

From a data perspective, the design includes:

Store
Keep the data in a secure repository so it retains its quality, stays consistent, and every delivery is identical.
Deliver
Every time the test runs, in whichever environment it runs in.
Validate
Verify the result from the data perspective.

6.5.4 The execution flow

1

Find

Locate the data in production that meets the required conditions

2

Protect

Move the data to an internal, protected database so no one can modify it

3

Deliver

On the test's request, deliver the data to the appropriate environment

4

Restore

Once the data has been consumed, restore it to its original state for reuse

5

Validate

Run the test and verify the outcome from the data perspective

6.5.5 Validation rules

Validation rules define the expected outcomes from a data perspective:

6.5.6 Pre-delivery rules

Pre-delivery rules ensure data quality before it reaches the test environment.

Example: Date of Birth

If the test requires the user to be under 18: icaria TDM reads the date, calculates how old the person would be at the time of delivery, and adjusts it to meet the condition. The delivered date of birth becomes today minus 17 years, ensuring the person is always 17 at the time of delivery.

6.5.7 Subsetting structure (data domain)

To deliver consistent data, a data domain is defined: what needs to be delivered (e.g. a complete partner record with at least one quotation in draft status across both applications).

The subsetting structure has a header entity (e.g. partner) with all its required related entities to complete the domain across every application involved.

07

TDM maturity model

TDM maturity is a journey, not a leap. icaria TDM has identified five levels, from reactive to continuous improvement.

Important

This model does not expect every organisation to reach the highest level. Its purpose is to help you identify where you stand, what maturity level each process has reached, and what the next logical step looks like.

Level Characteristics Indicators
1 — Initial Test data is managed reactively, manually, ad-hoc. Direct copies of production with no controls. No catalogue, no masking, reliance on "local heroes"
2 — Repeatable Some basic masking processes exist. Sensitive data has been identified in the main systems. First sensitive data map, basic masking rules
3 — Defined A Minimum Viable Architecture has been implemented. Automated provisioning is in place for at least one data domain. Automated pipeline, functional subsetting, defined archetypes
4 — Managed TDM integrated into CI/CD. Self-service operational. Usage metrics and lead time visible. Data as a service, full traceability, multiple domains
5 — Optimised Continuous improvement driven by metrics. Test data designed as part of the test case. Compliance by design across all environments. 100% coverage, measured ROI, zero real data in non-production

7.1 How to get started

The goal is not to do everything at once, but to demonstrate value quickly, learn, and then scale. The two levers that drive this are:

Regulation
If you face issues from full production environment copies, start by applying bulk masking mechanisms.
Efficiency
Pick a well-defined domain, build a simple pipeline, set up masking rules, and establish controlled provisioning.
08

CI/CD integration

By incorporating icaria TDM into the testing ecosystem, you complete the continuous testing pipeline.

8.1 The testing ecosystem

Testing teams commonly use tools such as:

Yet data management tooling is rarely part of the conversation. Adding icaria TDM completes the ecosystem.

8.2 The integrated execution cycle

Jenkins
(Orchestrator)
→
icaria TDM
(Provisioning)
→
Execution
Tool
→
icaria TDM
(Validation)
→
Results
Report
1

Query

The orchestrator (e.g. Jenkins) asks the test case manager which test to run and in which environment

2

Data provisioning

The orchestrator calls icaria TDM to provision the data into the target environment

3

Test execution

Once icaria TDM finishes, the orchestrator triggers the execution tool to run the test

4

Data validation

After the test, the orchestrator asks icaria TDM to verify results from the data perspective

5

Report

All information is consolidated into a report and the results are stored with data from every tool in the chain

8.3 End-to-end automation

When a test case is triggered, the system automatically:

The Result

This eliminates the typical bottlenecks in data preparation and synchronisation phases, enabling parallel test execution without conflicts or dependencies. It is the final step toward continuous testing, where data and tests advance at the same pace as development.

09

Benefits and ROI

The measurable impact of TDM on speed, coverage, costs, risk, and team satisfaction.

9.1 Key benefits

Greater Speed
  • Data is ready before the test begins
  • Fast recovery: restore and re-deliver
  • Environments that used to take weeks are ready in minutes
Greater Coverage
  • Legacy, new, and complex data always available
  • 100% realistic coverage
  • Happy-path and unhappy-path data alike
Lower Risk
  • Data looks real but is not
  • Regulatory compliance (GDPR, DORA)
  • No real data in non-production environments
Greater Efficiency
  • ROI exceeding 300%
  • Infrastructure cost reduction (up to 90%)
  • More autonomous and effective teams
>300% Return on investment
90% Infrastructure cost reduction

9.2 Additional operational benefits

9.3 Benefits for the team

Improved communication within and across teams:

We couldn't live without icaria (TDM)

— icaria Technology customer
10

icaria TDM: the test data management platform

Data management specialists. We help ensure data is governed, secure, and usable.

10.1 About icaria Technology

icaria Technology is a company specialising in data management. It helps organisations design and operationalise their data — from data governance and the protection of production environments through the automated enforcement of privacy rights (such as the right to erasure), to the reproducible provisioning of test data.

10.2 Gartner recognition

Gartner 2025 — Sample Vendor

icaria TDM is recognised by Gartner as a Sample Vendor in their 2025 report on three steps to optimise test data management.

TDM adoption, according to Gartner, is still at a very early stage in many organisations. Most companies — even large ones — still rely on manual processes. But TDM technology has matured: what the tools deliver matches what users expect.

10.3 About icaria TDM

icaria TDM is the Test Data Management platform built for mission-critical applications and OSS/BSS environments. It delivers realistic, secure, accurate, and consistent data exactly when and as many times as it is needed — ensuring that every test result is effective and reliable.

With icaria TDM, testers and developers dramatically reduce the hours spent producing and managing data, allowing them to focus on higher-value tasks. Leading companies in banking, insurance, and telecommunications already trust icaria TDM to transform their testing processes.

Three solutions, one complete suite

Test Data Masking

Security and compliance without compromise. Applies advanced anonymisation and pseudonymisation techniques to ensure data is representative without exposing sensitive information. GDPR-compliant and consistent across multiple environments.

On-Demand Test Data

Data ready for manual testing. Testers gain autonomy by obtaining optimised subsets on demand, finding the exact data for each test case through the data finder, and extracting datasets while preserving integrity and relationships.

Test Data for CI/CD

Automated data for continuous testing. Automates data provisioning in CI/CD pipelines, eliminating bottlenecks. Every test run gets the right, up-to-date data — no manual intervention required.

Compatible with leading platforms — Oracle, SAP, Salesforce, IBM, Hadoop, and many other data sources — icaria TDM integrates into any technology stack, delivering a robust and scalable solution. QA teams test more, test better, and test faster.

10.4 icaria TDM capabilities

Multi-technology
  • Oracle, Postgres, MySQL, SQL Server, DB2
  • Fixed-length files, CSV files
  • Other non-relational sources
Discovery
  • Static and dynamic profiling
  • Cross-application relationship identification
Masking algorithms
  • Standard algorithms for each data type
  • Extensible and customisable with Java
  • Bulk configuration
APIs and automation
  • Full REST API layer
  • Integration with orchestrators (Jenkins, etc.)

Deployment

Installed on the customer's own infrastructure (cloud or on-premise). icaria TDM is privacy-aware by design: the data belongs to the customer and stays within their infrastructure.

10.5 Implementation project

Getting a TDM project up and running requires:

  1. Obtaining database connections
  2. Installing the software on the customer's infrastructure
  3. Identifying requirements and key pain points
  4. Creating data structures
  5. Building the sensitive data map with masking rules
Fast Payback

The project takes a few months, but the return on investment is visible very quickly — even during the implementation itself, data quality issues are uncovered and smaller, leaner data deliveries already provide tangible benefits.

10.6 Customers and references

Large organisations with complex structures, demanding regulatory compliance requirements, and very large data volumes trust icaria TDM: banks, telecommunications companies, insurers. The platform scales from smaller teams working with hundreds of tables to organisations with thousands of tables, multiple technologies, and diverse teams.

icaria TDM

Ready to transform your test data management?

Discover how icaria TDM can help your organisation cut costs, increase test coverage, and achieve regulatory compliance — from day one.

Request a Demo

www.icariatechnology.com

icaria Technology — Data Management Specialists

© 2026 icaria Technology. All rights reserved.