To achieve this, teams traditionally copied large volumes of production data and then applied rudimentary masking techniques to sensitive information. This required significant computing resources and time. Today, this model is becoming obsolete thanks to artificial intelligence.
However, the high delivery speed enabled by artificial intelligence also creates a need for larger volumes of data in less time to prevent bottlenecks in pipelines. At the same time, AI itself offers a solution: today’s algorithms are capable of understanding business dependencies and logical correlations within real data to automatically generate large-scale synthetic environments.
The 3 architectures dominating the ecosystem
When evaluating how to integrate artificial intelligence into TDM, it is important to keep in mind that there is no single answer. However, there are three main techniques currently available in the market:
- Intelligent synthetic data generation: it does not interact with production data. The software uses Deep Learning to understand data behavior and replicate its logical structure in a 100% artificial way. Tonic.AI is a tool of this type, using neural networks to autonomously discover business dependencies. It is very easy to integrate into CI/CD pipelines thanks to its cloud-native architecture and is also low-maintenance, as it automatically adapts if the underlying data schema changes. It is ideal for modern cloud-based architectures and projects that require strict privacy compliance.
- Dynamic component-based modeling: it generates data flows using structured logical rules. It is ideal for simulating edge cases and large-scale stress testing. GenRocket is useful in this scenario for creating high volumes of data through logical rules. It uses AI to automate the initial import and modeling of the data schema. It requires a moderate implementation and maintenance effort, as teams need to be trained to properly model data scenarios. It also requires periodic schema updates if the business logic of the data changes. It is ideal for ecosystems with massive databases that require replicas with a high level of accuracy.
- Legacy data virtualization: it clones the physical storage structure, making it possible to create copies of production data and then apply superficial or deep masking techniques. One tool for this purpose is Delphix, which is useful for cloning physical storage blocks and masking data. It does not generate synthetic data, but instead relies on traditional data masking algorithms. It requires direct and ongoing coordination with system and infrastructure administrators, although it is operationally stable. At the same time, physical storage consumption must be continuously monitored. Its ideal use case is large monolithic ecosystems where identical cloning of production databases is a mandatory requirement.
Within this legacy data virtualization architecture, Informatica IDMC is a platform used in centralized governance and traditional enterprise data masking scenarios. Its value lies in providing centralized control supported by its enterprise AI engine (Claire). It requires a very high implementation effort, resulting in slow enterprise deployments that often involve several months of specialized consulting. It is also high-maintenance, requiring engineers and administrators dedicated exclusively to managing the platform. However, it is ideal for large corporations with highly complex hybrid architectures and critical dependencies on legacy systems or mainframes.
The strengths and limitations of each tool demonstrate that choosing a TDM strategy should not be based on market trends, but rather on carefully weighing the available resources and infrastructure. If the operational core is based on cloud microservices and the primary goal is to comply with privacy regulations, intelligent synthetic data generation (Tonic.AI) is the best option. On the other hand, if the main challenge is accelerating data provisioning for functional testing, dynamic modeling (GenRocket) emerges as the ideal alternative. Finally, legacy data virtualization tools (Delphix, Informatica IDMC) should be reserved for monolithic architectures or complex legacy systems, such as banking cores or mainframes.
By Ricardo García, Tester & QA of Baufest.


