Training Data
Datasets used for training, fine-tuning, and continued learning fail differently than the data traditional DLP was built to protect — poisoning and backdoor attacks target what a model learns, not just what it exposes.
This category covers dataset provenance, lineage tracking, and poisoning detection purpose-built for the data pipeline feeding your models.
Looking for help in this category?
Get matched with a vetted partner.
Tell us what you're trying to solve and we'll connect you with a vetted partner.
What Training Data covers
The subcategories, in practitioner terms.
Dataset provenance & lineage
Tracking where training data came from and how it has been transformed before it reaches a model.
Poisoning & backdoor detection
Screening datasets for manipulated samples designed to corrupt model behavior.
Access control for training pipelines
Governance over who can add to, modify, or trigger retraining against a dataset.
Compliance mappingMaps to NIST IR 8596's training-data concept, the CSA AI Controls Matrix's Data Security & Privacy Lifecycle Management and Model Security domains, ISO 42001 Annex A.7 (data for AI systems), and MITRE ATLAS's data-poisoning and dataset-membership-inference techniques.
Category structure adapted from the AI Defense Matrix by Lenny Zeltser and Sounil Yu, licensed CC BY-SA 4.0. For a broader view of the vendor landscape, see the AI Defense Matrix Catalog.
Not sure Training Data is your biggest gap? The assessment will tell you.
Take the free assessment →