ETSI announces the launch of its new Technical Report (TR) ETSI TR 104 180: Data Solutions; Development and identification of Data Quality Metrics, a standardised framework for quantitatively assessing the quality of data used across digital and AI ecosystems. The framework will help organisations determine whether data is fit for their intended purpose, while laying the groundwork for consistent, reproducible and standardised approaches to data quality testing.
The TR defines 18 ‘Data Quality Metrics’, with formal definitions and mathematical measurement methods, designed to help organisations assess whether datasets are sufficiently reliable, complete, accurate, representative and trustworthy for their intended purpose.
As AI adoption increases, organisations are becoming increasingly focused on the quality of the data underpinning AI ecosystems. ETSI’s 18 Data Quality Metrics enable organisations to determine potential quality, reliability and bias issues.
The eighteen identified metrics described in the document focus on:
- Fundamental data quality: including completeness, accuracy, reliability, consistency, precision, integrity, redundancy and uniqueness.
- Usability: including availability, coverage, lineage, traceability and timeliness.
- Fairness: including label quality, measurement bias and representation bias.
- Privacy and responsible data use: including anonymity and confidentiality.
To validate the framework and showcase its applicability, ETSI has applied the metrics to two contrasting datasets – industrial IoT sensor data and demographic data. This demonstrated how the metrics can be used across data contexts, and highlighted how different use cases require different combinations and weights of the Data Quality Metrics:
- Industrial IoT sensor data: where characteristics such as reliability, timeliness, accuracy, completeness, traceability and integrity are particularly important.
- Demographic data: representing human and social characteristics, considering completeness, anonymity, representation bias, label quality, confidentiality and coverage.
The metrics can enable dataset owners to self-verify and self-report data quality, helping them to transparently communicate reliability and build trust within digital and AI ecosystems.
“It is essential that data quality is measurable, especially for organisations who need to establish whether its data is fit to essential intents, like it would be the case of trustworthy AI,” said Diego Lopez, Chair of the ETSI Technical Committee DATA. “ETSI’s standardised metrics provide a common language for assessing data quality, giving quantitative evidence as to whether a dataset is fit for its intended purpose. This lays important groundwork for more consistent and repeatable approaches to data quality assessment, as AI and data-driven technologies continue to evolve.”
The proof of concept was performed by a team involving Sejong University, EGM, TTA, Daejeon University and CNIT, and it developed a functional, open-source Data Quality Validation System. The tool can provide a score to any dataset based on the eighteen Data Quality Metrics defined in the TR.
There’s plenty of other editorial on our sister site, Electronic Specifier! Or you can always join in the conversation by commenting below or visiting our LinkedIn page.
