ML-Based Data Efficacy Scorer for Automated Quality Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for measuring and improving data efficacy are time-consuming and require significant manual effort, as they rely on users defining and configuring metrics and thresholds, which are not adaptable to changing environments and user roles, leading to costly delays in data-driven decision-making.

Innovation Solution

An intelligent system using machine learning (ML) for automatic data efficacy scoring, anomaly monitoring, and personalized recommendations, which trains on meta-features and data quality metrics to provide adaptive and generalizable data quality scores across domains, reducing the need for manual configuration and static rules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual definition and configuration of data quality metrics and thresholds is used, then measurement precision can be improved, but device complexity and time consumption increase significantly

Engineering Contradiction:
Improvedata quality measurement precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system automatically selects and configures data quality metrics and thresholds without requiring manual user definition. The machine learning model autonomously analyzes data characteristics and determines appropriate metrics, enabling the system to serve itself rather than requiring extensive manual configuration by users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts data quality metrics and thresholds based on changing data characteristics and domain requirements. The machine learning model adapts parameters automatically, allowing the system to maintain high measurement precision across different domains without manual reconfiguration.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual curation of data efficacy measures is performed, then measurement precision for specific user roles can be improved, but loss of time increases

Engineering Contradiction:
Improveuser-specific data efficacy measurement precisionVSAvoidtime for manual configuration
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically determines which data quality metrics are most relevant for each user role by analyzing domain characteristics and data patterns. This eliminates the need for manual curation of metrics while maintaining user-specific precision through automated role-based adaptation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-configures data quality measures based on domain knowledge and historical data before actual measurement occurs. This preliminary setup allows the system to immediately provide user-specific metrics without requiring manual curation at the time of use.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If static data quality rules are used, then ease of operation can be improved, but adaptability to changing environments deteriorates

Engineering Contradiction:
Improveoperational simplicityVSAvoidadaptability to changing data environments
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system transitions from static rules to dynamic, adaptive data quality monitoring. The machine learning model continuously learns from changing data patterns and automatically adjusts metrics and thresholds, enabling the system to adapt to evolving environments while maintaining operational simplicity through automation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements continuous feedback loops where performance data is used to refine and improve data quality metrics over time. This feedback mechanism allows the system to automatically adapt to changing environments while maintaining ease of operation, as the system self-adjusts rather than requiring manual rule updates.

Inventive Principle:
Principle #23Feedback

4Reliability

If comprehensive data quality monitoring is implemented, then reliability of data-driven decisions can be improved, but loss of time for data cleaning increases

Engineering Contradiction:
Improvedata-driven decision reliabilityVSAvoidtime for data cleaning and efficacy issues
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system automatically detects, measures, and reports data quality issues without requiring manual intervention for data cleaning. The machine learning model autonomously identifies problems and provides recommendations, enabling comprehensive monitoring while minimizing time investment from users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces manual data cleaning processes with automated machine learning-based detection and analysis. This substitution eliminates the need for manual inspection and correction of data issues, maintaining comprehensive monitoring reliability while significantly reducing the time required for data cleaning activities.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20230136094A1Automatic, personalized, and explainable approach for measuring, monitoring, and improving data efficacy
Publication Date: 2023.05.04 ADOBE INC
  • US20230136094A1 patent drawing
  • US20230136094A1 patent drawing
  • US20230136094A1 patent drawing

AI summary

A method of determining efficacy of a dataset includes receiving data from a data source, wherein the data comprises a plurality of fields of unknown efficacy; mapping the data based on a plurality of data quality metrics and based on attributes of the plurality of fields wherein meta-features for the data are obtained; predicting a value for each of the plurality of data quality metrics using a ML model that takes the meta-features as input, wherein the value indicates whether a corresponding data quality metric is suitable for measuring efficacy of the fields; selecting a data quality metric based on the value, wherein the data quality metric measures an efficacy of the fields; and monitoring the efficacy of the fields in the data received from the data source based on the data quality metric.