Data Validation System for Algorithm Bias Certification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data systems lack access and agency for individual entities and people over their data, leading to systemic inaccuracies and biased predictions in various fields such as justice, economics, education, and health.

Innovation Solution

A system and method for data set validation, bias characterization, and valuation, which involves filtering data sets to extract core information, creating certified models to evaluate bias, and assigning value scores based on data quality, provenance, and bias characterization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data sets are used from data brokers without validation, then data availability and processing speed are improved, but data accuracy and reliability deteriorate due to systemic inaccuracies and lack of provenance understanding

Engineering Contradiction:
Improvedata processing speedVSAvoiddata accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary validation, certification, and bias characterization on data sets before they are used in machine learning workflows. A certification manager validates data sets against predefined criteria, and a bias characterization auditor assesses potential biases, ensuring data quality is established in advance rather than discovered later during model deployment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary components between data sources and machine learning models, including a certification manager that validates data provenance and quality, and a bias characterization auditor that assesses data biases. These intermediaries act as mediators that filter and certify data before it reaches the modeling stage, resolving the contradiction between speed and reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If comprehensive data validation and bias characterization processes are implemented, then data accuracy and fairness are improved, but system complexity and computational resources increase

Engineering Contradiction:
Improvedata qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The validation system is segmented into distinct modular components: a certification manager for validating data provenance and quality, and a separate bias characterization auditor for assessing biases. This segmentation allows each component to perform its specific function independently, making the overall complex system more manageable and maintainable while ensuring comprehensive data quality checks.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If data sets lack bias characterization and provenance information, then data processing simplicity is maintained, but algorithmic bias and unfair outcomes increase in critical fields such as justice, economics, and health

Engineering Contradiction:
Improvedata processing simplicityVSAvoidalgorithmic bias
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system implements feedback mechanisms where the bias characterization auditor continuously assesses data sets for biases and provides feedback to the certification process. This feedback loop ensures that data sets with problematic biases are identified and can be corrected or rejected, preventing algorithmic bias from propagating into machine learning models while maintaining a relatively simple interface for users.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250088542A1Data set and algorithm validation, bias characterization, and valuation
Publication Date: 2025.03.13 QOMPLX INC
  • US20250088542A1 patent drawing
  • US20250088542A1 patent drawing
  • US20250088542A1 patent drawing

AI summary

A system and method for providing access and agency to individual entities and people over their data for the purpose of data set validation to facilitate data set and algorithm bias certification and scoring. A first data set is filtered to extract its core information content and to create a certified data set. A certified model is created by training a machine learning algorithm on the certified data set, which certified model is then used to evaluate the bias of subsequent data sets. The data set may be given a value score which represents the overall validity of the data set and its bias characterization. A bias characterization audit can help identify the root causes of bias outcomes from predictive software and algorithms that perform third party tasks and services. The score can be used as a metric to further facilitate market transactions.