AI Training Data Transparency via Blockchain Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial Intelligence (AI) decision-making systems are influenced by non-objective influences in their training data, leading to biased outcomes due to geographic or diversity limitations, which can result in inappropriate decision-making.
Innovation Solution
A method and computer program product that provide transparency in training data by analyzing and rating the quality and non-objective influences of data sources, using metadata and publicly available information to curate a knowledge base for AI systems, excluding non-objective influences and weighting data sources based on predefined rules stored in a blockchain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If AI systems use diverse training data from multiple sources, then the coverage and representativeness of training data is improved, but non-objective influences and biases from certain sources can negatively impact decision-making reliability
Solution Approach 1:
The patent applies local quality by assigning different weights to different data sources based on their objectivity and quality characteristics. Each data source is evaluated individually for non-objective influences, and weighting factors are applied locally to each source rather than treating all data uniformly. This allows the system to maintain diversity while mitigating bias from specific sources.
Solution Approach 2:
The patent changes the parameter of data source weighting dynamically based on quality assessments. By introducing quality metrics and non-objective influence scores as adjustable parameters, the system can modify the effective contribution of each data source to the AI training process, thereby improving reliability while maintaining diversity.
2Quantity of substance
If AI systems incorporate data from all available sources, then the quantity and coverage of training data is improved, but the complexity of managing and evaluating data quality increases
Solution Approach 1:
The patent applies preliminary action by performing quality assessment and bias detection on data sources before they are fully integrated into the AI training process. The system pre-evaluates each data source for non-objective influences and assigns quality scores in advance, which simplifies subsequent data management and integration processes.
Solution Approach 2:
The patent introduces an intermediary layer between raw data sources and the AI training process. This intermediary system evaluates data quality, detects biases, and applies weighting factors, thereby managing the complexity of handling diverse data sources without requiring direct complex management of each individual source.
3Ease of manufacture
If AI systems use data without quality assessment, then the ease of implementation is improved, but the accuracy and objectivity of decision-making deteriorates
Solution Approach 1:
The patent applies self-service by enabling the AI system to automatically assess the quality and objectivity of its own training data sources. The system performs self-evaluation of data sources for non-objective influences and adjusts weighting accordingly, reducing the need for external manual quality assessment while maintaining decision accuracy.
Data Source
AI summary
A method, computer program product, and a system where a processor(s) obtains data from one or more sources, wherein a portion of the one or more sources comprise training data for a first artificial intelligence decision-making system. The processor(s) ingest data from each source into a corpus and ingest metadata corresponding to the data, into a volume accessible to a second artificial intelligence decision-making system. The processor(s) search public sources and obtain information describing quality and non-objective influence of the data from each source. The processor(s) provide as inputs to the second artificial intelligence decision-making system, the information and the metadata, to classify the quality and the non-objective influence of each source. The processor(s) obtains outputs comprising, for each source, a first rating classifying the quality of the source of the one or more sources. The processor(s) load the outputs into a blockchain, obtain rules, and designate, via the blockchain, based on the rules, the portion.


