Automated Data Enrichment Using Relevance and Confidence Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual data enrichment processes are subjective, error-prone, and inefficient due to the large number of data objects and variability in sources, leading to inconsistent results, especially when dealing with multiple sources.

Innovation Solution

An automated data enrichment system that uses an attribute relevance module to measure relevance, a source selection module to choose the best sources, an output value confidence module to calculate confidence, a source utility and adaptation module to assess source utility, and an ambiguity resolution module to handle multiple outputs, enabling objective and accurate enrichment of data objects across heterogeneous sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data enrichment is performed by searching sources and subjectively determining information pertinence, then data objects can be enriched with additional information, but the process generates erroneous results and is inefficient due to subjectivity and the large number of data objects

Engineering Contradiction:
Improveaccuracy of enrichment resultsVSAvoidenrichment processing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables automated self-enrichment of data objects by having the enrichment system automatically search sources, evaluate pertinence, and integrate information without human intervention. The system uses automated relevance evaluation and confidence scoring to make enrichment decisions independently, eliminating manual subjectivity while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of human enrichment with an automated computational system. The system uses algorithms to search sources, evaluate relevance, calculate confidence scores, and determine enrichment decisions, substituting human cognitive processes with automated mechanical-computational processes that are both faster and more consistent.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If the enrichment process is automated by obtaining information from a source, then processing efficiency improves, but changes in the content or format of the source require modification of the enrichment process

Engineering Contradiction:
Improveenrichment processing speedVSAvoidadaptability to source changes
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts its enrichment parameters and confidence thresholds based on source characteristics and performance. The automated relevance evaluation and confidence scoring mechanisms adapt to different source formats and content types, allowing the system to maintain high productivity across varying source conditions without requiring manual process modification.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The enrichment system is designed to handle multiple source types and formats through a universal automated process. The system's ability to evaluate relevance and calculate confidence scores applies across different data sources, making the automated enrichment process versatile and adaptable to various source changes without requiring source-specific customization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If multiple sources are available for enrichment, then data completeness can be improved, but selection of a particular source adds subjectivity to the enrichment process

Engineering Contradiction:
Improvecompleteness of enriched dataVSAvoidobjectivity of source selection
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system uses feedback from confidence scoring and relevance evaluation to objectively select among multiple sources. The automated process calculates confidence scores for information from different sources and uses this feedback to determine which sources provide the most reliable enrichment data, eliminating subjective source selection while maintaining data completeness.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The source selection process is dynamic and adaptive, automatically adjusting which sources are used based on their performance and the specific data object being enriched. The system dynamically evaluates source relevance and confidence for each enrichment task, providing an objective, data-driven approach to selecting among multiple sources that improves completeness without introducing subjectivity.

Inventive Principle:
Principle #15Dynamics

4Reliability

If manual enrichment is performed to ensure accuracy, then data quality can be maintained, but the large number of data objects requires excessive effort and cost

Engineering Contradiction:
Improvedata qualityVSAvoideffort and cost of enrichment
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system applies partial automation by using automated source selection and confidence scoring for all data objects, while reserving manual review only for cases where the automated system's confidence score falls below a threshold. This approach maintains data quality for the majority of objects through automated processes, reducing the effort and cost associated with manual enrichment while preserving accuracy where needed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8924407B2Data enrichment using heterogeneous sources
Publication Date: 2014.12.30 ACCENTURE GLOBAL SERVICES LTD
  • US8924407B2 patent drawing
  • US8924407B2 patent drawing
  • US8924407B2 patent drawing

AI summary

A data enrichment system may include an attribute relevance module to measure relevance of an attribute to a data object to be enriched. The data object may include the attribute including a known or an unknown value. An output value confidence module may calculate a confidence of an output value of a source used for enrichment of the data object. The output value may represent the known and/or unknown values of the attribute. The system may use the measured relevance of the attribute and the calculated confidence of the output value to determine assignment of the known or unknown values to the attribute.