Data Quality Improvement via Underrepresented Segment Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data sets are often unrepresentative or incomplete, leading to biased models that fail to accurately reflect population behavior, as they lack data from specific segments of the population, resulting in inaccurate and unreliable results.

Innovation Solution

A system that identifies underrepresented segments in a user population and dynamically generates tasks to gather feedback, using machine learning and artificial intelligence to improve data quality by labeling and storing user feedback, thereby modifying the presentation of items to better represent diverse user attributes and preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data sets are collected from general user populations, then data collection is efficient and straightforward, but the data becomes unrepresentative and biased toward overrepresented segments

Engineering Contradiction:
Improvedata collection efficiencyVSAvoiddata representativeness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system continuously monitors data set composition against population demographics, identifies underrepresented segments, and dynamically adjusts task distribution to solicit feedback from these groups. This closed-loop feedback mechanism ensures the data set progressively becomes more representative while maintaining collection efficiency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The task distribution system dynamically adapts based on real-time analysis of data set composition. Rather than static sampling, the system continuously modifies which user segments receive tasks to ensure proportional representation across all population segments, resolving the contradiction between efficient collection and representativeness.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If feedback is solicited from all user segments uniformly, then data collection is simple, but underrepresented segments remain insufficiently represented

Engineering Contradiction:
Improvedata collection simplicityVSAvoidsegment representation accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system applies different sampling strategies to different user segments based on their representation status. Overrepresented segments receive standard sampling while underrepresented segments receive targeted oversampling, allowing simple uniform collection where appropriate while applying precision correction where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system modifies the sampling parameter (task distribution probability) based on segment representation metrics. By dynamically adjusting this parameter, the system maintains operational simplicity for most segments while achieving precise representation control for underrepresented groups.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If behavioral models are trained on existing data sets, then model generation is efficient, but models inherit biases from unrepresentative data

Engineering Contradiction:
Improvemodel generation speedVSAvoidmodel accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis of data set composition before model training, identifies representation gaps, and solicits additional feedback from underrepresented segments. This preliminary action ensures the training data is representative before model generation begins, preventing bias inheritance while maintaining efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback loops that monitor data representativeness throughout the model generation process. If biases are detected in the training data, the system solicits additional targeted feedback to correct these issues before final model training, ensuring accuracy without sacrificing overall efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11531655B2Automatically improving data quality
Publication Date: 2022.12.20 GOOGLE LLC
  • US11531655B2 patent drawing
  • US11531655B2 patent drawing
  • US11531655B2 patent drawing

AI summary

Methods, systems, and computer readable medium include receiving, from a user device, a request for a digital component, determining an attribute of the user based information provided by the user or information contained in the request, identifying a behavioral model corresponding to the attribute, dynamically altering a presentation of an item depicted by the digital component based on the identified behavioral model, determining that the user corresponds to an underrepresented segment of a user population in a database containing information about the item, and in response, generating a digital component that includes the dynamically altered presentation of the item, solicits feedback from the user regarding the item, and includes a feedback mechanism, updating the database to include the feedback obtained, and modifying presentation of the item when distributed to other users having the attribute of the user based, at least in part, on the feedback obtained.