Influence-Function Training Sample Selection for Efficient Data Screening

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for selecting high-quality training data for machine learning models are inefficient and reliant on manual selection or external models, leading to suboptimal training efficiency.

Innovation Solution

A method and apparatus for determining training samples using a preset influence function to quantify the influence degree of candidate samples relative to a standard sample, allowing for the selection of high-quality training samples without external evaluation, thereby improving screening and training efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual selection or external quality evaluation models are used to select high-quality training data, then data quality can be maintained, but screening efficiency is low and reliance on external models increases

Engineering Contradiction:
Improvedata qualityVSAvoidscreening efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system uses the machine learning model itself to evaluate the quality of training samples through influence function calculation, eliminating the need for external evaluation models. The model autonomously determines which samples are high-quality by measuring their influence on model performance, thereby improving screening efficiency while maintaining data quality standards.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The influence function serves as an intermediary mechanism that bridges the gap between training samples and model performance evaluation. Instead of using external models or manual assessment, the influence function calculates the relationship between sample characteristics and model loss, providing an efficient automated quality assessment that resolves the contradiction between maintaining quality and improving efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual selection methods are used for training data, then data quality can be controlled, but the process is time-consuming and inefficient

Engineering Contradiction:
Improvedata qualityVSAvoidscreening time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual selection process with an automated computational system based on influence function calculation. This substitution eliminates the time-consuming nature of manual review while maintaining quality control through mathematical evaluation of sample influence on model performance, directly addressing the time loss issue.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If external quality evaluation models are used, then training data quality can be ensured, but the system becomes more complex and dependent on additional models

Engineering Contradiction:
Improvetraining data qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the quality evaluation capability from external models and integrates it directly into the training process through influence function calculation. By taking out the evaluation function and embedding it within the existing model framework, the system maintains data quality assurance while reducing overall system complexity and eliminating dependency on separate external models.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250299100A1Method for determining training sample, medium, electronic device and program product
Publication Date: 2025.09.25 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250299100A1 patent drawing
  • US20250299100A1 patent drawing
  • US20250299100A1 patent drawing

AI summary

The present disclosure provides a method for determining a training sample, a medium, an electronic device and a program product. The method includes acquiring a plurality of candidate training samples and a standard training sample, the candidate training sample including one of a text-type sample, an image-type sample, and an audio-type sample; for each candidate training sample, determining an influence degree of the candidate training sample relative to the standard training sample according to a preset influence function, the influence function being a function representing a relationship between the influence degree with a first loss of a machine learning model on the candidate training sample and a second loss of the machine learning model on the standard training sample; and determining a target training sample from the plurality of candidate training samples according to the influence degree, the target training sample being used for training the machine learning model.