Influence-Function Training Sample Selection for Efficient Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for selecting high-quality training data for machine learning models are inefficient and rely heavily on manual selection or external models, leading to low screening efficiency.

Innovation Solution

A method and apparatus for determining training samples using a preset influence function to quantify the influence degree of candidate samples relative to a standard sample, allowing for the selection of high-quality training samples without external models, utilizing gradient and Hessian matrix calculations to optimize the training process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual selection or external quality evaluation models are used to select high-quality training data, then data quality can be maintained, but screening efficiency becomes low and reliance on external models increases

Engineering Contradiction:
Improvescreening efficiencyVSAvoiddata quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements self-service by enabling the model to automatically evaluate and select its own training samples using influence function calculations. The system computes the influence degree of each candidate sample on the model's loss function and autonomously selects samples with positive influence, eliminating dependence on external evaluation models while maintaining data quality standards

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies parameter changes by transforming the sample selection criterion from qualitative assessment to quantitative measurement. It introduces the influence degree parameter, calculated through influence function mathematics, to objectively evaluate each candidate sample's impact on model performance. This parameter-based approach enables efficient automated screening while ensuring reliable selection of high-quality training data

Inventive Principle:
Principle #35Parameter changes

2Reliability

If more training samples are used to improve model performance, then training effectiveness increases, but training time and computational resources increase

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the most valuable training samples from the candidate pool by calculating and filtering based on influence degree. Instead of using all available samples, it selectively extracts samples with positive influence on the loss function, thereby reducing the total number of samples required for training while maintaining or improving model performance and reducing training time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by selecting a subset of training samples that provides sufficient training effectiveness without requiring the full dataset. By choosing samples with the highest positive influence degrees, the system achieves good model performance with fewer samples, avoiding the excessive use of computational resources and time associated with training on complete datasets

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4621656A1Method and apparatus for determining training sample, medium, electronic device and program product
Publication Date: 2025.09.24 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • EP4621656A1 patent drawingFigure 1~2
  • EP4621656A1 patent drawingFigure 3
  • EP4621656A1 patent drawingFigure 4

AI summary

The present disclosure relates to a method and apparatus for determining a training sample, a medium, an electronic device and a program product, which relate to the field of computer technologies. The method includes acquiring a plurality of candidate training samples and a standard training sample; for each candidate training sample, determining an influence degree of the candidate training sample relative to the standard training sample according to a preset influence function; and determining a target training sample from the plurality of candidate training samples according to the influence degree. High-quality training data beneficial to the training of the machine learning model can be obtained through screening without relying on an external evaluation model, thereby greatly improving the screening efficiency of the training sample and the training efficiency of the machine learning model.