Influence-Function Training Sample Selection for Efficient Data Screening
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for selecting high-quality training data for machine learning models are inefficient and reliant on manual selection or external models, leading to suboptimal training efficiency.
Innovation Solution
A method and apparatus for determining training samples using a preset influence function to quantify the influence degree of candidate samples relative to a standard sample, allowing for the selection of high-quality training samples without external evaluation, thereby improving screening and training efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual selection or external quality evaluation models are used to select high-quality training data, then data quality can be maintained, but screening efficiency is low and reliance on external models increases
Solution Approach 1:
The system uses the machine learning model itself to evaluate the quality of training samples through influence function calculation, eliminating the need for external evaluation models. The model autonomously determines which samples are high-quality by measuring their influence on model performance, thereby improving screening efficiency while maintaining data quality standards.
Solution Approach 2:
The influence function serves as an intermediary mechanism that bridges the gap between training samples and model performance evaluation. Instead of using external models or manual assessment, the influence function calculates the relationship between sample characteristics and model loss, providing an efficient automated quality assessment that resolves the contradiction between maintaining quality and improving efficiency.
2Reliability
If manual selection methods are used for training data, then data quality can be controlled, but the process is time-consuming and inefficient
Solution Approach 1:
The patent replaces the mechanical manual selection process with an automated computational system based on influence function calculation. This substitution eliminates the time-consuming nature of manual review while maintaining quality control through mathematical evaluation of sample influence on model performance, directly addressing the time loss issue.
3Reliability
If external quality evaluation models are used, then training data quality can be ensured, but the system becomes more complex and dependent on additional models
Solution Approach 1:
The patent extracts the quality evaluation capability from external models and integrates it directly into the training process through influence function calculation. By taking out the evaluation function and embedding it within the existing model framework, the system maintains data quality assurance while reducing overall system complexity and eliminating dependency on separate external models.
Data Source
AI summary
The present disclosure provides a method for determining a training sample, a medium, an electronic device and a program product. The method includes acquiring a plurality of candidate training samples and a standard training sample, the candidate training sample including one of a text-type sample, an image-type sample, and an audio-type sample; for each candidate training sample, determining an influence degree of the candidate training sample relative to the standard training sample according to a preset influence function, the influence function being a function representing a relationship between the influence degree with a first loss of a machine learning model on the candidate training sample and a second loss of the machine learning model on the standard training sample; and determining a target training sample from the plurality of candidate training samples according to the influence degree, the target training sample being used for training the machine learning model.


