Sample distribution adaptive adjustment method and system based on model training feedback driving

By collecting training feedback indicators in real time to identify weak categories, dynamically adjusting sampling weights and combining data enhancement, we solve the problem of sample distribution and model disconnection in deep learning model training, achieve real-time, refined and automated adjustment of model training effects, and improve the model's performance and generalization capabilities in weak and long-tail categories.

CN120804706APending Publication Date: 2025-10-17BEIJING OLA TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510920766.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

During the existing deep learning model training process, the sample distribution is disconnected from the model training, and there is a lack of real-time feedback mechanism, resulting in insufficient training of small category samples, long-tail categories or difficult-to-learn samples, limited model generalization ability, and existing methods making it difficult to achieve refined and automated sampling adjustment.

Method used

By collecting training feedback indicators in real time, identifying weak categories, dynamically adjusting sampling weights, and combining data enhancement or generation mechanisms, a training-data closed-loop linkage is formed to achieve real-time, refined, and automated adjustment of model training effects.

Benefits of technology

It improves the training effect and generalization ability of the model in weak categories and long-tail categories, improves the stability and efficiency of the training process, supports adaptability to multi-task scenarios, and prevents overfitting and sample redundancy disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804706A_ABST
    Figure CN120804706A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model training sample processing, in particular to a sample distribution adaptive adjustment method and system based on model training feedback driving, and the method comprises the following steps: S1, collecting feedback indexes in a training process in real time in a model training period, and obtaining a feedback index set; s2, weak item categories are analyzed and recognized based on the feedback index set, and an optimization instruction is generated; s3, dynamically adjusting the sampling weight to form an updated sampling strategy; s4, samples are collected according to a sampling strategy, and weak item category samples are supplemented; s5, putting the optimized sample set into the next round of training, and circulating the steps S1 to S4 until the training indexes of the learning effects of all key categories meet the termination condition; and acquiring a sampling strategy and a sample set. According to the method, real-time, refined and automatic sampling adjustment can be realized by taking the actual performance of model training as a core feedback index, so that the training effect and generalization ability of the model in key areas such as weak-term categories and long-tail categories are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of model training sample processing, and in particular to a sample distribution self-adaptive adjustment method and system based on model training feedback driving. BACKGROUND

[0002] In the current training process of deep learning models, the construction and sampling strategy of training data usually rely on static rules or artificial experience settings, such as fixed class ratio division, random sampling, etc. Such methods have the following problems in practical application:

[0003] 1. Sample distribution is disconnected from model training: Due to the lack of process feedback guidance, the construction of each class sample in the training set often cannot match the actual learning needs of the model, leading to insufficient training of small class samples, long-tail classes or difficult-to-learn samples;

[0004] 2. Lack of adjustment mechanism in training process: static sampling strategy cannot dynamically optimize the sampling distribution according to the current performance of the model, and some classes may be "ignored" or "invalidated" in training;

[0005] 3. Limited model generalization ability: the performance of the model in key classes or extreme scenarios is often restricted due to insufficient coverage of training samples.

[0006] In order to solve the above problems, the industry has gradually explored data construction methods based on quantitative indicators. For example, the number of steps each class sample is sampled in training is calculated in real time to ensure that all class samples reach a preset minimum exposure frequency; if the actual training frequency of a certain class sample is lower than the target value, the system can dynamically increase its sampling weight, automatically supplement samples or re-schedule the sampling strategy; thus a closed-loop data construction process based on "sample exposure feedback in training" is established.

[0007] This scheme effectively alleviates the problem of sample imbalance at the beginning of training, but its feedback mechanism is mainly based on the frequency signal of "whether the sample has been sampled" rather than the evaluation of the actual learning effect of the model on the sample. Although it has strong adaptive ability at the sample level and can make sampling adjustments according to the exposure in the training process, it still has deficiencies in the fine feedback and optimization of model learning, which are as follows:

[0008] 1. Unable to identify whether the learning effect of the model on the fully sampled class meets the requirements;

[0009] 2. Difficult to automatically diagnose and accurately fill in the "difficult-to-learn samples" or "weak classes" exposed in training;

[0010] 3. Lack of training performance integration mechanism across tasks and modalities, making it difficult to adapt to complex multi-task training scenarios.

[0011] In addition, existing strategies such as "active learning" or "curriculum learning" also introduce training feedback information, but the core focuses on sample uncertainty, confidence ranking or training sequence control, usually relies on manual configuration strategy, and fails to build a systematic and automated sampling optimization mechanism.

[0012] In summary, there is currently a lack of a general method and system that takes model training actual performance as the core feedback indicator and realizes real-time, refined and automated sampling adjustment. It is urgent to build a "training feedback-data distribution-sampling strategy" closed-loop linkage mechanism to improve the training effect and generalization ability of the model in key areas such as weak item categories and long-tail categories. SUMMARY

[0013] One of the purposes of the present application is to provide a sample distribution adaptive adjustment method driven by model training feedback, which can take model training actual performance as the core feedback indicator and realize real-time, refined and automated sampling adjustment to improve the training effect and generalization ability of the model in key areas such as weak item categories and long-tail categories.

[0014] In order to achieve the above purpose, a sample distribution adaptive adjustment method driven by model training feedback is provided, comprising the following steps:

[0015] S1. In the model training period, real-time acquisition of feedback indicators in the training process is performed, and structured analysis is performed according to the category or label dimension to obtain a feedback indicator set;

[0016] S2. Based on the feedback indicator set analysis, the learning effect of the model in each category or task dimension is judged, and the categories whose training effect does not meet the standard are identified and marked as weak item categories;

[0017] S3. The sampling weight of the weak item category is dynamically increased, and the sampling weight of the non-weak item category is reduced according to the demand or kept at the current sampling proportion to form an updated sampling strategy;

[0018] S4. After obtaining new samples according to the sampling strategy, it is judged whether the number of samples of the weak item category meets the preset minimum training requirement, and if not, the sample enhancement or generation mechanism is triggered to supplement the samples to obtain a sample set;

[0019] S5. The optimized sample set is put into the next round of training, and steps S1-S4 are cycled until the training indicators of the learning effect of all key categories meet the termination condition; the sampling strategy and sample set in the cycling process are obtained; if the training feedback of a certain category meets the standard for continuous multiple rounds, or the sampling strategy benefit is lower than the preset threshold, the supplement of samples is suspended and the previous sampling strategy is restored.

[0020] Further, the feedback indicators include, but are not limited to, at least two of loss, recall, precision, F1 score, class accuracy, and confusion matrix.

[0021] Further, the step of identifying the class whose training effect is not up to standard in the step S2 includes: weighting and fusing multiple feedback indicators in the feedback indicator set or performing rule judgment, comparing with a set threshold, and determining the underperforming class.

[0022] Further, the weight adjustment formula for dynamically increasing the sampling weight in the step S3 is as follows:

[0023] For the class c, the sampling weight is updated to:

[0024] wc(t+1)=wc(t)+α·(Tc-Mc(t))

[0025] wherein wc(t): current class sampling weight; Tc: class target indicator; Mc(t): current training feedback actual value; and a: adjustable coefficient, determining the adjustment amplitude.

[0026] Further, the sample supplement method in the step S4 specifically adopts a data enhancement algorithm or a sample generation model, and metadata information is attached to the enhanced sample or the generated sample, the metadata information including generation method, timestamp, and version number; and the trigger logic for the sample enhancement or automatic generation is:

[0027] If the current weight of a certain class satisfies:

[0028] wc(t)<γ·Tc

[0029] then the data enhancement or new sample generation logic is triggered to supplement the training sample, and gamma is a weak item determination coefficient set by the system.

[0030] The second purpose of the present application is to provide a sample distribution self-adaptive adjustment system based on model training feedback driving, comprising the following modules:

[0031] A training feedback collection module: used for collecting feedback indicators in the training process in real time during the model training period, and performing structured analysis according to the class or label dimension to obtain a feedback indicator set;

[0032] A feedback analysis and decision module: used for analyzing and judging the learning effect of the model in each class or task dimension based on the feedback indicator set, identifying the class whose training effect is not up to standard, and marking it as a weak item class;

[0033] A sampling adjustment module is configured to dynamically increase the sampling weight of the weak item category and decrease the sampling weight of the non-weak item category according to the demand or keep the current sampling proportion, so as to form an updated sampling strategy.

[0034] A data sampling and enhancement module is configured to determine whether the sample quantity of the weak item category meets the preset minimum training demand after obtaining new samples according to the sampling strategy, and trigger a sample enhancement or generation mechanism to supplement samples if the sample quantity does not meet the preset minimum training demand, so as to obtain a sample set.

[0035] A training-data closed-loop linkage module is configured to put the optimized sample set into the next round of training, and cycle the above modules until the training indicators of the learning effects of all key categories meet the termination condition; the sampling strategy and the sample set in the cycling process are obtained; if the training feedback of a category meets the standard for continuous multiple rounds, or the sampling strategy benefit is lower than a preset threshold, the supplement of samples is suspended and the previous sampling strategy is restored.

[0036] Further, the feedback indicators include, but are not limited to, at least two of loss, recall, precision, F1 score, category accuracy, and confusion matrix.

[0037] Further, the way of identifying the categories with substandard training effects in the feedback analysis and decision module is that multiple feedback indicators in the feedback indicator set are weighted and fused or judged by rules, and compared with a set threshold to determine the categories with substandard performance.

[0038] Further, the weight adjustment formula of dynamically increasing the sampling weight in the sampling adjustment module is as follows:

[0039] For a category c, the sampling weight is updated in the t+1th round as follows:

[0040] wc(t+1) = wc(t) + a · (Tc - Mc(t))

[0041] wherein wc(t) is the current category sampling weight; Tc is the category target indicator; Mc(t) is the current training feedback actual value; a is an adjustable coefficient that determines the adjustment amplitude.

[0042] Further, the way of supplementing samples in the data sampling and enhancement module specifically adopts a data enhancement algorithm or a sample generation model, and the enhanced samples or generated samples are attached with metadata information, and the metadata information includes the generation way, the timestamp, and the version number; and the trigger logic of the sample enhancement or automatic generation is as follows:

[0043] If the current weight of a category meets the following condition:

[0044] wc(t) < y · Tc

[0045] Then trigger data augmentation or new sample generation logic, supplement training samples, and gamma is the system's weak item judgment coefficient.

[0046] Principles and advantages:

[0047] Unlike traditional solutions, this solution focuses on: multi-dimensional feedback index fusion, not limited to a single performance dimension; complete adjustment closed loop in the same training cycle, support continuous optimization instead of discrete batch optimization; sampling adjustment, data augmentation, termination mechanism have linkage trigger and automatic return ability, adapt to complex task and model evolution demand.

[0048] Compared with existing static sampling or frequency control solutions, this solution has the following significant technical features and advantages in feedback granularity, adjustment mechanism, system linkage and task adaptability:

[0049] 1. Training feedback driven data optimization mechanism

[0050] This solution breaks through the limitations of traditional "static sampling strategy setting" or "single round feedback optimization", and builds an automatic evolution mechanism for data construction driven by training feedback, realizing the dynamic linkage of training feedback and data structure. Its key features include: 1) support collecting multi-dimensional training feedback indicators (such as loss, recall, precision, F1, etc.); 2) organize and analyze feedback indicators according to categories, tasks, etc.; 3) realize the automatic closed-loop linkage of sample sampling strategy, enhancement generation and training configuration. This mechanism not only identifies whether the sample is "sampled", but also judges whether it is "learned effectively", dynamically adjusts the sample data distribution through index convergence and deviation judgment.

[0051] 2. Real-time identification and active adjustment of weak categories, the system supports real-time evaluation of learning effect of each category within the training cycle, and through setting threshold, trend judgment or fusion analysis, etc. accurately locate the "short board category" or "difficult sample" in the model training process. Different from the traditional method of classifying frequency statistics, this solution: 1) uses real-time feedback indicators to locate performance bottlenecks; 2) can generate optimization priorities based on the learning effect of weak categories; 3) dynamically binds the adjustment target with sampling strategy and sample enhancement to form a linkage response.

[0052] 3. Sampling adjustment and sample enhancement dual strategy fusion. Unlike existing technologies that only update weights based on performance feedback, this solution introduces a "adjustment + enhancement" collaborative optimization mechanism: 1) weak categories not only increase the sampling probability, but also automatically trigger the enhancement or generation strategy; 2) enhanced samples are accompanied by metadata labels (such as generation method, timestamp, etc.) for version management; 3) decide whether to execute the enhancement strategy according to whether the sample meets the minimum training coverage requirement. This mechanism ensures that the data compensation strategy is accurate, automatic and transparent, solving the problem of training blind area caused by insufficient samples or category sparsity.

[0053] 4. Closed-loop control capability supporting automatic rollback of training strategy, with training indicator convergence detection mechanism, if the feedback of two consecutive training rounds reaches the preset target, the enhancement operation will be automatically paused and the sampling strategy will be restored to prevent overfitting or sample redundancy disturbance. This strategy rollback mechanism enhances the stability of model training (avoids training fluctuations), convergence (automatically terminates enhancement), and traceability (process can be recorded and rolled back).

[0054] 5. Support for universal adaptability of multi-task scenarios and platforms. This solution supports feedback granularity and adjustment methods for different types of tasks, including but not limited to: 1) image classification, natural language processing, recommendation system, multi-label learning; 2) can be connected with mainstream training frameworks (such as PyTorch, TensorFlow); each module has a decoupled structure, suitable for plug-in integration or containerized deployment. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1 A logic block diagram of a sample distribution adaptive adjustment system based on model training feedback driving according to an embodiment of the present application;

[0056] Figure 2 A flowchart of a sample distribution adaptive adjustment method based on model training feedback driving according to an embodiment of the present application. DETAILED DESCRIPTION

[0057] The following will be further described in detail through specific embodiments:

[0058] EMBODIMENT

[0059] A sample distribution adaptive adjustment method based on model training feedback driving, as shown in Figure 1 , Figure 2 , including the following steps:

[0060] S1, through the training feedback acquisition module, real-time acquisition of feedback indicators in the training process during the model training period, and structured analysis according to categories or label dimensions to obtain a feedback indicator set; the feedback indicators are obtained through training log information and model configuration information (batch_size, epoch); the feedback indicators include but are not limited to at least two of loss, recall, precision, F1 score, category accuracy, and confusion matrix.

[0061] In the initial training data preparation stage in the model training cycle: the original training data set is divided into three categories: main task samples, long tail samples and adversarial samples to support the subsequent training feedback driven sample optimization mechanism. In the initial stage, to ensure that the samples of each category are representative and effective for training, the sampling strategy is supported by configuring the task priority or category importance, and the training sample pool is constructed according to the preset proportion.

[0062] For example, the initial sampling distribution strategy can be set as follows:

[0063] 1. The proportion of main task samples is set to 80% to ensure the performance of the main task;

[0064] 2. The proportion of long tail samples is set to 15% to improve coverage and generalization ability;

[0065] 3. The proportion of adversarial samples is set to 5% to enhance the robustness of the model.

[0066] When the training configuration is batch_size=64, epoch=5, and the total training data amount is about 20000 samples, the corresponding number of samples is automatically extracted according to the above proportion, or the insufficient categories are supplemented by data collection / enhancement mechanism to complete the construction of the initial sample pool. At the same time, the sample pool supports subsequent on-demand update.

[0067] S2, the feedback analysis and decision module analyzes and judges the learning effect of the model in each category or task dimension based on the feedback index set, identifies the categories whose training effect does not meet the standard, and marks them as weak categories; the marked weak categories can be sorted into a weak category list, and the optimization priority is set. This not only facilitates the judgment of "whether to be sampled", but also facilitates the evaluation of "whether to be effectively learned". The step of identifying the categories whose training effect does not meet the standard in step S2 includes: weighting and fusing or rule judging multiple feedback indicators in the feedback indicator set, comparing with the set threshold, and determining the underperforming categories.

[0068] S3, the sampling adjustment module dynamically increases the sampling weight of the weak category, and reduces the sampling weight of the non-weak category according to the demand or keeps the current sampling proportion, forms an updated sampling strategy, and outputs; the weight adjustment formula for dynamically increasing the sampling weight in step S3 is as follows:

[0069] For category c, its sampling weight is updated to:

[0070] wc(t+1) = wc(t) + a · (Tc - Mc(t))

[0071] where wc(t): the current category sampling weight; Tc: the category target index; Mc(t): the current training feedback actual value; a: adjustable coefficient, determines the adjustment amplitude.

[0072] S4, after obtaining new samples according to the sampling strategy, judging whether the number of samples of the weak item category meets the preset minimum training requirement, if not, triggering the sample enhancement or generation mechanism to supplement the samples, and obtaining the sample set; the triggering logic of the sample enhancement or automatic generation is:

[0073] If the current weight of a certain category meets:

[0074] wc(t)<γ·Tc

[0075] Then trigger the data enhancement or new sample generation logic to supplement the training samples, and γ is the weak item judgment coefficient set by the system.

[0076] The sample supplement method in step S4 specifically adopts a data enhancement algorithm or a sample generation model, and the enhanced samples or generated samples are attached with metadata information, and the metadata information includes the generation method, timestamp, and version number.

[0077] S5, put the optimized sample set into the next round of training, and cycle steps S1-S4 until the training indicators of the learning effect of all key categories meet the termination condition; obtain the sampling strategy and sample set in the cycle process; write all sampling strategies and adjustment records, sample sets and version information into the data storage module to support subsequent auditing, model rollback and training reproduction. If the training feedback of a certain category meets the standard for continuous multiple rounds, or the sampling strategy benefit is lower than the preset threshold, then suspend the sample supplement and restore the previous sampling strategy (including the sampling strategy of the previous round or the initial sampling strategy). Support the automatic closed-loop process of training-feedback-adjustment-retraining, and have a strategy rollback mechanism to prevent overfitting or sample redundancy disturbance, improve the workflow efficiency. And the sampling strategy, sample set, triggering data enhancement or new sample generation logic in the cycle process can also be used as training data to optimize the sampling strategy, weak item judgment coefficient, etc. through a conventional learning model, to achieve fine-tuning and improve the efficiency of the entire process.

[0078] A sample distribution self-adaptive adjustment system based on model training feedback driving, the system is deployed as an independent module on a deep learning training platform (such as TensorFlow, PyTorch), which can be integrated with data management module, sample pool interface, training log collection interface and other subsystems to support the entire sample optimization process including the following modules:

[0079] The training feedback collection module is configured to collect feedback indexes in real time during the model training period, and perform structured analysis according to categories or label dimensions to obtain a feedback index set. The feedback indexes are obtained through training log information and configuration information, and the feedback indexes include, but are not limited to, at least two of loss, recall, precision, F1 score, category accuracy, and a confusion matrix.

[0080] The feedback analysis and decision module is configured to analyze and determine the learning effect of the model in various categories or task dimensions based on the feedback index set, identify categories with substandard training effects, and mark the categories as weak item categories. The marked weak item categories can be sorted into a weak item category list, and an optimization priority is set. In this way, it is not only convenient to determine whether to be sampled, but also convenient to evaluate whether to be effectively learned. The way in which the feedback analysis and decision module identifies categories with substandard training effects is that multiple feedback indexes in the feedback index set are weighted and fused or judged by rules, compared with a set threshold, and categories with substandard performance are determined. For example:

[0081] 1. The recall of category A (a long-tail category) is 0.65, which is lower than the target value 0.75.

[0082] 2. The loss of category B (an adversarial category) is 1.5, which is higher than the target 1.0.

[0083] The sampling adjustment module is configured to dynamically increase the sampling weight of the weak item category, and to decrease the sampling weight of the non-weak item category or maintain the current sampling proportion according to the demand, to form an updated sampling strategy. The weight adjustment formula for dynamically increasing the sampling weight in the sampling adjustment module is as follows:

[0084] For category c, the sampling weight is updated to:

[0085] wc(t+1)=wc(t)+α·(Tc-Mc(t))

[0086] wherein wc(t) is the current category sampling weight, Tc is the category target index, Mc(t) is the current training feedback actual value, and a is an adjustable coefficient that determines the adjustment amplitude.

[0087] For example, the weight of category A is increased from 2% to 5%, and is preferentially sampled in the next round of training.

[0088] The data sampling and enhancement module is configured to determine whether the number of samples of the weak item category meets the preset minimum training demand after obtaining new samples according to the sampling strategy. If not, a sample enhancement or generation mechanism is triggered to supplement the samples, and a sample set is obtained. The trigger logic of the sample enhancement or automatic generation is as follows:

[0089] If the current weight of a category meets:

[0090] wc(t) < γ · Tc

[0091] Then trigger data augmentation or new sample generation logic, supplement training samples, and γ is the weak item determination coefficient set by the system.

[0092] The sample supplement method in the data sampling and augmentation module adopts a data augmentation algorithm (such as T5, BART) or a sample generation model, and the augmented samples or generated samples are accompanied by metadata information, including generation method, timestamp, version number. For example, the augmentation process includes:

[0093] 1. Set the generation parameters (such as temperature = 0.9);

[0094] 2. Synthesize semantic diverse samples;

[0095] 3. Automatic labeling, including generation method, timestamp, augmented version, and other metadata.

[0096] 4. New samples will be included in the sample pool for subsequent training.

[0097] Training-data closed-loop module: used to put the optimized sample set into the next round of training, cycle the above 4 modules until the training indicators of all key categories meet the termination condition; obtain the sampling strategy and sample set in the cycle; write all sampling strategies and adjustment records, sample sets and version information into the data storage module, support subsequent audit, model rollback and training reproduction. If the training feedback of a certain category meets the standard for several consecutive rounds, or the sampling strategy yield is lower than the preset threshold, temporarily suspend the supplement of samples and restore the previous sampling strategy. This strategy rollback mechanism ensures that the training system has the ability of "self-diagnosis, self-convergence, and self-suspension" in the process of continuous optimization, which is an important mechanism that distinguishes it from traditional static sampling or batch adjustment methods.

[0098] The above modules can be independently encapsulated in the form of containerized deployment according to the actual training platform components, or used as integrated plug-in modules of the training platform, with good engineering adaptability and deployment expansion capability.

[0099] Comparison with prior art

[0100] To highlight the uniqueness of this scheme in terms of technical mechanism and system structure, this scheme will be compared with the current typical sampling strategy optimization scheme, which will be reflected in the following aspects:

[0101] 1. Different from static sampling and traditional frequency control sampling mechanism

[0102] 1) Static sampling strategy is usually based on preset category proportion or fixed sampling probability, which cannot respond to performance differences in the training process;

[0103] 2) Frequency control type sampling mechanism (such as the "minimum step coverage" method) can improve sample exposure, but only counts the number of samples, and cannot evaluate whether the model has truly "learned" the sample;

[0104] 3) The present scheme comprehensively measures the model training effect by fusing multi-dimensional feedback indicators (such as recall, loss, F1, etc.), and builds a dynamic linkage relationship between data distribution and learning feedback, thereby realizing effect-driven optimization of sample sampling instead of "exposure-driven optimization".

[0105] 2, Different from the existing "adjust sampling according to the feedback of the last batch"

[0106] In the prior art, the performance indicators of the last batch of trained models, such as clustering accuracy, are analyzed to dynamically adjust the sampling weight distribution of the current batch of training data. The core mechanism of this method is that after the model completes a round of training, the system evaluates its learning effect in each "domain category", and if some categories perform poorly, they are classified as "to-be-improved domains" and their sampling probability is increased in the next batch to achieve a balanced distribution of samples in different domains, thereby preventing loss shock or catastrophic forgetting phenomenon in the pre-training process.

[0107] Although this method introduces training feedback information to optimize the data sampling strategy, its technical mechanism has obvious limitations. First, it uses a cross-batch feedback-adjustment structure, that is, the sampling strategy of the current training data is adjusted based on the training results of the last batch of models, which has a long reaction period and a lagging response. Second, its sampling adjustment dimension is mainly the data proportion change of "domain categories", and it cannot be refined to dynamically optimize at the granularity of labels, tasks, and modalities. Third, this method usually only adjusts the sampling weight itself, lacks the coordinated design of data augmentation, sampling termination, and other mechanisms, and is difficult to build a complete data optimization closed-loop process.

[0108] In comparison, the technical solution proposed by the present scheme has significant differences and technical progress in mechanism. The present scheme supports real-time collection of multi-dimensional feedback information of the model within the training period, including but not limited to loss, recall, precision, F1 score, etc., and can identify weak item categories in the current training of the model by category or label dimension. These feedback information will directly drive the adjustment of sampling weight and can trigger sample enhancement strategies such as automatically synthesizing new samples or performing specific data completion operations. At the same time, the present scheme introduces a training state judgment mechanism that can automatically suspend enhancement operations and restore the initial sampling strategy when the training indicators continuously meet the standards, thereby controlling the training stability and preventing overfitting or sample structure disturbance.

[0109] In summary, although the present scheme has some commonality with the above-mentioned existing scheme in terms of "optimizing the sampling strategy using training feedback", there are essential differences between the two in terms of feedback processing period, complexity of adjustment mechanism, system linkage capability, and target orientation. The present scheme, with its continuous closed-loop and self-adaptive evolution data optimization capability, solves the problems of feedback response lag, coarse optimization granularity, and isolated process in existing methods, and has clear technical breakthroughs.

[0110] 3. Closed-loop, self-driven data construction system

[0111] The present scheme constructs a continuous closed-loop data construction process from training feedback collection, weak item identification, strategy adjustment, data enhancement, and retraining, and has the following technical advantages:

[0112] 1) Supports automatic identification of training bottlenecks and optimization of sampling strategy during training, rather than discrete static adjustment;

[0113] 2) Data enhancement and sampling adjustment linkage can dynamically fill in sample distribution gaps;

[0114] 3) Supports automatic judgment of training convergence and performs strategy rollback and enhancement pause operations to prevent training overfitting or redundant disturbance.

[0115] 4. Wide applicability across tasks and platforms

[0116] The present scheme supports image, text, speech, structured data, and other multi-modal training tasks, and the modules are decoupled, making it easy to integrate as a plug-in into existing training platforms. It is compatible with mainstream frameworks such as PyTorch and TensorFlow, and can work with active learning, curriculum learning, and data governance mechanisms to build a more flexible and evolutionary intelligent training system.

[0117] Advantages of the present scheme

[0118] The present scheme realizes dynamic coupling optimization of data construction and model performance through a training feedback-driven sample distribution self-adaptive adjustment method, and has obvious advantages in training efficiency, model stability, generalization ability, and system engineering compared to existing technologies, as follows:

[0119] 1. Significantly improves model performance on key categories

[0120] The present scheme continuously collects and analyzes performance feedback of the model during training, actively identifies categories and difficult-to-learn samples with weak performance, and dynamically improves their sampling weight and sample coverage. Verification in multiple real training tasks shows that:

[0121] 1) The recall of long-tail categories can be improved by an average of 5%-15%, with more significant advantages in extreme data distribution;

[0122] 2) Multi-label task, micro-average F1 score increased by 6%-10%, macro-average F1 score increased more obviously;

[0123] 3) After introducing the enhancement mechanism, the training loss of the adversarial sample decreased by about 30% from the original value, and the training stability was enhanced.

[0124] This mechanism avoids the phenomenon of "weak item classes being ignored" in traditional static sampling, effectively improving the model's learning ability in key areas.

[0125] 2、Enhance the controllability and diagnosability of the training process

[0126] By introducing multi-dimensional feedback indicators (such as loss, recall, F1, confusion matrix, etc.) as the basis for sampling adjustment, the system realizes continuous monitoring and strategy response to weak items in the training process. At the same time, by setting up training indicator convergence detection and enhancement strategy suspension mechanism, the training process has the following characteristics:

[0127] 1) Diagnosability: Identify learning bottlenecks and track sample structure defects based on feedback indicators from each round;

[0128] 2) Adjustability: Adjust sampling weights and sample generation strategies without manual intervention;

[0129] 3) Traceability: All sampling strategies and enhancements are automatically labeled with meta-information, supporting task version management and rollback.

[0130] Compared with the existing technology that only performs "result feedback → weight update" between batches, this solution realizes closed-loop adjustment and state-aware control within the training cycle, which is more suitable for building long-term maintainable large-scale training systems.

[0131] 3、Improve the overall engineering efficiency of the system and the robustness of the model

[0132] This solution integrates training feedback analysis, sampling strategy optimization, sample enhancement, and training scheduling modules, realizing an automated data-driven learning process that can reduce a large amount of manual tuning and debugging costs. In engineering applications, it has the following effects:

[0133] 1) The sample scheduling and supplement process is completely automatic and suitable for continuous training scenarios;

[0134] 2) The system can "fall back to the basic sampling strategy" at the intermediate stage of training, reducing the computational power and storage pressure;

[0135] 3) Enhanced samples have generation metadata, supporting subsequent training optimization, anomaly detection, and model interpretability analysis.

[0136] In addition, the system supports seamless docking with mainstream frameworks such as PyTorch and TensorFlow, facilitating integration with existing training pipelines and strong engineering adaptability.

[0137] 4. Generalization mechanism for adapting data evolution and model update

[0138] When the model is iterated or the training task is expanded, the traditional sampling strategy often needs to be reset or manually optimized. This scheme builds a training feedback-driven data structure evolution mechanism, enabling the system to have the following capabilities:

[0139] 1) Adaptive data distribution update: support for automatic discovery and adjustment of new categories and extreme data;

[0140] 2) Task switching scenario adaptation: unified feedback analysis framework for multi-modal, multi-label, and multi-task;

[0141] 3) Continuous learning capability guarantee: combined with the sampling strategy rollback mechanism, effectively preventing overfitting and catastrophic forgetting phenomenon.

[0142] The above features ensure the long-term applicability of the system in dynamic task environments, breaking through the structural limitations of previous static sampling methods.

[0143] The above is only an embodiment of the present application, and the common knowledge of the specific structure and characteristics of the scheme is not described too much. The person skilled in the art knows all the ordinary technical knowledge in the field of the application before the filing date or the priority date, can know all the prior art in this field, and has the ability to apply conventional experimental means before that date. The person skilled in the art can improve and implement the present scheme based on their own ability under the guidance of this application, and some typical known structures or known methods should not be an obstacle to the implementation of the present application by the person skilled in the art. It should be noted that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should be considered as the protection scope of the present application. These will not affect the effectiveness and practicality of the implementation of the present application. The scope of protection claimed in this application should be based on the content of its claims, and the specific implementation in the specification can be used to explain the content of the claims.

Claims

1. A sample distribution adaptive adjustment method based on model training feedback drive, characterized in that: The following steps are involved: S1. During the model training cycle, feedback indicators are collected in real time during the training process and structured analysis is performed by category or label dimension to obtain a set of feedback indicators. S2. Analyze and judge the learning effect of the model in each category or task dimension based on the feedback indicator set, identify categories where the training effect does not meet the standard, and mark them as weak categories; S3. Dynamically increase the sampling weight of weak categories and reduce the sampling weight of non-weak categories according to demand or maintain the current sampling ratio to form an updated sampling strategy; S4. After obtaining new samples according to the sampling strategy, determine whether the number of samples of the weak category meets the preset minimum training requirement. If not, trigger the sample enhancement or generation mechanism to supplement the samples to obtain a sample set; S5. Put the optimized sample set into the next round of training, and repeat steps S1-S4 until the training indicators of the learning effect of all key categories meet the termination conditions; Obtain the sampling strategy and sample set during the cycle; if the training feedback of a certain category meets the standard for multiple consecutive rounds, or the sampling strategy benefit is lower than the preset threshold, suspend the sample replenishment and resume the previous sampling strategy.

2. The method for adaptively adjusting sample distribution based on model training feedback drive according to claim 1, characterized in that: The feedback indicators include but are not limited to: at least two of loss, recall, precision, F1 score, category accuracy, and confusion matrix.

3. The method for adaptively adjusting sample distribution based on model training feedback drive according to claim 2, characterized in that: The step of identifying the category of training effect that does not meet the standard in step S2 includes: performing weighted fusion or rule judgment on multiple feedback indicators in the feedback indicator set, comparing with the set threshold, and determining the category of performance that does not meet the standard.

4. The method for adaptively adjusting sample distribution based on model training feedback drive according to claim 3, characterized in that: The weight adjustment formula for dynamically increasing the sampling weight in step S3 is as follows: For category c, its sampling weight is updated in round t+1 as: wc(t+1)=wc(t)+α·(Tc-Mc(t)) Among them, wc(t): is the current category sampling weight; Tc: is the category target indicator; Mc(t): is the actual value of the current training feedback; α: is the adjustable coefficient, which determines the adjustment range.

5. The method and system for adaptively adjusting sample distribution based on model training feedback drive according to claim 4, characterized in that: The method of sample supplementation in step S4 specifically adopts a data enhancement algorithm or a sample generation model, and metadata information is attached to the enhanced sample or generated sample, and the metadata information includes a generation method, a timestamp, and a version number; the triggering logic of the sample enhancement or automatic generation is: If the current weight of a category satisfies: wc(t)<γ·Tc This triggers data enhancement or new sample generation logic to supplement training samples, and γ is the weak item determination coefficient set by the system.

6. A sample distribution adaptive adjustment system based on model training feedback drive, characterized in that: Includes the following modules: Training feedback collection module: used to collect feedback indicators in real time during the model training cycle, and perform structured analysis by category or label dimension to obtain a feedback indicator set; Feedback Analysis and Decision-Making Module: This module analyzes and determines the model's learning performance in each category or task dimension based on a set of feedback indicators, identifies categories where training results do not meet the standards, and marks them as weak categories. Sampling adjustment module: used to dynamically increase the sampling weight of weak categories, reduce the sampling weight of non-weak categories according to demand or maintain the current sampling ratio, forming an updated sampling strategy; Data sampling and enhancement module: After obtaining new samples according to the sampling strategy, it is used to determine whether the number of samples in the weak category meets the preset minimum training requirements. If not, it triggers the sample enhancement or generation mechanism to supplement the samples and obtain the sample set; Training-data closed-loop linkage module: This module is used to put the optimized sample set into the next round of training, and repeat the above modules until the training indicators of the learning effect of all key categories meet the termination conditions; Obtain the sampling strategy and sample set during the cycle; if the training feedback of a certain category meets the standard for multiple consecutive rounds, or the sampling strategy benefit is lower than the preset threshold, suspend the sample replenishment and resume the previous sampling strategy.

7. The method and system for adaptively adjusting sample distribution based on model training feedback drive according to claim 6, characterized in that: The feedback indicators include but are not limited to: at least two of loss, recall, precision, F1 score, category accuracy, and confusion matrix.

8. The method and system for adaptively adjusting sample distribution based on model training feedback drive according to claim 7, characterized in that: The feedback analysis and decision module identifies the category of training effects that do not meet the standards by performing weighted fusion or rule judgment on multiple feedback indicators in the feedback indicator set, comparing them with the set threshold, and determining the category of performance that does not meet the standards.

9. The method and system for adaptively adjusting sample distribution based on model training feedback drive according to claim 8, characterized in that: The weight adjustment formula for dynamically increasing the sampling weight in the sampling adjustment module is as follows: For category c, its sampling weight is updated in round t+1 as: wc(t+1)=wc(t)+α·(Tc-Mc(t)) Among them, wc(t): is the current category sampling weight; Tc: is the category target indicator; Mc(t): is the actual value of the current training feedback; α: is the adjustable coefficient, which determines the adjustment range.

10. The method and system for adaptively adjusting sample distribution based on model training feedback drive according to claim 9, characterized in that: The method of sample supplementation in the data sampling and enhancement module specifically adopts a data enhancement algorithm or a sample generation model, and metadata information is attached to the enhanced sample or generated sample. The metadata information includes the generation method, timestamp, and version number. The trigger logic for the sample enhancement or automatic generation is: If the current weight of a category satisfies: wc(t)<γ·Tc This triggers data enhancement or new sample generation logic to supplement training samples, and γ is the weak item determination coefficient set by the system.

Citation Information

Cited By

  • Self-adaptive resampling method and device for synthetic data and real data, and electronic equipment

    CN121808748A

  • Adaptive resampling method and device for synthetic and real data, and electronic device

    CN121808748B