Model training optimization method and device based on course learning strategy and medium
By adopting multi-dimensional corpus evaluation index and dynamic weight adjustment technology in model training, the problem of not being able to fully utilize high-quality corpus under limited corpus samples is solved, and a more efficient and accurate model training effect is achieved.
Patent Information
- Application Number
- CN202510057374.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
In the model training process corresponding to the course learning strategy, application scenarios that lack massive corpus cannot fully explore high-quality corpus with limited corpus samples, and cannot fully utilize high-quality corpus during the training process, resulting in poor model training results.
Multi-dimensional corpus evaluation indicators (including loss assessment indicators, concentration scores and knowledge graph clustering indicators) are used to evaluate the corpus, and the target weight combination is determined through dynamic search, and the corpus sequence is optimized to achieve more efficient model training.
Through multi-dimensional evaluation and dynamic weight adjustment, high-quality corpus can be accurately screened under limited corpus samples, improving the efficiency and accuracy of model training, reducing resource consumption, and improving resource utilization efficiency.
Smart Images

Figure CN119988967A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model training technology, and in particular to a model training method, device and medium based on curriculum learning strategy. Background Art
[0002] With the rapid development of the field of natural language processing (NLP), technologies based on large-scale pre-trained models have achieved remarkable results in multiple application scenarios. The training effect of the model depends largely on the quality and diversity of the training corpus. In the curriculum learning strategy, presenting training samples in an orderly manner (loss sorting, attention score sorting, etc.) to improve the learning efficiency and effect of the model has been proven to be an effective training method.
[0003] Traditional model training methods usually rely on static and large-scale corpora, which makes it difficult to effectively screen out high-quality training samples, resulting in high training costs and inefficient resource utilization. There are high requirements for model accuracy and efficiency, but there is a lack of application scenarios with massive corpora. It is necessary to fully explore key samples with limited corpus samples. However, existing corpus screening methods are mostly based on a single indicator and cannot comprehensively measure the comprehensive quality of the corpus. In addition, in the process of corpus optimization in existing course learning methods, it is impossible to dynamically optimize the corpus sorting according to the real-time feedback of model training, resulting in insufficient data utilization in the training process and the inability to fully tap the potential of high-quality corpora.
[0004] Therefore, in the model training process corresponding to the course learning strategy, for application scenarios that lack massive corpus, it is impossible to fully mine high-quality corpus with limited corpus samples, and it is impossible to fully utilize high-quality corpus during the training process. The model training effect needs to be improved. Summary of the invention
[0005] One or more embodiments of the present specification provide a model training method, device and medium based on a course learning strategy, which are used to solve the following technical problems: in the model training process corresponding to the course learning strategy, for application scenarios lacking massive corpus, it is impossible to fully mine high-quality corpus with limited corpus samples, and it is impossible to fully utilize high-quality corpus during the training process, and the model training effect needs to be improved.
[0006] One or more embodiments of this specification adopt the following technical solutions:
[0007] One or more embodiments of the present specification provide a model training optimization method based on a course learning strategy, the method comprising: obtaining an initial corpus corresponding to a target model, performing a multidimensional corpus evaluation on multiple corpora in the initial corpus during each iterative training process, and determining a multidimensional corpus evaluation index corresponding to each corpus, wherein the multidimensional corpus evaluation index comprises a loss evaluation index, a concentration score, and a knowledge graph clustering index; performing a dynamic search within a pre-set weight parameter search range to determine a target weight combination corresponding to the multidimensional corpus evaluation index; sorting and optimizing the multiple corpora through the target weight combination corresponding to the multidimensional corpus evaluation index and the multidimensional corpus evaluation index corresponding to each corpus, and determining a current training corpus order, so as to perform model training according to the course learning strategy through the current training corpus order.
[0008] One or more embodiments of this specification provide a model training device based on a curriculum learning strategy, including:
[0009] at least one processor; and,
[0010] a memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0012] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.
[0013] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the technical solution provided in the embodiments of this specification, traditional model training relies on a static large corpus and it is difficult to identify high-quality samples. However, this solution breaks the limitation of single indicator screening by determining a multi-dimensional corpus evaluation indicator including a loss evaluation indicator, a focus score and a knowledge graph clustering indicator; the loss evaluation indicator accurately captures the degree to which the corpus causes model learning deviation, and can quickly locate low-quality corpora that may mislead model learning or increase training difficulty; the focus score reflects the degree of focus of the model on the corpus, highlighting the key information-bearing corpus in the model learning process; the knowledge graph clustering indicator analyzes the knowledge association between corpora based on clustering, and identifies corpora that are in the core or marginal position in the knowledge system; combining the three, accurately screen out high-quality corpora that have a great driving effect on model learning, and discard redundant and low-quality parts, so that even with limited corpus samples, the training process can be guaranteed to be efficient and accurate; through dynamic search in the preset weight parameter search interval, according to the multi-dimensional corpus evaluation feedback and model verification performance, adapt the target weight combination of the current model and corpus, and dynamically search to determine the target The target weight combination method is used to optimize and adjust in real time as the training progresses; dynamic search is performed within the pre-set weight parameter search range. Compared with the traditional manual adjustment of weights or simple fixed weight allocation methods, the most suitable target weight combination is automatically and targetedly found within the search range based on the multi-dimensional evaluation indicators of the corpus and the verification performance of the model under different weights, which greatly reduces the time cost of finding suitable weights and increases the possibility of finding the optimal weights; the corpus is sorted and optimized by the target weight combination and the multi-dimensional corpus evaluation indicators corresponding to each corpus, and then the current training corpus order is determined. The model can start learning from relatively simple, easy-to-understand corpora that are strongly related to existing knowledge. At the same time, the model can make more reasonable use of computing resources and storage resources during the training process, and use the corpus for training in an orderly manner from easy to difficult, avoiding the situation where the model repeatedly consumes resources on some more difficult corpora due to disordered or unreasonable corpus use but has poor results. The screening and sorting of the corpus also makes the storage of the corpus more targeted, reduces unnecessary storage overhead, and improves the resource utilization efficiency of the entire model training system. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:
[0015] Figure 1A flowchart of a model training method based on a course learning strategy provided in an embodiment of this specification;
[0016] Figure 2 A flowchart of another model training method based on a course learning strategy provided in an embodiment of this specification;
[0017] Figure 3 A schematic diagram of the structure of a model training device based on a course learning strategy provided in an embodiment of this specification. DETAILED DESCRIPTION
[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0019] The embodiment of this specification provides a model training method based on a course learning strategy. It should be noted that the execution subject in the embodiment of this specification can be a server or any device with data processing capabilities. Figure 1 A flow chart of a model training method based on a course learning strategy provided in an embodiment of this specification, such as Figure 1 As shown, it mainly includes the following steps:
[0020] Step S101, obtaining an initial corpus corresponding to the target model, performing multi-dimensional corpus evaluation on multiple corpora in the initial corpus during each iterative training process, and determining a multi-dimensional corpus evaluation index corresponding to each corpus.
[0021] Among them, the multi-dimensional corpus evaluation indicators include loss evaluation indicators, concentration scores, and knowledge graph clustering indicators;
[0022] In one embodiment of the present specification, when the target model needs to be trained, an initial corpus corresponding to the target model is obtained, and the initial corpus includes multiple corpora. In the discussion of the embodiments of the present specification, the target model can also be called a large model, and both represent models that need to be trained.
[0023] Existing corpus screening methods mostly use a single indicator and lack a multi-dimensional evaluation mechanism. They are unable to comprehensively measure the comprehensive quality of the corpus, which limits the further improvement of model performance. In one embodiment of the present specification, during each iterative training process, a multi-dimensional evaluation is performed on the corpus in the initial corpus, and the multi-dimensional corpus evaluation indicators include a loss evaluation indicator, a focus score, and a knowledge graph clustering indicator. Specifically, the situation of each corpus is measured from multiple different angles, such as calculating its loss evaluation indicator to determine the size of the deviation caused by the corpus when the model learns; evaluating the focus score to determine the degree of attention the model pays to the corpus; and using clustering algorithms to analyze the relationship between corpora, etc. In these ways, the corresponding multi-dimensional corpus evaluation indicators are finally determined for each corpus, so that further processing and analysis can be performed based on these indicators in the future, so as to help model training to be more efficient and higher quality.
[0024] A multi-dimensional corpus evaluation is performed on multiple corpora in the initial corpus to determine the multi-dimensional corpus evaluation indicators corresponding to each corpus, specifically including: setting the model mode to the evaluation mode, in which the forward propagation method is used to generate the predicted logits value, and based on the logits value, the predicted probability is determined; based on the predicted probability, the loss evaluation indicator and concentration score corresponding to each corpus in the current iterative training are generated; through the K nearest neighbor algorithm and the weakly connected component algorithm, the multiple corpora are clustered in the knowledge graph to determine the knowledge graph clustering indicators.
[0025] In one embodiment of the present specification, the large model is first placed in evaluation mode to ensure the accuracy and consistency of the calculation, and a forward propagation method is used to generate predicted logits, and a Softmax method is used to calculate the predicted probability, and the calculation formula is as follows:
[0026]
[0027] where p i is the predicted probability calculated by the Softmax function, N is the total number of categories, z i is the logit value of the ith category. The cross entropy loss function (CrossEntropyLoss) is used as the core indicator of loss evaluation, and its calculation formula is:
[0028]
[0029] Among them, y i is the actual label, p i It is the predicted probability calculated by the Softmax function, and N is the total number of categories. After each iteration of training, the large model obtains the Loss evaluation of each corpus under this training, which is the loss evaluation index.
[0030] Similarly, when determining the focus score, the large model is placed in evaluation mode and uses forward propagation to generate predicted logits for each corpus. The Softmax method is used to calculate the predicted probability, and the calculation formula is as follows:
[0031]
[0032] where p i is the predicted probability calculated by the Softmax function, N is the total number of categories, z i is the logit value of the ith category; the information entropy is calculated based on the probability distribution and used as the concentration score, and the formula is as follows:
[0033]
[0034] where p i It is the prediction probability calculated by the Softmax function. N is the total number of corpus categories. The higher the information entropy, the greater the model's prediction uncertainty for the word, and vice versa. At the end of each training, the big model obtains the Attention_score score of each corpus in this training, that is, the concentration score.
[0035] The multiple corpora are clustered in a knowledge graph by using the K nearest neighbor algorithm and the weakly connected component algorithm, and the knowledge graph clustering index is determined, specifically including: high-dimensional vector embedding is performed on each of the corpora in the initial corpus, and the embedding vector corresponding to each of the corpora is determined; the K nearest neighbor algorithm is used to construct a similarity graph between the multiple corpora in the initial corpus according to the embedding vector corresponding to each of the corpora; the weakly connected component algorithm is used to identify the similarity graph, and the weakly connected components in the similarity graph are determined to determine multiple corpus clusters; the number of similar corpora in each cluster of the corpus is obtained, and the knowledge graph clustering index of each corpus in the corpus cluster is determined according to the number of similar corpora corresponding to each cluster of the corpus.
[0036] Existing corpus optimization methods often lack effective clustering analysis methods when dealing with multi-dimensional evaluation indicators, and are unable to fully explore the potential connections between corpora, resulting in unsatisfactory optimization effects of the corpus structure. This not only affects the generalization ability of the model, but also increases resource consumption during training and reduces training efficiency.
[0037] In one embodiment of the present specification, the large model performs high-dimensional vector embedding on each piece of text in the corpus, which is implemented by the following formula:
[0038] e i =Embed(s i ),
[0039] Among them, e iRepresents the i-th corpus s i Embedding vector, Embed represents the embedding function of the pre-trained model, and the embedding function is used to capture the semantic relationship and similarity between corpora. According to the embedding vector corresponding to each corpus, the K nearest neighbor algorithm is used to construct a similarity graph between the corpora in the corpus. For each corpus, identify the k closest corpora in the embedding space to form an adjacency relationship. The formula is as follows: θ = (γ, ε), where γ is all corpus points in the corpus, and ε is the edge set constructed based on the K nearest neighbor algorithm. In other words, identify the K closest corpora in the embedding space for each corpus to form an adjacency relationship, so as to reflect the similarity between the corpora. Apply the weakly connected component algorithm on the constructed similarity graph, and divide multiple corpus clusters by identifying the weakly connected components in the graph:
[0040] C={C1,C2,C3,…,C m},
[0041] Among them, C represents the m clusters into which the corpus is divided, and each cluster C j Contains all corpus sets with high similarity, ensuring that the corpus similarity within the cluster is high and the differences between different clusters are large. Each corpus cluster contains all corpus sets with high similarity, and obtains the number of corpora corresponding to each corpus cluster, which can also be called the number of similar corpora. After determining the corresponding number of similar corpora, the knowledge graph clustering index of each corpus in the corpus cluster can be determined based on this number. The knowledge graph clustering index aims to measure the relevant characteristics and importance of the corpus at the knowledge level from the perspective of clustering. For example, if there are many similar corpora in the cluster where a corpus is located, this may mean that the corpus is in a relatively core position in this knowledge field (that is, the semantic category represented by the cluster), and is closely related to many other corpora, then its corresponding knowledge graph clustering index may be relatively high; conversely, if there are fewer similar corpora, it means that it is relatively marginal in this cluster and is not so closely related to other corpora, and its knowledge graph clustering index may be relatively low. In this way, each corpus in each corpus cluster is assigned a knowledge graph clustering indicator. These indicators can be used to further analyze the nature of the corpus, sort the corpus more finely, or determine the different uses and weights of the corpus based on the indicators during model training, etc., so that the entire corpus processing process can be carried out more scientifically and reasonably around the knowledge relevance of the corpus.
[0042] In addition, cluster analysis based on K nearest neighbor and WCC algorithm selects representative corpus within each cluster, and the formula is as follows:
[0043]
[0044] Among them, Representative (C j ) represents the cluster C j The representative corpus selected from the corpus. Optimize the sorting of the corpus to avoid data duplication and redundancy, ensure the diversity and coverage of the corpus, make the organization of the corpus more reasonable, and facilitate the use of subsequent models. After clustering, the clustered corpus is not used directly, but the entire corpus is optimized and sorted. The significance of this sorting is to arrange the corpus in a more reasonable order that can better reflect its internal connection and value, so that the subsequent model can better use these corpora for learning and other operations. In order to achieve optimized sorting and avoid some data problems (such as duplication, redundancy, etc.), representative corpus will be selected from each cluster that has been divided. The representative corpus can well reflect the core features and semantic information of the corpus in the cluster, which is equivalent to using them to "represent" the situation of the entire cluster. In this way, it can not only ensure that the corpus covers various semantics, features and other aspects represented by different clusters (ensuring diversity and coverage), but also remove possible duplicate and redundant data, making the quality and structure of the corpus more reasonable, and more conducive to the efficient use of these corpora by large models in the training process.
[0045] Through the above technical solutions, the loss evaluation index can identify the corpus that causes large deviations in the model; when optimizing the corpus, you can consider reducing the proportion of these high-loss corpora, or further processing them (such as cleaning, correction, etc.) to improve the overall quality of the corpus; the knowledge graph clustering index helps to determine the semantic associations and clustering structures between corpora; by selecting representative corpora in each cluster, it can ensure that the corpus covers various semantic categories and avoids duplication and redundancy of corpora. The optimized corpus is richer and more diverse in semantic content, and can better cover various knowledge areas required for model training; combined with the loss evaluation index and the focus score, the corpus can be sorted more reasonably, and the corpus with lower loss and higher model attention can be placed in front, so that the model can first contact these relatively easy-to-learn and important corpora in the early stage of training, which is in line with the learning process from easy to difficult, and helps to improve the learning efficiency of the model, so that the model can gradually use the information in the corpus; using the knowledge graph clustering index to sort the corpus, the corpus with close semantic associations can be placed in adjacent positions. In the process of model learning, the coherence and systematicity between related knowledge can be better understood, and the corpus structure can be further optimized to make the logical relationship between the corpora clearer; the multi-dimensional evaluation index enables the model to fully utilize various information of the corpus; in the learning process, the model not only considers the difficulty (loss) of the corpus and its own focus (focus), but also considers the semantic association between the corpus (knowledge graph clustering). The comprehensive learning method allows the model to better understand and master knowledge, thereby showing better performance in various tasks (such as classification, generation, etc.).
[0046] Step S102: Perform a dynamic search within a preset weight parameter search range to determine a target weight combination corresponding to the multi-dimensional corpus evaluation index.
[0047] In the existing model training method, scoring multidimensional corpus according to weights is a method that can comprehensively utilize information from all parties. However, most of the existing methods still manually adjust the weights, which is time-consuming and the results are unstable.
[0048] In one embodiment of the present specification, the target weight combination corresponding to the multidimensional corpus evaluation index is dynamically calculated and adjusted. After obtaining the multidimensional corpus evaluation index, an iterative search is performed for the multidimensional corpus evaluation index in combination with hyperparameter optimization (HPO) and K-fold cross validation. The average validation loss is determined by K-fold cross validation. Afterwards, the average loss is fed back to the TPE (Tree-structured Parzen Estimator) algorithm to update its internal probability distribution model, thereby gradually finding the target weight combination weights (α, β, and γ) that best match the multidimensional corpus evaluation index (loss evaluation index, concentration score, and knowledge graph clustering index).
[0049] Through the above technical solution, an iterative search method combining hyperparameter tuning and K-fold cross-validation is adopted to replace the manual adjustment of weights. The automated search method can systematically explore various possible weight combinations within the set range, greatly reducing the time and effort consumed by manual adjustment of weights. Since the weight search is automatically performed through an algorithm, computing resources can be better utilized. In each iteration, computing resources are used to specifically evaluate the model performance under different weight combinations, avoiding blind attempts and resource waste that may occur during manual adjustment. By reasonably arranging the interaction between K-fold cross-validation and TPE algorithms, a better weight combination can be found in a relatively short time, thereby improving the time efficiency of the entire model training process. Determine the average validation Loss can truly reflect the performance of the model on unseen data (validation set) under different weight combinations, and the adjustment of weights is based on the actual performance data of the model, rather than on subjective guesswork or empirical judgment. It can more accurately find the weight combination suitable for the model and improve the matching degree between the weight combination and the multi-dimensional corpus evaluation indicators. By dynamically calculating and adjusting the weight combinations (α, β and γ) corresponding to the multi-dimensional corpus evaluation indicators (loss evaluation indicators, focus scores and knowledge graph clustering indicators), it can make full use of the information of these different dimensions, so that the model can use the corpus for learning in a more reasonable way according to multiple factors such as the loss of the corpus, the degree of attention and the semantic clustering relationship during the training process.
[0050] A dynamic search is performed within a preset weight parameter search range to determine the target weight combination corresponding to the multi-dimensional corpus evaluation indicator, specifically including: in each iterative training process, determining the current weight candidate combination corresponding to the current iterative training; using the K-fold cross-validation method, verifying the current weight candidate combination to determine the current verification loss corresponding to the current weight candidate combination; through the hyperparameter tuning method, a dynamic search is performed within the weight parameter search range, and according to the current verification loss, the target weight combination with the minimum verification loss is determined.
[0051] The current weight candidate combination is verified by using a K-fold cross-validation method to determine the current verification loss corresponding to the current weight candidate combination, specifically comprising: dividing the initial corpus into K subsets, each subset being used as a verification set for one verification, and using K-1 subsets other than the subset as training sets; using the current weight candidate combination to perform cyclic training and verification on each of the subsets, and calculating the cross entropy loss based on the verification set to calculate the average cross entropy loss corresponding to the K verification losses, and determining the current verification loss corresponding to the current weight candidate combination.
[0052] In one embodiment of the present specification, the initial corpus is divided into K subsets. For example, if K=5, it is equivalent to evenly dividing the entire initial corpus into 5 parts, each of which is a subset. Each subset will serve as a validation set in turn, and when a certain subset is used as a validation set, the remaining K-1 subsets are combined as training sets, which can make full use of the data in the corpus, allow the model to be trained and verified under different data grouping conditions, and more comprehensively evaluate the performance of the model on different data subsets, reduce the deviation caused by a single division of the data set, and make the evaluation results more reliable and universal.
[0053] It should be noted that if the target model is the first iteration training, the current weight candidate combination is the manually set empirical weight value. If it is not the first iteration training, the current weight candidate combination is the weight parameter combination obtained by iterative search within the specified range using the TPE algorithm through hyperparameter tuning and K-fold cross validation in the previous iteration training process. For each current weight candidate combination determined (a weight combination composed of the loss ranking weight α, attention score weight β, and knowledge graph clustering weight γ corresponding to the multi-dimensional corpus evaluation index), the following operations must be performed: using the weight candidate combination, cyclic training and verification are performed under the K different training set and verification set combinations divided above. That is, for each division situation (a different subset is selected as the verification set each time), the model will be based on the corresponding training set (K-1 subsets) and the current weight candidate combination for training, and then verified on the corresponding verification set to observe the output effect of the model.
[0054] When the model is validated based on each set of training and validation sets, the cross entropy loss is calculated based on the validation set. Cross entropy loss is often used to measure the difference between the probability distribution of the classification model output and the probability distribution of the true label. The calculation formula is usually:
[0055]
[0056] Among them, N is the total number of categories (that is, the number of corpora in the validation set), y iis the actual label of the i-th sample (the true category), p i It is the probability that the model predicts that the sample belongs to each category (the probability value obtained by the model output after corresponding processing). For each K-fold cross-validation division, a cross entropy loss value can be obtained, and a total of K cross entropy loss values will be obtained. These K cross entropy loss values are averaged to obtain the average cross entropy loss. The calculation formula is:
[0057]
[0058] Loss here i It is the cross entropy loss value obtained from the i-th verification, and K is the total number of verifications (that is, the number of folds).
[0059] The average cross entropy loss value is the current validation loss corresponding to the current candidate weight combination, reflecting the overall performance of the model after the comprehensive evaluation method of K-fold cross validation under the currently selected weight combination. According to this current validation loss, the validation losses corresponding to other candidate weight combinations are compared, or fed back to the corresponding optimization algorithm to determine whether the weight combination is appropriate, and then decide whether to continue to use the weight combination or adjust the weights to find the optimal weight combination so that the model performs best on the validation set. For example, assuming there are three different candidate weight combinations A, B, and C, and they are operated according to the above K-fold cross validation process, and their corresponding average cross entropy losses (i.e., current validation losses) are calculated to be 0.3, 0.2, and 0.4 respectively, then it can be preliminarily judged that the candidate weight combination B performs relatively well in the current evaluation.
[0060] Through the hyperparameter tuning method, a dynamic search is performed within the weight parameter search range, and according to the current verification loss, a target weight combination with the minimum verification loss is determined, specifically including: according to the current verification loss and the current weight candidate combination, a weight parameter search is performed within the weight parameter search range, and iterative training verification is performed to determine the allocation verification loss corresponding to the multiple allocation weight candidate combinations searched; when the number of iterations corresponding to the iterative training verification meets the preset iteration number threshold, the minimum verification loss is determined among the multiple allocation verification losses and the current verification loss, and the target weight combination is determined with the weight candidate combination corresponding to the minimum verification loss.
[0061] In one embodiment of the present specification, through a dynamic search method, exploration is performed within a set weight parameter search range based on the current verification loss, and after multiple rounds of iterative training and verification, the target weight combination with the smallest verification loss is finally determined. The current verification loss reflects the performance of the model on the verification set under the currently used weight candidate combination. The current weight candidate combination is a parameter configuration at a starting point or intermediate state in the search process. Based on these two, the search is started within a pre-set weight parameter search range. If the current verification loss is high, it means that the current weight candidate combination may not be ideal, then it is necessary to try other weight parameter values within the search range, and explore new weight candidate combinations in the direction of reducing the verification loss.
[0062] First, we need to define the search range of the weight parameters α, β, and γ to ensure that the parameter adjustment is performed within a reasonable range. The range is defined as:
[0063] α∈[a,b],β∈[a_01,b_01],c∈[a_02,b_02],
[0064] And these ranges are limited to [0, 1]. This is because the weight represents the relative importance of each ranking result in the comprehensive evaluation, and its sum should be within a reasonable range, and each weight cannot be negative. The Bayesian optimization (Tree-Structured Parzen Estimator, TPE) algorithm in the Optuna framework is applied to predict the performance of different parameter combinations by building a probability-based model, and efficiently search for the optimal weight combination within a preset range. First, according to the distribution of the objective function, the threshold is defined, and the hyperparameter combination is divided into excellent performance (y = 1) and poor performance (y = 0) combinations. The specific formula is:
[0065] Π=quantile(f(x),τ),
[0066] Where Π represents the threshold, f(x) represents the objective function, τ represents the quantile (a fixed percentage written in advance), the part exceeding τ represents the part with excellent performance (y=1), and the part below τ represents the part with poor performance (y=0). For example, if the objective function is the validation loss, then the parameter combination corresponding to the validation loss below a certain threshold can be considered to have excellent performance.
[0067] Construct probability distribution functions for excellent and poor performance combinations respectively:
[0068] p(x|y=1)
[0069] p(x|y=0)
[0070] These two probability distribution functions are constructed based on the existing parameter combinations and their performance classifications, and are used to estimate the performance probabilities of different parameter combinations in the future. Based on the principle of Bayesian optimization, the next hyperparameter combination is selected by maximizing the expected improvement. The specific formula is:
[0071] p(x)=Πp(x|y=1)+(1-Π)p(x|y=0),
[0072]
[0073] Π represents the prior probability of the part with excellent performance. This means that when selecting the next hyperparameter combination, the probability distribution of the combination with excellent performance and the combination with poor performance will be comprehensively considered, and the parameter combination that is more likely to improve the model performance will be selected first. This process will be iterated continuously to gradually find a better parameter combination.
[0074] Every time a new set of candidate weight combinations is found, iterative training and verification must be performed. In this way, each round of iteration will obtain a verification loss corresponding to a new candidate weight combination, which is called the allocation verification loss. As the iterations continue, the allocation verification losses corresponding to multiple candidate weight combinations will be collected. Together, they constitute the data basis for the subsequent determination of the target weight combination. Each allocation verification loss represents the performance of the corresponding candidate weight combination on the verification set. The preset iteration number threshold is a termination condition for the entire search process. It is set after comprehensively considering multiple factors such as search efficiency, computing resource consumption, and the possibility of finding a better solution. For example, if the iteration number threshold is set to 5 times, the entire iterative training and verification process will continue until 5 iterations are completed, and the candidate weight combination corresponding to the minimum verification loss is determined, which is the target weight combination to be determined in the end.
[0075] Step S103, through the target weight combination corresponding to the multi-dimensional corpus evaluation index and the multi-dimensional corpus evaluation index corresponding to each corpus, the multiple corpora are sorted and optimized, and the current training corpus order is determined, so as to perform model training according to the course learning strategy through the current training corpus order.
[0076] In one embodiment of the present specification, the multi-dimensional corpus evaluation index is an index obtained by measuring the corpus from different angles, including a loss evaluation index, a concentration score, and a knowledge graph clustering index. The target weight combination is the weight value corresponding to each dimensional index determined through the hyperparameter tuning process. The role of the target weight combination is to be able to reasonably weight the corpus evaluation indicators of different dimensions according to their importance to model training, so that these indicators can be combined to work together and more accurately reflect the overall value of the corpus and its impact on model training. Based on the target weight combination and the multi-dimensional corpus evaluation index corresponding to each corpus, multiple corpora can be sorted and optimized.
[0077] It should be noted that the curriculum learning strategy is a training method that imitates the human learning process, that is, to let the model learn knowledge step by step from simple to complex. The current training corpus order determined above meets the requirements of the curriculum learning strategy. During model training, according to this optimized corpus order, the model first contacts the corpus that is in a better position after evaluation and sorting. These corpora are often relatively simple, easy to understand and highly concerned by the model, and have reasonable relevance to other corpus knowledge. The model first learns basic patterns, features, and knowledge from these corpora. As the training progresses, it gradually learns the corpora that are ranked later. These corpora may be more complex, but because the model already has the previous learning foundation, it can better grasp the content, thereby achieving a more efficient and high-quality learning process.
[0078] For example, in the text classification task of natural language processing, after the corpus is sorted and optimized, the front corpus may be some texts with clear themes, simple sentence structures, and obvious emotional tendencies. The model first learns basic vocabulary semantics, grammatical rules, and simple classification logic based on these texts; the back corpus can be texts with richer content, more complex semantics, and multiple emotions and themes. At this time, the model can try to understand and handle these more complex situations with the previous learning accumulation, thereby improving the performance in the entire text classification task. At the same time, it also helps to improve the generalization ability of the model, so that it can accurately classify new and unseen texts. Using a combination of multi-dimensional corpus evaluation indicators and target weights to optimize the corpus order and carry out model training according to the course learning strategy can give full play to the value of the corpus, improve the effect of model training and the final performance; dynamically optimize the corpus sorting according to the real-time feedback of the model training, realize the full utilization of the corpus data during the training process, and give full play to the potential of high-quality corpus.
[0079] The target weight combination corresponding to the multidimensional corpus evaluation index and the multidimensional corpus evaluation index corresponding to each corpus are used to sort and optimize the multiple corpora to determine the current training corpus order, specifically including: scoring the corpus through the target weight combination corresponding to the multidimensional corpus evaluation index and the multidimensional corpus evaluation index corresponding to each corpus to generate a current corpus score corresponding to each corpus; and sorting and optimizing the multiple corpora according to the current corpus score to determine the current training corpus order.
[0080] The corpus is scored by combining the target weights corresponding to the multidimensional corpus evaluation indicators and the multidimensional corpus evaluation indicators corresponding to each corpus, and a current corpus score corresponding to each corpus is generated, specifically including: determining the evaluation weight corresponding to each corpus evaluation indicator by combining the target weights corresponding to the multidimensional corpus evaluation indicators; taking a weighted sum of the evaluation weight corresponding to each corpus evaluation indicator and the multidimensional corpus evaluation indicators corresponding to each corpus, scoring the corpus, and generating a current corpus score corresponding to each corpus.
[0081] In one embodiment of the present specification, the multi-dimensional corpus evaluation index covers the consideration of the corpus from multiple different angles. The loss evaluation index, the focus score and the knowledge graph clustering index respectively reflect the characteristics of the corpus from the deviation brought by the corpus to the model learning, the degree of attention paid by the model to the corpus and the relevance of the corpus at the knowledge level. The target weight combination includes the loss ranking weight α, the attention score weight β and the knowledge graph clustering weight γ, which gives the corresponding importance to the evaluation indicators of each dimension, so that the information of different dimensions can be integrated in a reasonable proportion to measure the comprehensive value of the corpus.
[0082] For each corpus, the corpus score is calculated by combining the target weight and its corresponding multi-dimensional corpus evaluation index. The corpus score can be calculated using the following formula:
[0083] Final Corpus Ranking=α×Loss Ranking+β×Attention Ranking+γ×Clustering Ranking
[0084] Among them, Final Corpus Ranking represents the current corpus score of the corpus, α is the loss ranking weight, β is the attention score weight, and γ is the knowledge graph clustering weight. Loss Ranking is the loss evaluation index value, AttentionRanking is the attention score weight, and Clustering Ranking is the knowledge graph clustering weight. Through the above weighted calculation method, the index of each dimension is multiplied by the corresponding weight and then added to obtain a score value that comprehensively reflects the performance of the corpus in multiple important aspects. This score comprehensively considers multiple factors such as the difficulty of the corpus in making the model learn, the degree of attracting the model's attention, and the position in the knowledge system, and can more comprehensively measure the value of the corpus for model training. After obtaining the current corpus score corresponding to each corpus, multiple corpora are sorted and optimized according to these scores. Arrange the corpus according to the size of the corpus score. For example, a corpus with a higher score means that it has more advantages in model training after comprehensively considering the factors of various dimensions, so it will be ranked first; and a corpus with a lower score will be ranked behind.
[0085] After the sorting operation, the corpus that may have been originally disorganized or arranged according to a single dimension (for example, only according to simple standards such as corpus source and length) is sorted into a more reasonable and scientific order from the perspective of model training. This order is the current training corpus order. This makes the corpus in the corpus present a step-by-step arrangement that conforms to the model learning rules. For example, starting from corpus that is easy for the model to learn and has strong relevance, gradually transitioning to relatively complex corpus that the model can also master well based on the previous learning foundation, which helps the model to use the corpus for learning more efficiently and better realize the accumulation of knowledge and improvement of capabilities.
[0086] Through the technical solution provided by the embodiments of this specification, traditional model training relies on a static large corpus, which makes it difficult to identify high-quality samples. However, this solution breaks the limitation of single indicator screening by determining a multi-dimensional corpus evaluation indicator including a loss evaluation indicator, a focus score, and a knowledge graph clustering indicator. The loss evaluation indicator accurately captures the degree to which the corpus causes model learning deviation, and can quickly locate low-quality corpora that may mislead model learning or increase training difficulty. The focus score reflects the degree of focus of the model on the corpus, highlighting the key information-bearing corpus in the model learning process. The knowledge graph clustering indicator analyzes the knowledge association between corpora based on clustering, and identifies corpora that are in the core or marginal position in the knowledge system. Combining the three, high-quality corpora that have a great driving effect on model learning are accurately screened out, and redundant and low-quality parts are discarded, so that the training process can be efficient and accurate even with limited corpus samples. By dynamically searching in the preset weight parameter search interval, the target weight combination of the current model and corpus is adapted according to the multi-dimensional corpus evaluation feedback and model verification performance, and the target weight combination determined by dynamic search is determined in real time as the training progresses. Optimization and adjustment; dynamic search is performed within the pre-set weight parameter search range. Compared with the traditional manual adjustment of weights or simple fixed weight allocation methods, the most suitable target weight combination is automatically and targetedly found within the search range based on the multi-dimensional evaluation indicators of the corpus and the verification performance of the model under different weights, which greatly reduces the time cost of finding suitable weights and increases the possibility of finding the optimal weights; the corpus is sorted and optimized through the target weight combination and the multi-dimensional corpus evaluation indicators corresponding to each corpus, and then the current training corpus order is determined. The model can start learning from relatively simple, easy-to-understand corpora that are strongly related to existing knowledge. At the same time, the model can make more reasonable use of computing resources and storage resources during the training process, and use the corpus for training in an orderly manner from easy to difficult, avoiding the situation where the model repeatedly consumes resources on some more difficult corpora due to disordered or unreasonable corpus use but has poor results. The screening and sorting of the corpus also makes the storage of the corpus more targeted, reduces unnecessary storage overhead, and improves the resource utilization efficiency of the entire model training system.
[0087] Figure 2 Another flow chart of a model training optimization method based on a course learning strategy provided in the present specification is applied to a model training optimization system, which includes a multi-dimensional corpus evaluation module and a weight automatic calculation module. Figure 2As shown in the figure, the corpus to be trained is submitted to the multi-dimensional corpus evaluation module for analysis. For example, when faced with a corpus that contains both simple samples and complex samples, the model will perform a forward propagation operation on each corpus, calculate its loss value (Loss) and attention score (Attention Score) respectively, and then sort these results. At the same time, based on the high-dimensional vector embedding of the corpus, the system uses the K nearest neighbor and weakly connected component (WCC) algorithm to cluster the samples and divide the corpus into groups of different categories or topics.
[0088] Next, after obtaining the multi-dimensional evaluation results, the automatic weight calculation module performs an iterative search for different rankings of Loss, AttentionScore, and clustering results, combined with hyperparameter tuning and K-fold cross-validation. The system first divides the corpus into K subsets, selects 1 subset as the validation set each time, and the remaining (K-1) subsets as training sets for cyclic training, and records the average validation loss. Afterwards, the average loss will be fed back to the TPE algorithm to update its internal probability distribution model, so as to gradually find the weight parameter combination α, β, and γ that best matches the Loss, Attention, and clustering effects. It should be noted that after reaching the preset threshold of iterative training times, the training is judged to be over.
[0089] After obtaining the relatively optimal weight combination, the system generates the final corpus ranking scheme based on the weighted integration formula of α×Loss Ranking+β×AttentionRanking+γ×Clustering Ranking. At this time, the course learning strategy will give priority to high-priority samples or increase their sampling frequency, thereby achieving significant improvement in training results with less high-quality corpus.
[0090] Through the synergy of the above-mentioned multi-dimensional corpus evaluation and automatic weight adjustment, the model can fully mine key samples under limited data, improve training efficiency and generalization performance. At the same time, it also provides a flexible and feasible way for subsequent incremental training or new task adaptation, which can greatly reduce training costs and optimize resource utilization efficiency in different application scenarios (such as text classification, machine translation, dialogue systems, etc.). It is suitable for application scenarios that have high requirements for model accuracy and efficiency but lack massive corpora. The multi-dimensional evaluation method combining the K nearest neighbor and weakly connected component clustering algorithms can not only capture the potential correlation between corpora, but also effectively balance the loss ranking, attention ranking and clustering results through the automatic weight search strategy, and achieve refined control of corpus priority during course learning, thereby significantly enhancing the performance and scalability of the model.
[0091] The embodiment of this specification also provides a model training optimization device based on a course learning strategy, such as Figure 3As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0092] The embodiments of the present specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.
[0093] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0094] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0095] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0096] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0097] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0100] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0101] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0102] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0103] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0104] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.
Claims
1. A model training optimization method based on curriculum learning strategy, characterized in that: The method comprises: Obtain an initial corpus corresponding to the target model, and in each iterative training process, perform multi-dimensional corpus evaluation on multiple corpora in the initial corpus to determine the multi-dimensional corpus evaluation index corresponding to each corpus, wherein the multi-dimensional corpus evaluation index includes a loss evaluation index, a concentration score, and a knowledge graph clustering index; Perform a dynamic search within a preset weight parameter search range to determine a target weight combination corresponding to the multi-dimensional corpus evaluation index; Through the target weight combination corresponding to the multi-dimensional corpus evaluation index and the multi-dimensional corpus evaluation index corresponding to each corpus, the multiple corpora are sorted and optimized, and the current training corpus order is determined, so as to perform model training according to the course learning strategy through the current training corpus order.
2. According to claim 1, a model training optimization method based on curriculum learning strategy is characterized in that: Performing multi-dimensional corpus evaluation on multiple corpora in the initial corpus to determine the multi-dimensional corpus evaluation index corresponding to each corpus, specifically including: Set the model mode to evaluation mode, in which a forward propagation method is used to generate predicted logits values, and a predicted probability is determined based on the logits values; According to the predicted probability, a loss evaluation index and a concentration score corresponding to each of the corpora in the current iterative training are generated; The multiple corpora are clustered into knowledge graphs using the K-nearest neighbor algorithm and the weakly connected component algorithm to determine knowledge graph clustering indicators.
3. According to claim 2, a model training optimization method based on curriculum learning strategy is characterized in that: The multiple corpora are clustered into knowledge graphs by using the K nearest neighbor algorithm and the weakly connected component algorithm to determine the knowledge graph clustering indicators, which specifically include: Performing high-dimensional vector embedding on each of the corpora in the initial corpus to determine an embedding vector corresponding to each of the corpora; Using a K-nearest neighbor algorithm, constructing a similarity graph between the plurality of corpora in the initial corpus according to the embedding vector corresponding to each corpus; Using a weakly connected component algorithm, the similarity graph is identified to determine weakly connected components in the similarity graph to determine multiple corpus clusters; The number of similar corpora of each of the corpus clusters is obtained, and the knowledge graph clustering index of each corpus in the corpus cluster is determined based on the number of similar corpora corresponding to each of the corpus clusters.
4. According to the model training optimization method based on curriculum learning strategy in claim 1, it is characterized in that: Dynamically search within the preset weight parameter search range to determine the target weight combination corresponding to the multi-dimensional corpus evaluation index, specifically including: During each iterative training process, a current weight candidate combination corresponding to the current iterative training is determined; Using a K-fold cross-validation method, verify the current weight candidate combination to determine the current verification loss corresponding to the current weight candidate combination; Through the hyperparameter tuning method, a dynamic search is performed within the weight parameter search range, and the target weight combination with the minimum verification loss is determined according to the current verification loss.
5. According to claim 4, a model training optimization method based on curriculum learning strategy is characterized in that: The current weight candidate combination is verified by using a K-fold cross-validation method to determine the current verification loss corresponding to the current weight candidate combination, specifically including: The initial corpus is divided into K subsets, each subset is used as a validation set for one validation, and K-1 subsets other than the subset are used as training sets; Each of the subsets is cyclically trained and validated using the current weight candidate combination, and a cross entropy loss is calculated based on the validation set to calculate an average cross entropy loss corresponding to K validation losses, thereby determining a current validation loss corresponding to the current weight candidate combination.
6. A model training optimization method based on curriculum learning strategy according to claim 4, characterized in that: Through the hyperparameter tuning method, a dynamic search is performed within the weight parameter search range, and a target weight combination with the minimum verification loss is determined according to the current verification loss, specifically including: According to the current verification loss and the current candidate weight combination, a weight parameter search is performed within the weight parameter search range, and iterative training verification is performed to determine the allocation verification losses corresponding to the searched multiple allocation weight candidate combinations; When the number of iterations corresponding to the iterative training verification meets the preset iteration number threshold, the minimum verification loss is determined among the multiple allocated verification losses and the current verification losses, and the target weight combination is determined by the weight candidate combination corresponding to the minimum verification loss.
7. The model training optimization method based on curriculum learning strategy according to claim 1 is characterized in that: By combining the target weights corresponding to the multi-dimensional corpus evaluation indicators and the multi-dimensional corpus evaluation indicators corresponding to each corpus, the plurality of corpora are sorted and optimized to determine the current training corpus order, specifically including: The corpus is scored by combining the target weights corresponding to the multidimensional corpus evaluation indicators and the multidimensional corpus evaluation indicators corresponding to each corpus, and a current corpus score corresponding to each corpus is generated; According to the current corpus score, the plurality of corpora are sorted and optimized to determine the current training corpus order.
8. The model training optimization method based on curriculum learning strategy according to claim 7 is characterized in that: The corpus is scored by combining the target weights corresponding to the multi-dimensional corpus evaluation indicators and the multi-dimensional corpus evaluation indicators corresponding to each corpus, and a current corpus score corresponding to each corpus is generated, specifically including: Determine the evaluation weight corresponding to each corpus evaluation indicator by combining the target weights corresponding to the multi-dimensional corpus evaluation indicators; The evaluation weight corresponding to each corpus evaluation indicator and the multi-dimensional corpus evaluation indicator corresponding to each corpus are weighted and summed, and the corpus is scored to generate a current corpus score corresponding to each corpus.
9. A model training optimization device based on curriculum learning strategy, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Strategy generation model training method and device, strategy generation method and device, and terminal
CN120494031A