A data cleaning method
By calculating the loss function of the pre-trained language model to evaluate the word element value and filter the target word element, the problem of the impact of internal noise and redundant information in the sample is solved, and data quality and model performance are improved.
Patent Information
- Application Number
- CN202510376678.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the supervision and fine-tuning process of pre-trained language models, the prior art ignores the quality differences at the lexical level within the sample, resulting in data quality being affected by noise and redundant information, reducing model training efficiency and performance.
By calculating the loss function of the pre-trained language model, evaluating the word element value, filtering out the target word element in the preset interval, removing noise and redundant information, and improving data quality.
It improves data quality, shortens the convergence time of the model, improves the generalization ability and resource utilization efficiency of the model, and optimizes the training efficiency and performance of the model.
Smart Images

Figure CN119884611B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a data cleaning method. Background Art
[0002] In the field of natural language processing, the birth of large language models (LLMs) has brought about a major revolution, demonstrating powerful multi-task processing capabilities. As a key link, supervised fine-tuning (SFT) can make LLMs more suitable for specific tasks and human expectations, and is widely used in fields such as intelligent customer service, intelligent writing, and knowledge Q&A. With the development of LLMs, the importance of data in model training has become increasingly prominent. Research shows that the impact of data quality on model performance far exceeds that of data quantity.
[0003] Currently, various data cleaning and screening methods have been proposed in the SFT stage of LLM, such as screening based on perplexity, completion length, quality scores generated by the model, etc. at the sample level. However, these methods have obvious limitations. They only focus on the overall quality assessment of samples and ignore the quality differences within samples, resulting in reduced model training efficiency, learning wrong patterns or information, affecting the performance of downstream tasks, being difficult to fully utilize the value of high-quality data, and limiting the performance improvement of LLMs in SFT tasks. Summary of the Invention
[0004] This application provides a data cleaning method to at least solve the problem that in the process of supervised fine-tuning of pre-trained language models, data quality is affected by noise and redundant information at the token level within samples.
[0005] This application provides a data cleaning method, including:
[0006] Obtain the token to be processed;
[0007] Based on the loss function of the pre-trained language model, calculate the value evaluation score of the token to be processed, where the value evaluation score is used to measure the prediction accuracy of different pre-trained language model parameters for the token to be processed;
[0008] Arrange the tokens to be processed according to the value evaluation score, and screen to obtain the target tokens within a preset interval.
[0009] This application also provides a data cleaning device, including:
[0010] An obtaining module, configured to obtain the token to be processed;
[0011] A calculation module, configured to calculate a value evaluation score of a to-be-processed token based on a loss function of a pre-trained language model, where the value evaluation score is used to measure the prediction accuracy of different pre-trained language model parameters for the to-be-processed token;
[0012] A screening module, configured to rank the to-be-processed tokens according to the value evaluation scores, and screen out target tokens within a preset interval.
[0013] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above data cleaning methods when executing the computer program.
[0014] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above data cleaning methods are implemented.
[0015] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above data cleaning methods are implemented.
[0016] Through this application, a value evaluation score of a to-be-processed token is calculated based on a loss function of a pre-trained language model, and this score measures the prediction accuracy of different pre-trained language model parameters for this token. During the supervised fine-tuning process of a pre-trained language model, noise and redundant information at the token level within the sample often cause the model to have biases in predicting these worthless or low-value tokens. By calculating the value evaluation score, it is possible to identify which tokens have less impact on the model's prediction accuracy, and these tokens are likely to be the sources of noise or redundant information.
[0017] This application ranks the to-be-processed tokens according to the value evaluation scores, and screens out target tokens within a preset interval. This operation can exclude tokens with lower value evaluation scores (i.e., tokens that may contain noise and redundant information). Through this screening method, these tokens that have a negative impact on the model fine-tuning can be effectively removed, thereby improving the data quality and reducing their interference with the supervised fine-tuning of the pre-trained language model.
[0018] Therefore, this application can solve the technical problem that the data quality is affected by noise and redundant information at the token level within the sample during the supervised fine-tuning process of a pre-trained language model, thereby improving the data quality, accelerating the convergence of the model, enhancing the generalization ability of the model, and optimizing the resource utilization. Description of the Drawings
[0019] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 Schematic diagram of the application environment architecture for running a data cleaning method provided by an embodiment of the present application;
[0021] Figure 2 Schematic diagram of the hardware architecture for running a data cleaning method provided by an embodiment of the present application;
[0022] Figure 3 Schematic diagram of the process of a data cleaning method provided by an embodiment of the present application;
[0023] Figure 4 Schematic diagram of the process of iteratively training a pre-trained language model provided by an embodiment of the present application;
[0024] Figure 5 Schematic diagram of the structure of a data cleaning device provided by an embodiment of the present application;
[0025] Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0027] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0028] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further details the present application with reference to the drawings and specific implementation manners.
[0029] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the data cleaning method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0030] As Figure 1 shown, Figure 1 FIG. 1 is a schematic diagram of an application environment architecture in which a data cleaning method provided by an embodiment of the present application runs. The application environment architecture includes a data storage layer, a computing framework layer, a task scheduling and management system, etc.
[0031] Among them, the data storage layer needs to have high-capacity and high-reliability storage devices, such as a distributed file system, to store a large amount of corpus data to be cleaned, so as to meet the requirements of the cleaning method for large-scale data processing.
[0032] The computing framework layer takes a deep learning framework as the core and provides computing support for operations such as calculating the token value evaluation score based on a loss function. These frameworks can efficiently calculate the loss function of the pre-trained language model, thereby obtaining the token value evaluation score.
[0033] The task scheduling and management system is responsible for scheduling cleaning tasks, reasonably allocating computing resources, and managing the task execution of each computing node, so as to ensure the efficient and stable progress of the cleaning work.
[0034] As Figure 2 shown, Figure 2 FIG. 2 is a schematic diagram of a hardware architecture in which a data cleaning method provided by an embodiment of the present application runs. The hardware architecture includes a high-performance central processing unit (CPU) cluster, a graphics processing unit (GPU) acceleration card, a high-speed network device, etc.
[0035] Among them, the high-performance CPU cluster undertakes basic computing tasks such as model initialization, data preprocessing, and task coordination. Its multi-core and high-frequency characteristics help to quickly process data reading, simple calculations, and logical control, etc.
[0036] The GPU acceleration card is the core computing hardware, especially for the complex matrix operations of the pre-trained language model. It accelerates the model training and the calculation process of the token value evaluation score, and greatly improves the cleaning efficiency.
[0037] The high-speed network device needs to be equipped with a 10 Gigabit Ethernet or even an InfiniBand network to achieve fast data transmission between the CPU and the GPU and between each computing node. When cleaning a large-scale corpus, data is frequently exchanged between different devices and nodes. The high-speed network can reduce data transmission latency, ensure the full utilization of computing resources, and avoid data transmission from becoming a performance bottleneck.
[0038] Supervised Fine-Tuning (SFT) is a process of further training on a pre-trained model using labeled data. By fine-tuning the pre-trained model on a labeled dataset for a specific task, the model can better adapt to the requirements of the specific task, thereby improving its performance on that task. During the supervised fine-tuning stage of pre-trained language models, the noise and redundant information at the token level within samples seriously interfere with data quality. Noise includes mislabeling, information unrelated to the task, etc., which cause the model to receive incorrect signals during learning and confuse the acquisition of key knowledge. Redundant information such as common words and phrases that appear repeatedly but are not very helpful for the task increases the training load of the model and reduces training efficiency. The interweaving of these negative factors makes it difficult to guarantee data quality, hinders the pre-trained language model from fully leveraging the advantages of supervised fine-tuning, and urgently requires effective means to solve.
[0039] To fully or partially solve the problems existing in related technologies, embodiments of the present application provide a data cleaning method, and the method will be described in detail in combination with the execution process of the data cleaning method.
[0040] As Figure 3 shown, the data cleaning method mainly includes the following steps S300~S302:
[0041] S300. Obtain the token to be processed.
[0042] In some embodiments, the present application is applicable to multi-modal data cleaning. When performing step S300, first obtain multi-modal data, and then segment and / or encode the multi-modal data to obtain the token to be processed. Thus, the corpora of different modalities are uniformly represented as a continuous token sequence. Among them, multi-modal data includes but is not limited to text, images, and videos. For text data in the present application, sub-word segmentation methods such as WordPiece or Byte-Pair Encoding (BPE) can be used to encode and obtain the token to be processed, so as to control the vocabulary size while retaining semantics. For image data, it can be divided into image patches (patches) and encoded to obtain the token to be processed, which can capture local features. For video data, it can be decomposed into a frame sequence and spatio-temporal encoded to obtain the token to be processed, taking into account both inter-frame relationships and spatial information.
[0043] The above embodiments convert multi-modal data into a unified token sequence, enabling the pre-trained language model to process the corpus in a unified manner, facilitating the model to learn different modal information, and improving the compatibility of the pre-trained language model. Different segmentation and encoding means are used for different modal data to form effective representations, helping the pre-trained language model understand and associate different modal semantics, thereby improving the model performance.
[0044] S301. Calculate the value evaluation score of the token to be processed based on the loss function of the pre-trained language model.
[0045] Among them, the value evaluation score is used to measure the prediction accuracy of different pre-trained language model parameters for the token to be processed. Different pre-trained language model parameters can be the model parameters of different pre-trained language models, or different model parameters of the same pre-trained language model, which can be understood as the model parameters at different stages during the training of the pre-trained language model. The higher the prediction accuracy of the token to be processed, the smaller the value evaluation score.
[0046] S302. Arrange the tokens to be processed according to the value evaluation score, and filter out the target tokens within the preset interval.
[0047] The above steps S300 - S302 can effectively remove the noise and redundant tokens in the corpus by calculating the value evaluation score of the tokens, improve the quality of the corpus data, and thus improve the training efficiency and performance of the pre-trained language model.
[0048] In some embodiments, the pre-trained language models can be different pre-trained language models, including the first model and the second model, and the performance of the second model is better than that of the first model. Correspondingly, the different pre-trained language model parameters include the first model parameters and the second model parameters. It should be noted that the number of different pre-trained language models in this application is not limited, and here two are taken as examples for illustration.
[0049] When executing step S301, first calculate the first loss function of the token to be processed under the preset context window and the first model parameters, and the second loss function of the token to be processed under the preset context window and the second model parameters. Then, calculate the value evaluation score based on the difference between the second loss function and the first loss function.
[0050] Optionally, in the process of calculating the value evaluation score based on the difference between the second loss function and the first loss function, the value evaluation score Score of the token to be processed can be calculated according to the following formula (1):
[0051] (1)
[0052] In formula (1), x i,j represents the token to be processed, x i,0:j-1 represents the preset context window where the token to be processed is located. θ and θ' represent different pre-trained language model parameters, such as the first model parameter θ and the second model parameter θ'. l ( xi,j | x i,0:j-1 ; θ ) is the loss function of the token to be processed in the preset context window x i,0:j-1 and the model parameters θ, such as the first loss function, ; l ( x i,j | x i,0:j-1 ; θ )' is the loss function of the token to be processed in the preset context window x i,0:j-1 and the model parameters θ', such as the second loss function. The change in the loss function can measure the information gain or loss of the pre-trained language model for the token to be processed under different model parameters θ and θ', reflecting the data quality related to the token to be processed. The pre-trained language model x i,0:j-1 in the given context window x i,j has a higher prediction accuracy for the token to be processed, the smaller the value of Score.
[0053] In the above embodiments, the value evaluation score of the token to be processed is calculated through the model parameters of different pre-trained language models. Due to the differences in architecture, training data, training objectives, etc. of different pre-trained language models, the understanding and focus on the same token are different, and the value of the token can be considered from multiple perspectives, avoiding the limitations of a single model, so as to more accurately judge the importance of the token for model training. By combining the judgments of different pre-trained language models, various noise tokens in the corpus can be more comprehensively identified. Different pre-trained language models can judge whether a token is redundant from different perspectives, so as to more accurately screen out redundant tokens and improve the quality and information density of the corpus. In addition, different pre-trained language models learn different language knowledge and patterns during the training process. By using the parameters of the two models for token cleaning, these diverse knowledge can be integrated into the corpus, enabling the model to access richer language information during the training process, thereby enhancing the generalization ability of the model.
[0054] On the basis of the above embodiments, in the process of calculating the value evaluation score based on the difference between the second loss function and the first loss function, first obtain the occurrence frequency of the token to be processed in the corpus, and then calculate the value evaluation score based on the occurrence frequency, the difference between the first loss function and the second loss function.
[0055] Specifically, the value evaluation score Score is calculated as shown in formula (2).
[0056] (2)
[0057] In formula (2), in addition to the difference between the first loss function and the second loss function, it includes a penalty term regarding the frequency of occurrence of the to-be-processed token in the corpus. The hyperparameter α is used to control the weight of the penalty term. N x,i,j is the number of times the to-be-processed token appears in the corpus. N total is the total number of tokens in the corpus. This penalty term indicates that the higher the frequency of occurrence of the to-be-processed token in the corpus, the higher the possibility of being cleaned. This enables the pre-trained language model to pay attention to those relatively uncommon but more critical tokens when processing the corpus.
[0058] In the above embodiments, by using the frequency-of-occurrence penalty term, the scores of high-frequency tokens are suppressed, increasing the possibility of being discarded, avoiding the model's over-reliance on high-frequency common tokens, and prompting it to learn more diverse and discriminative token information. When facing new data, the model can better adapt to different contexts and semantics, reduce the problem of insufficient generalization caused by the dominance of high-frequency words, and improve its performance and generalization ability in different scenarios.
[0059] The value evaluation parameter can also be calculated based on accuracy, log-likelihood, entropy, etc. The above is only an exemplary illustration, and the present application does not specifically limit the calculation method of the value evaluation parameter.
[0060] In some embodiments, when performing step S302, first, the to-be-processed tokens are clustered according to the part of speech of the to-be-processed tokens to obtain at least one part-of-speech clustering result; then, the to-be-processed tokens in at least one part-of-speech clustering result are arranged according to the value evaluation score, and the target tokens in the preset interval are screened out according to a preset deletable threshold.
[0061] Among them, the preset deletable threshold represents the maximum number of tokens that can be deleted for each part of speech. It is used to prevent excessive deletion of tokens of a certain part of speech at one time to avoid the situation of semantic discontinuity and semantic loss.
[0062] Specifically, the parts of speech of the tokens to be processed include, but are not limited to: nouns, pronouns, adjectives, adverbs, verbs, numerals, articles, prepositions, conjunctions, interjections, participles, and infinitives. Cluster the tokens to be processed according to these part-of-speech types to obtain clustering results for different parts of speech. Then, for each part-of-speech clustering result, sort the tokens to be processed in descending order according to their value evaluation scores, and then clean the tokens according to a preset deletable threshold and a preset interval to obtain the target tokens. Compare whether the number of other tokens in the preset interval in each part-of-speech clustering result is greater than the preset deletable threshold. If so, it means that if these other tokens are deleted at once, semantic problems will occur. Therefore, in this application, the sorted part-of-speech clustering results are cleaned according to the preset deletable threshold, and the number of remaining tokens is greater than the size of the preset interval. If the number of other tokens is less than or equal to the preset deletable threshold, it means that if these other tokens are deleted at once, no semantic problems will occur. Then, first delete these other tokens from the sorted part-of-speech clustering result to obtain the target tokens in the preset interval.
[0063] Through part-of-speech clustering, the above embodiments enable the large model to clearly distinguish vocabulary with different grammatical functions, thereby improving the accuracy of text semantic understanding, limiting the maximum deletable quantity, retaining tokens with high value, facilitating the model to focus on semantic information, avoiding learning too much redundant or low-value information, and reducing the risk of overfitting. This screening mechanism enables the model to learn more representative features and, when facing new data, can better generalize and improve the performance of the model in various scenarios.
[0064] In some embodiments, when performing step S302, first, sort the tokens to be processed according to their value evaluation scores to obtain a first ascending result. Then, according to the total number of tokens to be processed and a preset proportionality coefficient, calculate a preset interval, and further set target labels for the tokens to be processed in the preset interval in the first sorting result to screen out the target tokens. The target label indicates that the target token is "valuable", indicating that the target label is retained.
[0065] Specifically, sort the value evaluation scores of all tokens to be processed in descending order to obtain a sorting result. Then, taking the total number of tokens to be processed as the base number, combined with a preset proportionality coefficient (such as k , which can be a fixed constant), determine the range boundaries of the preset interval by multiplying the total number of tokens to be processed by the preset proportionality coefficient. Further, set target labels for the tokens to be processed ranked in the top k % of the first sorting result to set the "valuable" label and retain these tokens to be processed as target tokens. Otherwise, set the "worthless" label for the tokens to be processed and remove these tokens to be processed.
[0066] In the above embodiments, by setting a proportionality coefficient, lemmas with higher scores and greater values are selected, a large number of low-value lemmas are removed, and the model training efficiency is improved. At the same time, the model is prevented from being interfered by a large number of meaningless or low-value lemmas, and the pertinence and effectiveness of data processing are enhanced.
[0067] In the process of training the same pre-trained language model, adjusting the model parameters is to enable the pre-trained language model to better fit the training data, thereby improving the prediction ability for new data. In some embodiments, for the training of the same pre-trained language model, when performing step S300, first divide the data sample set of the pre-trained language model into multiple data sample subsets, and then use any one of the data sample subsets to fine-tune the basic model parameters of the pre-trained language model to obtain the initial model parameters. Any one of the data sample subsets is randomly selected from multiple data sample subsets. Then, obtain the to-be-processed lemmas from other data sample subsets, where the other data sample subsets are the data sample subsets other than the aforementioned any one of the data sample subsets among the multiple data sample subsets. Then, for these to-be-processed lemmas obtained from other data sample subsets, perform the aforementioned steps S301 - S302.
[0068] In the process of performing steps S301 - S302 on the to-be-processed lemmas obtained from other data sample subsets, it includes first calculating the loss function of the to-be-processed lemmas in other data sample subsets under their context window and basic model parameters, and calculating the loss function of the to-be-processed lemmas in other data sample subsets under their context window and initial model parameters. Then, based on the difference between these two loss functions, obtain the value evaluation score of the to-be-processed lemmas in other data sample subsets, referring to formula (1) or (2).
[0069] It should be noted that when the model is updated from the basic model parameters θ to the initial model parameters θ', this change in parameters will cause the prediction ability of the model for each lemma to change. At this time, l ( x i,j | x i,0:j-1 ; θ ) and l ( x i,j | x i,0:j-1 ; θ ') respectively represent under the basic model parameters θ and the initial model parameters θ', x i,j this lemma in the given context window x i,0:j-1The negative value of the predicted probability. When the model is updated from parameter θ to θ', the predicted probability distribution of the pre-trained language model for the token to be processed will change, thus affecting the prediction accuracy. A negative value of the value evaluation score indicates that the pre-trained language model is more accurate in predicting the token to be processed after the model parameter update; a positive value of the value evaluation score indicates that the prediction accuracy of the pre-trained language model for the token to be processed decreases after the model parameter update. For example, if the pre-trained language model is more accurate in predicting a certain token to be processed after the model parameter update, it means that the information carried by the token to be processed is more valuable for the training of the pre-trained language model, and the data quality related to the token to be processed is relatively high; conversely, if the prediction accuracy of the pre-trained language model for a certain token to be processed decreases after the model parameter update, it means that there may be problems with the data related to the token to be processed.
[0070] Further, sort the tokens to be processed in other subsets of data samples according to these value evaluation scores, and screen out the tokens ranked in the top k % as the target tokens of other subsets of data samples. Here, the tokens to be processed in other subsets of data samples can be clustered by part of speech first to obtain at least one part-of-speech clustering result, and then the tokens to be processed in at least one part-of-speech clustering result are sorted in ascending order according to the value evaluation scores, and then screened according to a preset deletable threshold to obtain the tokens ranked in the top k % as the target tokens of other subsets of data samples.
[0071] After that, use the target tokens of other subsets of data samples to fine-tune the initial model parameters to obtain the reference model parameters. Then calculate the third loss function of the target tokens under the preset context window and the basic model parameters, and the fourth loss function of the target tokens under the preset context window and the reference model parameters. Then, based on the difference between the fourth loss function and the third loss function, obtain the value evaluation score, as shown in formula (1) or (2). Arrange the target tokens according to these value evaluation scores, and screen out the tokens in the preset interval for iterative training of the pre-trained language model until the model converges.
[0072] It should be noted that in the above embodiments, the preset proportional coefficient can be a fixed constant. In different data sets or training tasks, as long as k remains unchanged, the tokens can be screened according to the same rule to ensure the consistency of data processing. This helps to improve the stability in the model training process, avoid training fluctuations caused by changes in the token screening rules, and enable the model to learn data features more robustly.
[0073] Optionally, the preset proportionality coefficient can vary with model iteration. Thus, in the above embodiments, in the process of arranging the target tokens according to these value evaluation scores, screening out the tokens within the preset interval for iterative training of the pre-trained language model until the model converges, the following steps may be included: First, arrange the target tokens according to the value evaluation scores to obtain a second sorting result, and then calculate the preset interval according to the total number of target tokens and the preset proportionality coefficient, where the preset proportionality coefficient decreases with the increase of the iteration step during the iterative training of the pre-trained language model. Further, screen out the tokens within the preset interval from the second sorting result for subsequent iterative training of the pre-trained language model until a converged pre-trained language model is obtained.
[0074] Specifically, in the initial stage of iterative training, in the process of performing steps S301 - S302 on other subsets of data samples, it includes first calculating the loss function of the tokens to be processed in other subsets of data samples under their context windows and the basic model parameters, and calculating the loss function of the tokens to be processed in other subsets of data samples under their context windows and the initial model parameters. Then calculate the difference between these two loss functions to obtain the value evaluation scores of the tokens to be processed in other subsets of data samples, and further sort the tokens to be processed in other subsets of data samples according to these value evaluation scores. Calculate the preset interval according to the total number of tokens to be processed and a relatively large preset proportionality coefficient to screen out the tokens to be processed within this preset interval as target tokens from the first sorting result. In the second iteration, fine-tune the initial model parameters using the target tokens in other subsets of data samples to obtain reference model parameters. Then calculate the third loss function of the target tokens under the preset context window and the basic model parameters, and the fourth loss function of the target tokens under the preset context window and the reference model parameters. Based on the difference between the fourth loss function and the third loss function, obtain the value evaluation scores. Arrange the target tokens according to these value evaluation scores to obtain a second sorting result, and then calculate another preset interval according to the total number of target tokens and a new preset proportionality coefficient smaller than the preset proportionality coefficient in the previous iteration to screen out the target tokens within this new preset interval as samples for the next iterative training. The subsequent iterative training process is similar to the foregoing, where the preset proportionality coefficient gradually decreases, ensuring the cleaning degree of the tokens and the model training speed until a converged pre-trained language model is obtained.
[0075] Exemplarily, as Figure 4 shown, divide the data sample set D of the pre-trained language model into multiple subsets of data samples { D 0, D 1,..., D N}, and then use the data sample subset D0 to fine-tune the base model parameters θ of the pre-trained language model to obtain the initial model parameters θ0. Then calculate the value evaluation scores of the to-be-processed tokens of other data sample subsets { D 1,..., D N}. Use the value evaluation scores to clean the worthless tokens to obtain the target tokens of the to-be-processed tokens. Then use the target tokens of the to-be-processed tokens to fine-tune the initial model parameters θ0 to obtain the reference model parameters θ1, and then perform the second iteration. During the second iteration, calculate the value evaluation scores of the target tokens based on the reference model parameters θ1 to further clean the low-value tokens, and use the cleaned and fine-tuned reference model parameters θ1 to obtain the new model coefficients θ 2……n for subsequent iterations. During the subsequent iteration process, based on the new model coefficients θ 2……n calculate the value evaluation parameters and clean the tokens until n = N, and the iteration ends to obtain the cleaned data sample set and the pre-trained language model with converged parameter models.
[0076] In the above embodiments, at the initial stage of training, a relatively large k allows more tokens to be retained, avoiding the loss of potential useful information due to premature screening, providing rich and diverse information for the model, enabling the model to quickly learn the basic patterns and extensive features in the corpus, and constructing a preliminary knowledge system. As the iteration progresses, k gradually shrinks, and the model begins to focus on a small number of more valuable tokens, deeply learning complex and key semantic information and detailed features, enabling the model to concentrate on processing core information, improving the data utilization efficiency, and at the same time reducing the computational cost and training time. This progressive learning from broad to fine improves the model's understanding depth and accuracy of the corpus.
[0077] Based on the above embodiments, the present application provides an alternative implementation manner. The preset proportionality coefficient decays cosine-like as the number of iteration steps increases during the iterative training of the pre-trained language model. Correspondingly, after arranging the target tokens according to the value evaluation scores to obtain the second sorting result, and before calculating the preset interval according to the total number of the target tokens and the preset proportionality coefficient, it further includes: first determining the current iteration step and the total number of iteration steps of the pre-trained language model, and then calculating the preset proportionality coefficient according to the first proportionality coefficient boundary value, the second proportionality coefficient boundary value, and the current iteration step and the total number of iteration steps. The first proportionality coefficient boundary value and the second proportionality coefficient boundary value are preset. The first proportionality coefficient is greater than the second proportionality coefficient boundary value, and the first proportionality coefficient boundary value can be expressed as k max , and the second proportionality coefficient boundary value can be expressed as k min . So that the preset proportionality coefficient in the iterative process of the pre-trained language model starts from the first proportionality coefficient boundary valuek max Gradually decrease to the second proportionality coefficient boundary value k min .
[0078] Specifically, according to the first proportionality coefficient boundary value, the second proportionality coefficient boundary value, the current iteration step number, and the total number of iteration steps, calculate the preset proportionality coefficient as shown in formula (3):
[0079] (3)
[0080] In formula (3), k max is the first proportionality coefficient boundary value, k min is the second proportionality coefficient boundary value, n is the current iteration step number, N is the total number of iteration steps. k Decays cosine-like as the iteration steps progress.
[0081] The above embodiments utilize the smooth change characteristics of cosine decay to avoid k the impact of sudden changes in the value on the model, which helps the model to be stably trained and enables it to better transfer and apply the learned knowledge when facing new data.
[0082] In some embodiments, the present application establishes a feedback mechanism, which specifically includes using the converged pre-trained language model to perform downstream tasks and determining the accuracy of the converged pre-trained language model in the downstream tasks; then, adjusting the aforementioned first proportionality coefficient boundary value and the second proportionality coefficient boundary value according to the accuracy to achieve feedback on the iterative training of the large language model. The direction and amplitude of the adjustment of the first proportionality coefficient boundary value and the second proportionality coefficient boundary value are related to the type of the downstream task.
[0083] Among them, the downstream task can be a text classification task, an image classification task, a mathematical logic analysis task, etc., and the present application does not specifically limit this. Different downstream tasks have different characteristics and requirements. Dynamically adjusting the first proportionality coefficient boundary value and the second proportionality coefficient boundary value according to the task performance can enable the model to find the most suitable token screening strategy in different tasks.
[0084] Exemplarily, for text classification tasks and image classification tasks, considering that the tokens to be processed contain more noise and the influence of the tokens to be processed on the task results is relatively uniform, if the accuracy of the pre-trained language model in the task decreases after training converges based on the cleaned tokens, then increase the first proportionality coefficient boundary value k max and the second proportionality coefficient boundary value k minFor mathematical logic analysis tasks, the missing of a few key words will lead to a sudden change in the reasoning results. If the accuracy of the pre-trained language model in the task decreases after the training converges based on the cleaned word units, the first proportional coefficient boundary value is reduced. k max and the second proportional coefficient boundary value k min .
[0085] The above embodiment can directly understand the quality of the model under the first and second scale factor boundary values by evaluating the performance of the model on the downstream tasks. If the task performance is poor, you can make targeted adjustments. k , in order to change the way the model selects words, so that the model can focus on information that is more conducive to completing downstream tasks, improve the model's key indicators such as accuracy and recall rate on the task, and optimize the overall performance of the model. The model can flexibly adapt to the needs of different tasks by adjusting the first and second scale coefficient boundary values, maintain good performance, and help improve the practicality and scalability of the model.
[0086] In summary, the present application provides a data cleaning method, which calculates the value assessment score of the word to be processed based on the loss function of the pre-trained language model. This score measures the prediction accuracy of the word by different pre-trained language model parameters. During the supervised fine-tuning process of the pre-trained language model, the noise and redundant information at the word level in the sample often cause the model to deviate from the prediction of these worthless or low-value words, increasing the value of the loss function. By calculating the value assessment score, it is possible to identify which words have less impact on the model's prediction accuracy, and these words are likely to be the source of noise or redundant information. For example, some high-frequency function words that do not contribute much to semantics (such as "的", "了", etc.) may get lower scores when calculating the value assessment score due to their limited effect on improving the model's prediction accuracy.
[0087] In addition, the present application arranges the tokens to be processed according to the value assessment scores, and screens the target tokens in the preset range. This operation can exclude tokens with low value assessment scores (i.e., tokens that may contain noise and redundant information). Because the tokens corresponding to noise and redundant information often have low value assessment scores, this screening method can effectively remove these tokens that have a negative impact on model fine-tuning, thereby improving data quality and reducing their interference with the supervised fine-tuning of the pre-trained language model.
[0088] Compared with the related art, this application produces the following technical effects:
[0089] (1) By removing noise and redundant information, the data used for supervised fine-tuning of the pre-trained language model becomes cleaner and more valuable. High-quality data enables the model to learn more accurate and useful knowledge, thereby improving the performance and performance of the model. For example, in text classification tasks, after removing noise and redundant information, the model can focus more on the key semantic information in the text and improve the accuracy of classification.
[0090] (2) Due to the improved quality of the cleaned data, the model can converge to the optimal solution faster during the fine-tuning process. Reducing the interference of noise and redundant information, the model can more effectively learn the patterns and rules in the data, avoid making too many adjustments in the wrong direction, thereby shortening the training time and improving the training efficiency.
[0091] (3) Using the cleaned data for fine-tuning, the model can learn more representative and generalizable features, rather than overfitting the noise and redundant information in the training data. This helps the model to better predict and infer when facing new and unseen data, improving the generalization ability of the model and enabling it to have a more stable performance in different scenarios and tasks.
[0092] (4) After removing noise and redundant information, the amount of data is relatively reduced, which means that the computing resources (such as memory, CPU, GPU, etc.) required during model training will also be reduced accordingly. It is possible to process more tasks under the same hardware conditions, or use smaller computing devices to complete the training, thereby reducing the computing cost and improving the resource utilization efficiency.
[0093] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner.
[0094] The embodiments of the present application also provide a data cleaning device, as Figure 5 shown, the device includes:
[0095] An acquisition module 500, configured to acquire to-be-processed tokens;
[0096] A calculation module 501, configured to calculate a value evaluation score of the to-be-processed tokens based on the loss function of the pre-trained language model, where the value evaluation score is used to measure the prediction accuracy of different pre-trained language model parameters for the to-be-processed tokens;
[0097] A screening module 502, configured to arrange the to-be-processed tokens according to the value evaluation score, and screen out target tokens within a preset interval.
[0098] As an alternative implementation provided in the embodiments of the present application, the acquisition module 500 is specifically configured to: acquire multimodal data; segment and / or encode the multimodal data to obtain to-be-processed tokens.
[0099] As an alternative implementation provided in the embodiments of the present application, the pre-trained language model includes a first model and a second model, and the performance of the second model is better than that of the first model;
[0100] The calculation module 501 is specifically configured to: calculate a first loss function of the to-be-processed tokens under a preset context window and first model parameters, and a second loss function of the to-be-processed tokens under the preset context window and second model parameters; obtain a value evaluation score based on the difference between the second loss function and the first loss function.
[0101] As an alternative implementation provided in the embodiments of the present application, the calculation module 501 is specifically configured to: obtain the occurrence frequency of the to-be-processed tokens in the corpus; calculate a value evaluation score according to the difference between the second loss function and the first loss function, and the occurrence frequency.
[0102] As an alternative implementation provided in the embodiments of the present application, the screening module 502 is specifically configured to: cluster the to-be-processed tokens according to the part of speech of the to-be-processed tokens to obtain at least one part-of-speech clustering result; arrange the to-be-processed tokens in at least one part-of-speech clustering result according to the value evaluation score, and screen out target tokens within a preset interval according to a preset deletable threshold.
[0103] As an alternative implementation provided in the embodiments of the present application, the screening module 502 is specifically configured to: arrange the to-be-processed tokens according to the value evaluation score to obtain a first sorting result; calculate a preset interval according to the total number of the to-be-processed tokens and a preset proportionality coefficient; use the to-be-processed tokens within the preset interval in the first sorting result as target tokens.
[0104] As an alternative implementation provided in the embodiments of the present application, the acquisition module 500 is specifically configured to: divide the data sample set of the pre-trained language model into multiple data sample subsets; fine-tune the basic model parameters of the pre-trained language model by using any one of the data sample subsets to obtain initial model parameters; acquire to-be-processed tokens from other data sample subsets, where the other data sample subsets are data sample subsets other than any one of the multiple data sample subsets.
[0105] As an alternative implementation provided in the embodiments of the present application, the apparatus further includes an iteration module, configured to: fine-tune the initial model parameters by using the target tokens corresponding to other data sample subsets to obtain reference model parameters;
[0106] Calculate the third loss function of the target token under the preset context window and the basic model parameters, and the fourth loss function of the target token under the preset context window and the reference model parameters; calculate the value evaluation score based on the difference between the fourth loss function and the third loss function; arrange the target tokens according to the value evaluation score, and filter out the tokens within the preset interval for the iterative training of the pre-trained language model until a converged pre-trained language model is obtained.
[0107] As an alternative implementation provided in the embodiments of the present application, the iterative module is specifically configured to: arrange the target tokens according to the value evaluation score to obtain a second sorting result; calculate the preset interval according to the total number of target tokens and the preset proportionality coefficient; filter out the tokens within the preset interval from the second sorting result for the iterative training of the pre-trained language model until a converged pre-trained language model is obtained; wherein, the preset proportionality coefficient gradually decreases as the number of iterative steps increases during the iterative training of the pre-trained language model.
[0108] As an alternative implementation provided in the embodiments of the present application, the iterative module is further configured to: determine the current iterative step and the total number of iterative steps of the pre-trained language model; calculate the preset proportionality coefficient according to the first proportionality coefficient boundary value, the second proportionality coefficient boundary value, the current iterative step and the total number of iterative steps, and the first proportionality coefficient boundary value is greater than the second proportionality coefficient boundary value, so that the preset proportionality coefficient gradually decreases from the first proportionality coefficient boundary value to the second proportionality coefficient boundary value during the iterative training of the pre-trained language model.
[0109] As an alternative implementation provided in the embodiments of the present application, the apparatus further includes a feedback module, configured to: execute a downstream task using the converged pre-trained language model, and determine the accuracy of the converged pre-trained language model in the downstream task; adjust the first proportionality coefficient boundary value and the second proportionality coefficient boundary value according to the accuracy to perform feedback iteration on the pre-trained language model.
[0110] For the description of the features in the corresponding embodiments of the data cleaning device, reference may be made to the relevant descriptions in the corresponding embodiments of the data cleaning method, which will not be elaborated here one by one.
[0111] Embodiments of the present application also provide an electronic device, as Figure 6 shown, including a memory 601 and a processor 602. A computer program is stored in the memory 601, and the processor 602 is configured to run the computer program to execute the steps in any of the above embodiments of the data cleaning method.
[0112] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described embodiments of the data cleaning method when running.
[0113] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media capable of storing computer programs such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs.
[0114] An embodiment of the present application also provides a computer program product, where the computer program product includes a computer program, and the computer program implements the steps in any of the above-described embodiments of the data cleaning method when executed by a processor.
[0115] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the computer program implements the steps in any of the above-described embodiments of the data cleaning method when executed by a processor.
[0116] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0117] The above has introduced in detail a data cleaning method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data cleaning method, characterized in that, including: obtaining a token to be processed; calculating a value evaluation score of the token to be processed based on a loss function of a pre-trained language model, where the value evaluation score is used to measure the prediction accuracy of different pre-trained language model parameters for the token to be processed; arranging the tokens to be processed according to the value evaluation score, and screening out target tokens within a preset interval; The arranging the tokens to be processed according to the value evaluation score and screening out target tokens within a preset interval includes: arranging the tokens to be processed according to the value evaluation score of the tokens to be processed to obtain a first sorting result; calculating the preset interval according to the total number of the tokens to be processed and a preset proportionality coefficient; taking the tokens to be processed within the preset interval in the first sorting result as the target tokens; arranging the target tokens according to the value evaluation score of the target tokens to obtain a second sorting result; calculating the preset interval according to the total number of the target tokens and the preset proportionality coefficient; screening out the tokens within the preset interval from the second sorting result for iterative training of the pre-trained language model until a converged pre-trained language model is obtained; wherein, the preset proportionality coefficient gradually decreases as the number of iterative steps increases during the iterative training of the pre-trained language model.
2. The method according to claim 1, characterized in that The obtaining a token to be processed includes: obtaining multimodal data; segmenting and / or encoding the multimodal data to obtain the token to be processed.
3. The method according to claim 1, characterized in that, The pre-trained language model includes a first model and a second model, and the performance of the second model is better than that of the first model; The calculating a value evaluation score of a token to be processed based on a loss function of a pre-trained language model includes: calculating a first loss function of the token to be processed under a preset context window and first model parameters, and a second loss function of the token to be processed under the preset context window and second model parameters; calculating the value evaluation score based on the difference between the second loss function and the first loss function.
4. The method according to claim 3, wherein The calculating the value evaluation score based on the difference between the second loss function and the first loss function includes: obtaining the occurrence frequency of the token to be processed in the corpus; calculating the value evaluation score according to the difference between the second loss function and the first loss function and the occurrence frequency.
5. The method according to claim 1, characterized in that, The arranging the tokens to be processed according to the value evaluation score and screening out target tokens within a preset interval includes: clustering the tokens to be processed according to the part of speech of the tokens to be processed to obtain at least one part-of-speech clustering result; arranging the tokens to be processed in the at least one part-of-speech clustering result according to the value evaluation score, and screening out target tokens within a preset pre-order interval according to a preset deletable threshold.
6. The method according to claim 1, characterized in that The obtaining a token to be processed includes: dividing the data sample set of the pre-trained language model into multiple data sample subsets; fine-tuning the basic model parameters of the pre-trained language model by using any one of the data sample subsets to obtain initial model parameters; Obtain the to-be-processed tokens from other subsets of data samples, where the other subsets of data samples are the subsets of data samples other than any one of the multiple subsets of data samples.
7. The method according to claim 6, characterized in that, After arranging the to-be-processed tokens according to the value evaluation scores and screening out the target tokens within a preset interval, the method further includes: Using the target tokens corresponding to the other subsets of data samples within the preset interval to finely tune the initial model parameters to obtain reference model parameters; Calculating a third loss function of the target tokens under a preset context window and basic model parameters, and a fourth loss function of the target tokens under the preset context window and reference model parameters; Calculating a value evaluation score based on the difference between the fourth loss function and the third loss function; Arranging the target tokens according to the value evaluation score, screening out the tokens within the preset interval for iterative training of the pre-trained language model until a converged pre-trained language model is obtained.
8. The method according to claim 1, wherein After arranging the target tokens according to the value evaluation score to obtain a second sorting result, before calculating the preset interval according to the total number of the target tokens and a preset proportionality coefficient, the method further includes: Determining the current iteration step and the total number of iteration steps of the pre-trained language model; Calculating the preset proportionality coefficient according to a first proportionality coefficient boundary value, a second proportionality coefficient boundary value, the current iteration step, and the total number of iteration steps, where the first proportionality coefficient boundary value is greater than the second proportionality coefficient boundary value, so that the preset proportionality coefficient gradually decreases from the first proportionality coefficient boundary value to the second proportionality coefficient boundary value during the iterative training of the pre-trained language model.
9. The method according to claim 8, wherein After arranging the target tokens according to the value evaluation score, screening out the tokens within the preset interval for iterative training of the pre-trained language model until a converged pre-trained language model is obtained, the method further includes: Using the converged pre-trained language model to perform a downstream task and determining the accuracy of the converged pre-trained language model in the downstream task; Adjusting the first proportionality coefficient boundary value and the second proportionality coefficient boundary value according to the accuracy for feedback iteration of the pre-trained language model.
10. A data cleaning device, characterized in that, Including: An acquisition module for acquiring to-be-processed tokens; A calculation module for calculating a value evaluation score of the to-be-processed tokens based on the loss function of the pre-trained language model, where the value evaluation score is used to measure the prediction accuracy of different pre-trained language model parameters for the to-be-processed tokens; A screening module for arranging the to-be-processed tokens according to the value evaluation score and screening out the target tokens within a preset interval; Specifically, the screening module arranges the to-be-processed tokens according to the value evaluation score of the to-be-processed tokens to obtain a first sorting result; Calculating the preset interval according to the total number of the to-be-processed tokens and the preset proportionality coefficient; Regarding the to-be-processed tokens within the preset interval in the first sorting result as the target tokens; Arrange the target tokens according to the value evaluation scores of the target tokens to obtain a second sorting result; Calculate the preset interval according to the total number of the target tokens and a preset proportionality coefficient; Filter out the tokens within the preset interval from the second sorting result for iterative training of the pre-trained language model until a converged pre-trained language model is obtained; Wherein, the preset proportionality coefficient is gradually reduced as the number of iterative steps increases during the iterative training of the pre-trained language model.
11. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the data cleaning method according to any one of claims 1 to 9 when executing the computer program.
12. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the data cleaning method according to any one of claims 1 to 9 when executed by a processor.
13. A computer program product comprising a computer program, characterized in that, The computer program implements the steps of the data cleaning method according to any one of claims 1 to 9 when executed by a processor.
Citation Information
Patent Citations
Language model fine tuning method, text classification method, device and equipment
CN114896395A