LLMs pre-training data set optimization method and device and storage medium
By introducing a confusion algorithm based on hidden Markov model optimization, sliding window and weighted average calculation in the LLMs pre-trained data set, the problems of inefficiency and noise impact of traditional data preparation methods are solved, and high-quality data screening and model performance are improved.
Patent Information
- Application Number
- CN202510184713.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
In the pre-training stage of LLMs large-scale model, traditional data preparation methods are inefficient and difficult to adapt to the needs of large-scale data processing. Directly using the existence of noise and irrelevant information when exposing large-scale data sets, it may dilute the attention and efficiency of the model and affect the model performance.
The confusion algorithm based on the hidden Markov model optimization is adopted, and the confusion calculation of the text fragment by fragment is performed on the text through a sliding window. Combined with the weighted average calculation, the comprehensive confusion score of the entire text is obtained, which is used to filter the semantic confusion statements inside the data set.
Effectively filter out semantic confusing sentences in the dataset, convert freely obtained low-quality corpus text into high-quality and valuable corpus, save the pre-training cost of LLMs and improve model capabilities.
Smart Images

Figure CN120105100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of LLMs pre-training data processing, and in particular to a LLMs pre-training data set optimization method, device and storage medium. Background Art
[0002] In the pre-training stage of LLMs large models, the preparation and optimization of data sets are the key to improving model performance. Traditional data preparation methods mainly include manual screening of data sets and direct use of public large-scale data sets. Although manual screening of data sets can ensure data quality, it is inefficient and difficult to adapt to large-scale data processing needs. The method of directly using public large-scale data sets reduces the workload of manual screening, but because the data sets contain a lot of noise and irrelevant information, it may dilute the model's attention and efficiency, affecting the performance of the final model.
[0003] The current conventional data preparation and optimization method is to preprocess and screen massive data through automated tools and algorithms to improve data quality and relevance. In addition, for proprietary data in specific fields or languages, such as the creation of Colossal Clean Crawled Corpus (C4), high-quality text data is extracted by applying a series of filters to a single snapshot of the CommonCrawl dataset to enhance the performance of the model on specific tasks. This approach not only improves the quality of the data, but also facilitates a better understanding of the source, structure and potential bias of the data through detailed documentation and analysis.
[0004] Among them, efficient data preparation and precise data optimization are the key to improving model performance. In the pre-training stage, the use of automated tools and algorithms for data preprocessing and screening, as well as the use of proprietary data, have been proven to effectively improve data quality and model performance. In the fine-tuning stage, selectively using a small amount of high-quality data rather than relying on large-scale data sets has become an effective strategy to improve model targeting and efficiency.
[0005] In order to improve the speed and accuracy of LLMs training, it is necessary to accurately identify and utilize data that can comprehensively cover key dimensions and have the highest value. However, conventional LLMs pre-training processes generally face a lack of corpus sources and a variety of low-quality problems in large-scale web crawling data. These problems lead to high pre-training costs and unsatisfactory results.
[0006] To this end, this application specifically proposes a LLMs pre-training data set optimization method, device and storage medium, which can achieve technical in-depth screening of large-scale low-quality data, improve data quality, and effectively optimize the LLMs training process to solve the above-mentioned technical problems. Summary of the invention
[0007] The main purpose of the present invention is to provide a method, device and storage medium for optimizing LLMs pre-training data sets. By introducing a perplexity algorithm based on hidden Markov model optimization, sentences with confusing semantics in the data set can be effectively screened out, and low-quality corpus texts obtained for free on the Internet can be converted into high-quality and valuable corpus, effectively saving LLMs pre-training costs and improving model capabilities, so as to solve the technical problems raised in the background technology.
[0008] The present invention adopts the following technical solutions to solve the above technical problems:
[0009] A method for optimizing a LLMs pre-training data set, comprising:
[0010] Select the data in the dataset and use a sliding window to calculate the perplexity of the hidden Markov model for each text segment;
[0011] Based on the hidden Markov model perplexity calculation result, a comprehensive perplexity score of the entire text is obtained by perplexity weighted average calculation, which is used to filter out sentences with semantic confusion within the data set.
[0012] Preferably, the sliding window size is set to 10, and the step size of the sliding window is set to 1 word.
[0013] Preferably, the hidden Markov model perplexity calculation is used to select n words closest to the current pointing position for perplexity calculation, and the specific calculation formula is:
[0014]
[0015] in represents the cross entropy loss, where W represents the word sequence in the test set, N represents the total number of words in the sequence, and P(w i |w i-1 ,…,w i=n+1 ) means that the model predicts that given all the previous words, the next word is w i probability.
[0016] Preferably, the specific calculation formula for the perplexity weighted average calculation is:
[0017]
[0018] Where M is the total number of sliding windows, PPL(w j ) is the perplexity of the jth window, w j is the corresponding window weight coefficient.
[0019] Preferably, the method further includes multiplying one-hot encoding with the probability matrix of the entire word list to calculate the probability distribution of the next word sequence based on the words appearing in the current window to predict the probability of the next word appearing.
[0020] Preferably, a preset perplexity threshold is also included, and data that potentially interferes with model training is eliminated based on the comprehensive perplexity score result and the perplexity threshold.
[0021] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0022] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0023] It can be seen from the above technical solution that the present invention provides a method, device and storage medium for optimizing LLMs pre-training data set. Compared with the prior art, the present invention has the following advantages:
[0024] 1. The present invention uses a sliding window to perform perplexity evaluation on local areas of the text, which can more accurately identify abnormal or complex fragments in the text and effectively filter out potential negative impacts on model training. Compared with the traditional method of directly evaluating the perplexity of the entire text, the present method analyzes the local complexity of the text fragments, and can more carefully identify and eliminate data that potentially interferes with model training. While reducing the computational complexity, it also more specifically grasps the key elements for evaluating text quality.
[0025] 2. The present invention adopts the weighted average method to integrate the perplexity scores of each window, which can provide a comprehensive quality assessment index for the entire text. Compared with a single perplexity calculation, it can more comprehensively reflect the quality status of the text.
[0026] 3. Based on the traditional perplexity calculation, the present invention introduces sliding window technology and weighted average strategy to more accurately evaluate the language model fitness of the text, effectively screen out semantically confusing sentences in the data set, and convert low-quality corpus texts obtained for free on the Internet into high-quality and valuable corpus, effectively saving the LLMs pre-training cost and improving the model capabilities.
[0027] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become easy to understand through the following description. Of course, it is not necessary to achieve all of the advantages described above simultaneously for any product implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0029] Figure 1 It is a logical schematic diagram of the overall steps of the method of the present invention;
[0030] Figure 2 It is a schematic diagram of the operation of calculating the perplexity of the hidden Markov model by sliding the window in the method of the present invention;
[0031] Figure 3 The figure is a schematic diagram of the operation flow of the perplexity weighted average calculation in the method of the present invention. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0033] In the embodiment, see Figures 1 to 3 .
[0034] Efficient data preparation and accurate data optimization are the key to improving model performance. In the LLMs pre-training stage, the use of automated tools and algorithms for data preprocessing and screening, as well as the use of proprietary data, have been proven to effectively improve data quality and model performance. In the fine-tuning stage, selectively using a small amount of high-quality data rather than relying on large-scale data sets has become an effective strategy to improve model targeting and efficiency.
[0035] Therefore, the embodiment of the present invention proposes a method for optimizing LLMs pre-training data sets, which effectively screens out semantically confusing sentences in the data set by introducing a perplexity algorithm based on hidden Markov model optimization, and can convert low-quality corpus texts obtained for free from the Internet into high-quality and valuable corpus, effectively saving LLMs pre-training costs and improving model capabilities.
[0036] refer to Figure 1 , the method specifically comprises:
[0037] (1) Select the internal data of the dataset, such as Figure 2 As shown, a sliding window is used to calculate the hidden Markov model perplexity of the text segment by segment;
[0038] (2) Based on the perplexity calculation results of the hidden Markov model, such as Figure 3 As shown in the figure, the comprehensive perplexity score of the entire text is obtained by weighted average calculation of perplexity, which is used to filter out sentences with confusing semantics within the dataset.
[0039] Here, the perplexity of local areas of the text is evaluated through a sliding window, which can more accurately identify abnormal or complex fragments in the text and effectively filter out potential negative impacts on model training. At the same time, the weighted average method is used to integrate the perplexity scores of each window, providing a comprehensive quality evaluation indicator for the entire text. Compared with a single perplexity calculation, it can more comprehensively reflect the quality status of the text.
[0040] It should be noted that perplexity, as an indicator to measure the predictive ability of a language model, is generally used to reflect the model's prediction accuracy for the next word under given contextual conditions. That is, perplexity is based on the overall probability of the model's prediction of the test data set, and is used to quantify the model's effectiveness in processing natural language processing tasks. An efficient language model should have a lower perplexity. A lower perplexity PPL means that the model can better understand the text and can more accurately predict or generate natural language text.
[0041] Specifically, the calculation of perplexity is based on the cross-entropy loss of the model for the test set. An ideal language model should have a lower perplexity, indicating that it has a stronger ability to understand and predict text data. The calculation of perplexity here depends on the cross-entropy loss of the model for the test set, which in turn reflects the average branching factor of the model when predicting the next word. The cross-entropy loss, as the loss function used for back-propagation of the language model, measures the similarities and differences between the probability distribution predicted by the model and the true probability distribution.
[0042] Therefore, this application introduces sliding window technology and weighted average strategy on the basis of traditional perplexity calculation to facilitate more accurate evaluation of the language model fitness of the text. Compared with directly evaluating the perplexity of the entire text, the method of this application analyzes the local complexity of text fragments, which can more carefully identify and eliminate data that potentially interferes with model training. It reduces the computational complexity while more specifically grasping the key elements of evaluating text quality.
[0043] Among them, the hidden Markov model perplexity calculation is optimized based on the traditional PPL algorithm, integrating the hidden Markov model. Not all words in the whole paragraph are involved in the calculation, but it is used to select the n words closest to the current pointing position. Based on this, it is defined as the perplexity of the given language model L for the test set T. The specific calculation formula is:
[0044]
[0045] in represents the cross entropy loss, where W represents the word sequence in the test set, N represents the total number of words in the sequence, and P(w i |w i-1 ,…,w i=n+1 ) means that the model predicts that given all the previous words, the next word is w i probability.
[0046] In addition, in a specific embodiment, for each text segment with a window size of 10, its perplexity is calculated, and the comprehensive perplexity score of the entire text is obtained by weighted average, and the step size of the sliding window is set to 1 word to ensure comprehensive coverage of the text.
[0047] At this time, the specific calculation formula for the weighted average calculation of confusion is:
[0048]
[0049] Where M is the total number of sliding windows, PPL(w j ) is the perplexity of the jth window, w j is the corresponding window weight coefficient.
[0050] In the further specific application implementation process, it also includes using one-hot encoding and the full word list probability matrix to do the product, based on the words appearing in the current window, calculate the probability distribution of the next word sequence. At this time, it is necessary to ensure that the sum of the probability values is 1. When the sum of the probability values is not 1, it is also necessary to perform normalization processing. After that, the probability obtained by calculation can be used to predict the probability of the next word appearance.
[0051] In the further specific application implementation process, it also includes presetting a perplexity threshold and eliminating data that potentially interferes with model training based on the comprehensive perplexity score results and the perplexity threshold.
[0052] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0053] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0054] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the LLMs pre-training data set optimization methods in the above embodiments.
[0055] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above method.
[0056] The embodiment of the present application also provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus.
[0057] Memory, used to store computer programs;
[0058] The processor is used to implement the above-mentioned LLMs pre-training data set optimization method when executing the program stored in the memory.
[0059] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0060] The communication interface is used for communication between the above electronic device and other devices.
[0061] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0062] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components.
[0063] It should also be noted that electronic devices also include terminal devices, which can also be called terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablet computers (Pad), computers with wireless transceiver functions, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of the present application do not limit the specific technology and specific device form adopted by the terminal devices.
[0064] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk SolidState Disk), etc.
[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
[0066] In addition, it should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components in a certain specific posture. If the specific posture changes, the directional indication will also change accordingly.
[0067] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. In addition, the meaning of "and / or" appearing in the full text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or a scheme in which A and B are satisfied at the same time. In addition, in the embodiments of the present invention, "multiple" refers to more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in the field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
Claims
1. A method for optimizing a LLMs pre-training data set, characterized in that: include: Select the data in the dataset and use a sliding window to calculate the perplexity of the hidden Markov model for each text segment; Based on the hidden Markov model perplexity calculation result, a comprehensive perplexity score of the entire text is obtained by perplexity weighted average calculation, which is used to filter out sentences with semantic confusion within the data set.
2. The LLMs pre-training data set optimization method according to claim 1, characterized in that: The sliding window size is set to 10, and the step size of the sliding window is set to 1 word.
3. The LLMs pre-training data set optimization method according to claim 1, characterized in that: The hidden Markov model perplexity calculation is used to select the n words closest to the current pointing position for perplexity calculation. The specific calculation formula is: in represents the cross entropy loss, where W represents the word sequence in the test set, N represents the total number of words in the sequence, and P(w i |w i-1 ,…,w i=n+1 ) means that the model predicts that given all the previous words, the next word is w i probability.
4. The LLMs pre-training data set optimization method according to claim 3, characterized in that: The specific calculation formula for the perplexity weighted average calculation is: Where M is the total number of sliding windows, PPL(w j ) is the perplexity of the jth window, w j is the corresponding window weight coefficient.
5. The LLMs pre-training data set optimization method according to claim 1, characterized in that: It also includes the use of one-hot encoding and the full word list probability matrix to do the product, based on the words appearing in the current window, calculate the probability distribution of the next word sequence to predict the probability of the next word appearing.
6. The LLMs pre-training data set optimization method according to claim 1, characterized in that: It also includes a preset perplexity threshold, which eliminates data that potentially interferes with model training based on the comprehensive perplexity score results and the perplexity threshold.
7. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.
Citation Information
Cited By
Screening method and equipment of instruction data, medium and product
CN120653995A
Pancreatic cancer prediction method and system based on local and global confusion weighted pruning
CN121839088A