Pre-training data processing method and device, medium, equipment and program product

By performing quality and relevance predictions on the initial dataset of a large language model, and then filtering and refining the results based on the evaluation, the problem of insufficient performance of large language models in professional knowledge domains is solved, the model's task capability in this domain is improved, and the training cost is reduced.

CN122175036APending Publication Date: 2026-06-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411796365.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing large-scale language models perform poorly on tasks in the domain of professional knowledge. Manually constructed supervised instruction fine-tuning and reinforcement learning training data are costly and limited by the scale of data, resulting in technical bottlenecks.

Method used

By acquiring the initial dataset of the target knowledge domain, the first evaluation model and the second evaluation model are used to predict data quality and domain relevance, respectively. The data is filtered by combining the quality evaluation results and the relevance prediction results to obtain the intermediate dataset. The target pre-training set is obtained through fine-tuning and filtering. A progressive funnel strategy is adopted to reduce the cost of constructing training data.

Benefits of technology

It significantly improves the model's ability to perform tasks in the domain of professional knowledge, reduces the resource overhead of training data, balances model performance, and mines high-quality pre-training data in the domain of professional knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122175036A_ABST
    Figure CN122175036A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, medium, device, and program product for processing pre-training data, relating to the field of artificial intelligence technology. The method includes: acquiring an initial dataset for a target knowledge domain, the initial dataset comprising multiple pre-training data retrieved for the target knowledge domain; performing data quality prediction on the pre-training data in the initial dataset based on a first evaluation model to obtain a quality evaluation result; performing domain relevance prediction on the pre-training data in the initial dataset based on a second evaluation model to obtain a domain relevance prediction result, the relevance prediction result indicating the relevance between the pre-training data and the target knowledge domain, the first and second evaluation models being constructed based on a large-scale language model; filtering the initial dataset based on the quality evaluation result and the relevance prediction result to obtain an intermediate dataset; and performing fine-tuning and filtering on the intermediate dataset to obtain a target pre-training set. This application can efficiently acquire a large amount of high-quality pre-training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to methods, apparatus, media, equipment, and program products for processing pre-training data. Background Technology

[0002] The impressive performance of Large Language Models (LLMs) across various tasks largely stems from the rich world knowledge contained in the pre-training data. Generally, pre-training data comes from a wide range of sources, with open-source web data being a major one. This data contains web pages of varying quality and covers a broad range of topics, allowing LLMs to acquire a broad range of capabilities. However, within open-source web data, the proportion of content related to specialized domains, such as mathematical logic, is often very small. Therefore, training LLMs solely with general pre-training data results in poor performance on tasks requiring specialized knowledge.

[0003] To further improve the performance of LLMs in specialized domains such as mathematical logic reasoning, related solutions often add Supervised Fine-Tuning (SFT) and reinforcement learning stages after the pre-training phase. These stages use manually labeled, high-quality data to teach the model how to handle tasks in specialized domains, such as teaching it how to solve mathematical reasoning problems through manually labeled mathematical problem-solving exercises. However, the training data costs for manually constructed SFT and reinforcement learning are high. Limited by data scale, the SFT and reinforcement learning stages face technical bottlenecks in improving the performance of LLMs in specialized domain tasks. Summary of the Invention

[0004] This application provides a method, apparatus, medium, device, and program product for processing pre-training data. The technical solution is as follows:

[0005] On the one hand, this application provides a method for processing pre-trained data, the method comprising:

[0006] Obtain an initial recall dataset for the target knowledge domain, wherein the initial recall dataset includes multiple pre-trained data recalled for the target knowledge domain;

[0007] Based on the first evaluation model, the pre-trained data of the initial recruitment dataset is used to predict data quality and obtain quality evaluation results.

[0008] Based on the second evaluation model, the pre-training data of the initial recall dataset is used to predict the domain relevance, and the domain relevance prediction result is obtained. The relevance prediction result is used to indicate the relevance between the pre-training data and the target knowledge domain. The first evaluation model and the second evaluation model are constructed based on a large language model.

[0009] Based on the quality evaluation results and the correlation prediction results, the initial dataset is filtered to obtain an intermediate dataset.

[0010] The intermediate dataset is then sorted and filtered to obtain the target pre-training set.

[0011] On the other hand, this application provides a pre-training data processing apparatus, the apparatus comprising:

[0012] Acquisition module: used to acquire the initial recall dataset for the target knowledge domain, the initial recall dataset including multiple pre-trained data recalled for the target knowledge domain;

[0013] Quality evaluation module: used to predict the data quality of the pre-trained data of the initial dataset based on the first evaluation model, and obtain the quality evaluation result;

[0014] The relevance prediction module is used to predict the domain relevance of the pre-training data of the initial dataset based on the second evaluation model, and obtain the domain relevance prediction result. The relevance prediction result is used to indicate the relevance of the pre-training data to the target knowledge domain. The first evaluation model and the second evaluation model are constructed based on a large language model.

[0015] Data filtering module: used to filter the initial dataset based on the quality evaluation results and the correlation prediction results to obtain an intermediate dataset;

[0016] Fine-ranking and filtering module: used to fine-rank and filter the intermediate dataset to obtain the target pre-training set.

[0017] On the other hand, this application provides a computer-readable storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the pre-trained data processing method as described above.

[0018] On the other hand, this application provides a computer device including a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or at least one program being loaded and executed by the processor to implement the pre-training data processing method as described above.

[0019] On the other hand, this application provides a computer program product, which includes computer instructions that, when executed by a processor, implement the pre-training data processing method as described above.

[0020] The pre-training data processing method, apparatus, medium, equipment, and program products provided in this application have the following technical effects:

[0021] The technical solution of this application first obtains an initial dataset for the target knowledge domain, and then uses a first evaluation model and a second evaluation model to perform data quality prediction and relevance prediction respectively to evaluate the quality of the pre-training data and its relevance to the target knowledge domain. Then, the initial screening of the initial dataset is achieved by combining the quality evaluation results and the relevance prediction results to obtain an intermediate dataset. This improves the quality of the data to be finely ranked while reducing the amount of data, so as to facilitate fine ranking and screening of the intermediate dataset to obtain the target pre-training set. This allows for the mining of high-quality pre-training data in the professional knowledge domain, thereby increasing the proportion of data in that knowledge domain during the pre-training process and significantly improving the model's task capability in the professional knowledge domain. Furthermore, a progressive funnel strategy is adopted in the pre-training data construction process to reduce the cost of training data construction, improve the data filtering effect, and balance resource consumption and model performance.

[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0023] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application;

[0025] Figure 2 This is a flowchart illustrating a method for processing pre-trained data provided in an embodiment of this application;

[0026] Figure 3 This is a flowchart illustrating another method for processing pre-trained data provided in an embodiment of this application;

[0027] Figure 4 This is a flowchart illustrating another method for processing pre-trained data provided in an embodiment of this application;

[0028] Figure 5This is a flowchart illustrating another method for processing pre-trained data provided in an embodiment of this application;

[0029] Figure 6 This is a flowchart illustrating another method for processing pre-trained data provided in an embodiment of this application;

[0030] Figure 7 This is a flowchart illustrating a method for processing pre-training data according to a specific embodiment of this application.

[0031] Figure 8 This is a schematic diagram of the structural framework of a pre-training data processing device provided in an embodiment of this application;

[0032] Figure 9 This is a schematic diagram of the hardware structure of a device for implementing a pre-training data processing method, provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0035] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application, such as... Figure 1As shown, the application environment may include at least server 01. In this embodiment, server 01 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0036] Specifically, the server 01 mentioned above may include a physical device, which may include a network communication submodule, a processor and a memory, etc., and may also include software running on the physical device, which may include applications, etc.

[0037] In this embodiment of the application, server 01 can obtain the initial dataset of the target knowledge domain, and then use the first evaluation model and the second evaluation model to perform data quality prediction and relevance prediction respectively, to obtain quality evaluation results and relevance prediction results. Then, the initial screening of the initial dataset is achieved by combining the quality evaluation results and relevance prediction results to obtain the intermediate dataset. Then, the intermediate dataset is finely sorted and screened to obtain the target pre-training set.

[0038] Furthermore, it is understandable that Figure 1 The illustration only depicts an application environment for a pre-training data processing method. This application environment may include more or fewer nodes, and this application is not limited thereto. For example, it may also include a terminal, and the terminal and server 01 can be directly or indirectly connected via wired or wireless communication, and this application is not limited thereto. Specifically, the terminal may include physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart voice interaction devices, smart home appliances, smart wearable devices, and in-vehicle terminal devices, and may also include software running on the physical device, such as applications.

[0039] It is understood that, in the specific embodiments of this application, the open-source web page data, pre-training data and other related data involved need to obtain the permission or consent of the user or related parties when the embodiments of this application are applied to specific products or technologies, and the collection, use and processing of related data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0040] The following description, in conjunction with the accompanying drawings, introduces a method for processing pre-trained data provided in this application, which can be applied to the server side or the terminal side. Figure 2This is a flowchart illustrating a method for processing pre-trained data according to an embodiment of this application. This application provides method operation steps as shown in the embodiments or flowcharts, but based on conventional or non-inventive methods, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server product execution, the method can be executed sequentially according to the embodiments or drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Please refer to... Figure 2 The pre-training data processing method provided in this application embodiment may include the following steps S201-S209:

[0041] S201: Obtain the initial dataset for the target knowledge domain.

[0042] Specifically, the initial recall dataset includes multiple pre-trained datasets for recalling the target knowledge domain. This recall can be implemented using a recall model; for example, the recall model may include, but is not limited to, the FastText model. The target knowledge domain can be a specialized knowledge domain, such as a professional academic subject; for example, it could be the domain of mathematical logic reasoning or the domain of materials synthesis.

[0043] In some embodiments, S201 may include S301-S303:

[0044] S301: Obtain the initial dataset.

[0045] Specifically, the initial dataset includes multiple datasets to be recalled. The data sources for these datasets include, but are not limited to, open-source websites and platforms. The datasets can be open-source web page data. Optionally, a database containing a general pre-trained data pool can be maintained to store the initial dataset. The initial recall dataset can be obtained by recalling data from this database.

[0046] In some embodiments, the data in the initial dataset is obtained after data preprocessing. Accordingly, the process of obtaining the initial dataset includes: obtaining multiple open-source web page data; performing deduplication on the multiple open-source web page data; and generating the initial dataset based on the deduplicated open-source web page data.

[0047] Specifically, a data candidate pool can be constructed to store the initially acquired open-source web page data. This data is then extracted and deduplicated to filter out homogeneous web page data, which is then stored as the data to be recalled in a pre-training data pool, resulting in the initial dataset. The open-source web page data can be obtained by extracting the main content of open-source pages using page extraction and parsing tools. For example, the open-source web page data can come from platforms such as CC, refinedWeb, fineweb, and datacamp. In one example, the initial dataset size is close to the petabyte scale. Thus, by acquiring and deduplicating the open-source web page data, retrievable data is initially screened to avoid excessive homogeneity in the recalled data.

[0048] S303: Based on the recall model, retrieve the data to be recalled from the initial dataset in the target knowledge domain to obtain the initial recall dataset.

[0049] Specifically, the recall model is trained under constraints based on a pre-trained data recognition system for the target knowledge domain, using a preset ratio of positive to negative training samples. For example, the preset ratio of positive to negative training samples could be 1:5, and the recall model could be a FastText model, etc. By further filtering the initial dataset for the target knowledge domain using the recall model, irrelevant data is filtered out, reducing the amount of data required for subsequent data evaluation and improving the efficiency and accuracy of data filtering. Furthermore, the use of a certain ratio of positive and negative examples during training effectively enhances the recall capability of the recall model.

[0050] Specifically, positive training samples are data belonging to the target knowledge domain and meeting data quality conditions, while negative training samples are obtained by randomly sampling from the initial dataset. For example, if the target knowledge domain is mathematical logic reasoning, positive training samples are high-quality open-source mathematical webpage data, and negative training samples are randomly sampled from general domain data to ensure that negative examples cover data from almost all domains, improving the recognition performance of the recall model. Positive and negative training samples can include data in multiple languages, such as Chinese and English data, or some Chinese data can be translated into English.

[0051] In one example, during the initial data retrieval phase for the mathematical logic reasoning domain, the expected targets for recall include, but are not limited to, pages containing rich mathematical knowledge such as mathematical theorems, definitions, proofs, mathematical problems and answers, traditional mathematics documents, or interdisciplinary documents with mathematical formulas in fields such as physics, chemistry, biology, economics, and finance. The FastText model is used to rapidly recall data in the data logic reasoning domain on a p-level initial dataset, achieving high recall of math-related pages with low computational overhead and resource consumption. Testing and evaluation show that for major known open-source mathematical platforms (such as math.stackexchange), the FastText model achieves a recall rate of approximately 85%, with unrecalled pages primarily being weakly math-related. Furthermore, the proportion of math-related pages in the initial dataset's general pre-training data is generally around 0.6%, meaning approximately 8TB of pre-training data can be recalled from a few p-level of general pre-training data. Considering that some general pre-training data includes data in less commonly spoken languages, further filtering of this language data can be performed after the initial retrieval if training on such data is not performed.

[0052] S203: Based on the first evaluation model, perform data quality prediction on the pre-trained data of the initial dataset to obtain the quality evaluation results.

[0053] Specifically, the first evaluation model is built upon a large-scale language model, which can be fine-tuned using open-source, general-purpose web data combined with instruction learning. For example, it can be fine-tuned based on a 1B-scale LLM for data quality prediction, the size of which can be chosen based on the size of the initial dataset. The first evaluation model predicts data quality on the pre-training data of the initial dataset based on discrete tiered evaluations. Correspondingly, the quality evaluation result is obtained by fusing multiple discrete tiered evaluation results; the quality evaluation result is used to indicate the data quality of the pre-training data. Thus, low-quality data is filtered out through quality tiering.

[0054] In some embodiments, reference is made to Figure 3 S203 may include S401-S403:

[0055] S401: Obtain the first instruction text;

[0056] S403: Input the pre-trained data of the first instruction text and the initial recall dataset into the first evaluation model to perform quality predictions for each of the multiple data quality evaluation items on the pre-trained data, and fuse the obtained first prediction scores to obtain the quality evaluation result.

[0057] Specifically, the first instruction text provides guidance information needed for the first evaluation model to perform data quality prediction, including multiple data quality evaluation items. These multiple data quality evaluation items can be set based on the needs of multiple dimensions of data quality evaluation, enabling the first evaluation model to perform the evaluation tasks corresponding to each data quality evaluation item and output corresponding first prediction scores. The total score obtained after summing these first prediction scores is used as the quality evaluation result. This summing process can be simple summing or weighted summing, etc. Understandably, the fusion method is not limited to the above summing method; other methods that can generate quality evaluation results, such as averaging or multiplication, can also be used, which will not be enumerated here. By using multiple data quality evaluation items to guide the large language model in performing multi-dimensional data quality evaluation, low-quality data is effectively filtered, improving the accuracy of high-quality web page data selection.

[0058] In one example, the initial dataset for mathematical logic reasoning is approximately 8TB in size. A first evaluation model was built based on 1B-scale LLMs. The first prompt text can be exemplified as follows:

[0059] "The following is a webpage. To assess whether this page is of high quality, use the additive 5-point scoring system described below. Accumulate points based on the satisfaction of each difficulty:"

[0060] - If the webpage content contains serious quality issues such as toxicity, bias, or privacy breaches, or if the content is severely fragmented, contains a large number of serious grammatical errors or noisy segments, add 0 points.

[0061] - If the webpage content has very little information or almost no actual content, such as being very short or mostly consisting of repetitive short sentences, add 0 points.

[0062] - If the webpage content has a certain amount of information, but contains a lot of repetitive, simple or templated expressions, add 1 point.

[0063] - If the webpage touches on a specific topic, provides basic information but the content is rather superficial. It contains some grammatical errors or a small amount of irrelevant content, but has a certain level of formatting and organization, add 1 point.

[0064] - If the webpage content is detailed, covers multiple aspects of the topic, the information is accurate and useful, the layout is relatively clear, and there are only minor inconsistencies or small grammatical errors that do not affect overall comprehension, add 1 point.

[0065] - If the webpage content is very detailed and in-depth, grammatically correct, neatly formatted, non-toxic, unbiased, and does not involve privacy breaches, and provides unique insights or a degree of innovation, add 1 point.

[0066] - If a webpage excels in content, offering valuable and unique insights or innovative solutions that provide high guidance for users and has a clear layout, add 1 point.

[0067] Web page:

[0068] %s

[0069] After checking the webpage:

[0070] Briefly describe your total score in no more than 100 words.

[0071] - Use the following format to derive the score: "Score:<Total Score>"

[0072] S205: Based on the second evaluation model, perform domain relevance prediction on the pre-trained data of the initial dataset to obtain the domain relevance prediction results.

[0073] Specifically, the second evaluation model is built upon a large-scale language model; it can be fine-tuned based on open-source general web page data combined with instruction learning, such as fine-tuning training based on 1B-scale LLMs for domain relevance prediction. The size of the LLMs can be selected based on the data size of the initial dataset. The second evaluation model predicts domain relevance of the pre-training data in the initial dataset based on discrete tiered evaluations. Accordingly, the relevance prediction result is obtained by fusing multiple discrete tiered evaluation results; the relevance prediction result is used to indicate the relevance between the pre-training data and the target knowledge domain. In this way, low-relevance data is filtered out through relevance tiering.

[0074] In some embodiments, reference is made to Figure 4 S203 may include S501-S503:

[0075] S501: Obtain the second instruction text;

[0076] S503: Input the pre-trained data of the second instruction text and the initial recall dataset into the second evaluation model to predict the relevance of each of the multiple relevance evaluation items of the pre-trained data, and fuse the obtained second prediction scores to obtain the relevance prediction result.

[0077] Specifically, the second instruction text provides guidance information needed for the second evaluation model to perform relevance prediction, including multiple relevance evaluation items. These multiple relevance evaluation items can be set based on the needs of multiple dimensions of domain relevance evaluation, enabling the second evaluation model to perform the evaluation tasks corresponding to each relevance evaluation item and output corresponding second prediction scores. The total score obtained after summing these second prediction scores is used as the relevance prediction result. This summing process can be simple summing or weighted summing, etc. Understandably, the fusion method is not limited to the above summing method; other methods that can generate relevance prediction results, such as averaging or multiplication, can also be used, which will not be enumerated here. By using multiple relevance evaluation items to guide the large language model in performing multi-dimensional relevance evaluation, weakly relevant data can be effectively filtered, improving the accuracy of filtering high-quality domain-specific web page data.

[0078] In one example, the initial dataset for mathematical logic reasoning is approximately 8TB in size. A second evaluation model was built based on 1B-scale LLMs. The second prompt text can be exemplified as follows:

[0079] "The following is a webpage. Evaluate whether this page has high mathematical value and whether it is useful in a mathematical context for teaching from elementary to university. Use the 5-point grading system for addition described below. Accumulate points based on the satisfaction of each difficulty level:"

[0080] - If a webpage provides basic information related to mathematical concepts, problems, or methods, but may also include content from other non-mathematical disciplines, such as advertisements or information from other subjects, this information, while related to mathematics, is not the primary focus. Add 1 point.

[0081] - Add 1 point if a webpage contains elements related to mathematics, even if these elements are not explained in detail according to strict mathematical standards. For example, a webpage discusses financial markets and mentions some mathematical concepts such as generals and statistics, but these concepts are not explained in detail according to mathematical standards; they are simply mentioned without going into depth.

[0082] - If a webpage is suitable for math learning and application, introducing key concepts relevant to the school curriculum, it provides valuable math resources for learners overall, even if the content is not comprehensive or contains some irrelevant information. It might resemble an introductory section of a textbook or a basic tutorial, suitable for learning but with significant limitations, earning 1 point. For example, a webpage introduces basic algebraic concepts but also includes some discussions unrelated to the topic.

[0083] - If a webpage is highly relevant to mathematics and beneficial for teaching, providing a wealth of detailed mathematical content, including exercises and solutions, and containing almost no irrelevant information, it's similar to a chapter in a textbook or a detailed tutorial. Add 1 point. For example, a webpage detailing a chapter of geometry, including definitions, theorems, examples, and exercises.

[0084] - Add 1 point if a webpage excels in mathematical value, follows detailed reasoning, and is written in a clear and easy-to-understand style. It provides a deep and comprehensive understanding of the topic, focusing entirely on mathematical content without any non-mathematical elements. For example, a webpage that comprehensively and thoroughly explains the basic concepts and applications of calculus, with clear and easy-to-understand content and no non-mathematical elements.

[0085] Web page:

[0086] %s

[0087] After checking the webpage:

[0088] Briefly describe your total score in no more than 100 words.

[0089] - Use the following format to derive the score: "Mathematical Score: <Total Score>"

[0090] S207: Based on the quality assessment results and correlation prediction results, the initial dataset is filtered to obtain the intermediate dataset.

[0091] Specifically, based on the quality evaluation results and the correlation prediction results, low-quality pre-training data are filtered out, and the remaining pre-training data is used to generate an intermediate dataset. For example, any pre-training data with a score lower than 3 in either the quality evaluation result or the correlation prediction result can be filtered out when the total score is 5.

[0092] Understandably, the initial recall model only needs to guarantee the recall rate and cannot guarantee the quality of the recalled pages. The first evaluation model and the second evaluation model built by the large language model can quickly classify and evaluate the initial recall dataset through two classification tasks: data quality and knowledge relevance, so as to filter out low-quality pages and provide high-quality domain data for subsequent data ranking.

[0093] In one example, based on the first and second instruction texts from the example above, a large-scale language model of 1B scale was fine-tuned using approximately 200,000 training data points to obtain the first and second evaluation models. During inference, the first and second instruction texts were concatenated to form the first 1000 characters of the webpage text, which were then used to request the trained first and second evaluation models. After filtering out low-scoring data, approximately 500G of high-quality mathematical pre-training data was selected from the aforementioned 8T initial dataset.

[0094] S209: Perform fine sorting and filtering on the intermediate dataset to obtain the target pre-training set.

[0095] Specifically, the intermediate dataset's pre-training data can be finely sorted based on the fine-sorting model, and a target pre-training set can be generated based on a certain number of top-ranked data.

[0096] In summary, the technical solution of this application first obtains an initial dataset for the target knowledge domain, and then uses a first evaluation model and a second evaluation model to perform data quality prediction and relevance prediction respectively to evaluate the quality of the pre-training data and its relevance to the target knowledge domain. Then, the initial screening of the initial dataset is achieved by combining the quality evaluation results and the relevance prediction results to obtain an intermediate dataset. This improves the quality of the data to be finely ranked while reducing the amount of data, so as to facilitate fine ranking and screening of the intermediate dataset to obtain the target pre-training set. This allows for the mining of high-quality pre-training data in the professional knowledge domain, thereby increasing the proportion of data in that knowledge domain during the pre-training process and significantly improving the model's task capability in the professional knowledge domain. Furthermore, the pre-training data construction process adopts a progressive funnel strategy, which reduces the cost of training data construction, improves the data filtering effect, and balances resource consumption and model performance.

[0097] In some embodiments, reference is made to Figure 5 S209 may include S601-S605:

[0098] S601: Obtain the third instruction text;

[0099] S603: Input the pre-trained data of the third instruction text and the intermediate dataset into the fine ranking model for data evaluation, and obtain the fine ranking evaluation result;

[0100] S605: Based on the ranking evaluation results, the intermediate dataset is filtered to obtain the target pre-training set.

[0101] Specifically, the third instruction text provides guidance information needed for evaluating the performance of the fine-ranking model. The fine-ranking model is built upon a large language model, using the third instruction text as task guidance. Pre-trained data is input into the fine-ranking model to output a fine-ranking evaluation result corresponding to the task information in the third instruction text. This evaluation result indicates the suitability of the pre-training data with the target knowledge domain and the model to be trained. The model to be trained is a large language model within the target knowledge domain.

[0102] Understandably, while the tiered evaluation of the first and second evaluation models can filter out low-quality data, the model performance has an upper limit. At the same time, discrete scoring cannot more finely distinguish the quality of each webpage. However, by using a large language model to perform fine-grained evaluation of the intermediate dataset in conjunction with instructions, the data can be further distinguished and sorted to further filter out high-quality domain training data. This data can then be used as pre-training data to teach the model to learn domain knowledge, thereby increasing the proportion of domain knowledge in the pre-training stage while ensuring data quality and improving the final training effect of the model.

[0103] In some embodiments, the third instruction text includes first task information identified by the first task and second task information identified by the second task; correspondingly, S603 may include S6031-S6032:

[0104] S6031: Based on the fine ranking model, the pre-trained data is used to perform first task identification by combining first task information and second task identification by combining second task information to obtain first indicator data and second indicator data.

[0105] S6032: By integrating the data from the first indicator and the second indicator, the ranking evaluation result is obtained.

[0106] Specifically, the first task identification is used to match the pre-training data with the target knowledge domain to determine whether it belongs to the target knowledge domain. The second task identification is used to match the pre-training data with the model to be trained to determine whether it is suitable for the domain task of the model to be trained in the target knowledge domain. The first task information is used to provide guidance information for the fine-ranking model to perform the first task identification, and the second task information is used to provide guidance information for the fine-ranking model to perform the second task identification.

[0107] Specifically, the first indicator data indicates the probability that the pre-training data belongs to the target knowledge domain, and the second indicator data indicates the probability that the pre-training data is suitable for the model to be trained, that is, the probability that it is suitable for the domain task of the model to be trained. Next, the two probabilities are fused, and the specific fusion method may include, but is not limited to, multiplication. The resulting refined ranking evaluation result is used to comprehensively evaluate the pre-training data. Based on the refined ranking evaluation result, the pre-training data in the intermediate dataset is re-ranked, and a certain number of the top-ranked pre-training data are selected to generate the target pre-training set, which is added to the pre-training dataset of the model to be trained to improve its domain knowledge content during the pre-training stage. In this way, the suitability of the pre-training data with the knowledge domain and model domain task is further evaluated in combination with different tasks to achieve refined ranking filtering and further improve data quality.

[0108] In some embodiments, the fine-ranking model uses a continuous scoring mechanism to perform the fine-ranking evaluation task, while using a stronger large-scale language model (such as 70B LLMs) to score the pre-training data. By flexibly adjusting the threshold, smaller and higher-quality pre-training data can be obtained.

[0109] Specifically, the third instruction text may include multiple tasks, which can be selected based on their suitability with the target knowledge domain and their suitability with the trained model. Combining the previous examples, the third instruction text can be exemplified as follows:

[0110] "Please determine whether the following webpage contains mathematical knowledge, and whether it is suitable for teaching language models to learn mathematical skills. For both tasks, answer only Yes / No."

[0111] Web page:"

[0112] In some embodiments, the first indicator data can characterize the probability likelihood of positive solutions among the preset number of solutions with the highest probability of the fine-ranking model's output for the first task. Here, a positive solution refers to a solution that represents the pre-training data belonging to the target knowledge domain. The second indicator data can characterize the probability likelihood of positive solutions among the preset number of solutions with the highest probability of the fine-ranking model's output for the second task. Here, a positive solution refers to a solution that represents the pre-training data that is adapted to the domain task of the model to be trained.

[0113] In one example, a VLLM sampling method is used to construct a fine-grained ranking model for fine-grained ranking evaluation inference. When decoding the token for each task, the model is required to provide the five words with the highest probabilities. If "Yes" or "No" appears among the five words with the highest probabilities, the likelihood of these two words being decoded is recorded as logp(Yes) and logp(No). If neither "Yes" nor "No" appears among the five words with the highest decoding probabilities, the default decoding likelihood is -100. Then, the probability of each task (e.g., first task identification and second task identification) being predicted as "Yes" is calculated using the following formula: p i =exp logpi(yes) / (exp logpi(yes) +exp logpi(No) ), p i The indicator data representing the i-th task, such as the first indicator data or the second indicator data, are used to obtain the quality score of the fine ranking evaluation. For example, the first indicator data p1 and the second indicator data p2 are multiplied together to obtain the fine ranking evaluation result p1*p2. Then, the data with lower p1*p2 are filtered out, and the high-scoring data is used as the final pre-training data output.

[0114] Based on some or all of the above embodiments, in the embodiments of this application, reference is made to Figure 6Following S209, the method also includes S701-S707:

[0115] S701: Obtain the data source to which each pre-training data in the target pre-training set belongs;

[0116] S703: Based on the ranking evaluation results of each pre-trained data, perform data source distribution statistics to obtain the data source evaluation results of each data source.

[0117] S705: Select the target data source from each data source corresponding to the target pre-training set based on the data source evaluation results;

[0118] S707: Obtain updated open-source pre-trained data from the target data source to serve as pre-trained data for recall.

[0119] Specifically, the data source refers to the source website of the pre-training data, such as the source website for publishing mathematical knowledge pages. The data source for each piece of pre-training data in the target pre-training set is determined. Then, the proportion of pre-training data in the target pre-training set and the ranking evaluation results of each data source are statistically analyzed to identify data sources with higher data proportions and / or higher data scores. These target data sources are then used as fixed data crawling sites to selectively acquire their updated open-source content, supplementing the initial dataset and further improving the quantity and efficiency of high-quality professional domain data acquisition in the pre-training data. Understandably, the evaluation criteria for higher data proportions and / or higher data scores can be set based on actual needs, such as setting proportion thresholds, score thresholds, or setting the extraction ratio for data source ranking.

[0120] Taking the field of mathematical logic reasoning as an example, based on the ranking evaluation results, the domain names of high-scoring web pages are statistically analyzed, and data from these mathematical vertical sites are crawled in a targeted manner. The newly added page data is then added to the general pre-training data pool to enrich the initial recall dataset and expand the recallable data resources.

[0121] The following combination Figure 7This application describes a process for pre-trained data mining in the field of mathematics, including: 1) maintaining a data candidate pool, obtaining an initial dataset through page data extraction and deduplication, and storing it in a maintained general pre-trained data pool database; 2) retrieving mathematics-related page data from the initial dataset in the general pre-trained data pool using the FastText model; 3) fine-tuning two first and second evaluation models using 1B-scale LLMs to evaluate the quality and mathematical relevance of the web page data, while filtering low-quality data to obtain an intermediate dataset; 4) performing fine-tuning and scoring on the initially filtered page data using 70B-scale LLMs, and sorting all data in the intermediate dataset; 5) filtering out low-scoring data, using high-scoring data as the final mathematical pre-training data output to obtain the target pre-training set, while simultaneously counting the domain names of high-scoring web pages and crawling these vertical sites to enrich the pre-training data pool. Thus, a progressive data mining strategy is employed to extract high-quality mathematical pre-training data from general web page data in a low-cost and efficient manner. Simultaneously, an asymptotic funnel strategy continuously improves the computational cost and filtering effectiveness of the data filtering model, maximizing the balance between resource consumption and model performance. Furthermore, a vertical website data supplementation mechanism continuously expands the mathematical pre-training data recall pool, ultimately obtaining a large amount of high-quality mathematical pre-training data, thus overcoming the limitations of existing solutions in terms of data scale and quality. Using this mathematical pre-training data to train LLMs can significantly improve the model's mathematical logic reasoning capabilities.

[0122] The technical solution of this application can be applied to large-scale model-related applications, such as large-scale model apps for mathematics teaching and tutoring, and has achieved significant performance improvements in mathematical logic reasoning.

[0123] This application embodiment also provides a pre-training data processing device 800, such as... Figure 8 As shown, Figure 8 The diagram shows a structural schematic of a pre-training data processing apparatus according to an embodiment of this application. The apparatus may include the following modules:

[0124] Acquisition Module 10: Used to acquire the initial recall dataset for the target knowledge domain. The initial recall dataset includes multiple pre-trained data retrieved for the target knowledge domain.

[0125] Quality evaluation module 20: used to predict the data quality of the pre-trained data of the initial dataset based on the first evaluation model, and obtain the quality evaluation results;

[0126] Relevance prediction module 30: It is used to predict the domain relevance of the pre-training data of the initial dataset based on the second evaluation model, and obtain the domain relevance prediction result. The relevance prediction result is used to indicate the relevance between the pre-training data and the target knowledge domain. The first evaluation model and the second evaluation model are constructed based on a large language model.

[0127] Data filtering module 40: Used to filter the initial dataset based on the quality assessment results and correlation prediction results to obtain an intermediate dataset;

[0128] Fine-ranking and filtering module 50: Used to fine-rank and filter the intermediate dataset to obtain the target pre-training set.

[0129] In some embodiments, the first evaluation model predicts the data quality of the pre-training data of the initial recruitment dataset based on discrete tiered evaluation; the second evaluation model predicts the domain relevance of the pre-training data of the initial recruitment dataset based on discrete tiered evaluation.

[0130] In some embodiments, the quality evaluation module 20 may be specifically used for:

[0131] Obtain the first instruction text, which is used to provide guidance information for the first evaluation model to perform data quality prediction, including multiple data quality evaluation items;

[0132] The first instruction text and the pre-trained data of the initial recall dataset are input into the first evaluation model to perform quality predictions for each of the multiple data quality evaluation items on the pre-trained data, and the obtained first prediction scores are fused to obtain the quality evaluation result.

[0133] In some embodiments, the correlation prediction module 30 may be specifically used for:

[0134] Obtain the second instruction text, which provides guidance information for the second evaluation model to perform correlation prediction, including multiple correlation evaluation items;

[0135] The second instruction text and the pre-trained data of the initial recall dataset are input into the second evaluation model to predict the relevance of each of the multiple relevance evaluation items on the pre-trained data, and the obtained second prediction scores are fused to obtain the relevance prediction result.

[0136] In some embodiments, the fine-sorting and screening module 50 may be specifically used for:

[0137] Obtain the third instruction text, which provides guidance for evaluating the performance of the fine-ranking model based on a large language model;

[0138] The pre-trained data of the third instruction text and the intermediate dataset are input into the fine ranking model for data evaluation to obtain the fine ranking evaluation result. The fine ranking evaluation result is used to indicate the fit between the pre-trained data and the target knowledge domain, as well as the model to be trained, which is a large language model of the target knowledge domain.

[0139] The intermediate dataset is filtered based on the ranking evaluation results to obtain the target pre-training set.

[0140] In some embodiments, the third instruction text includes first task information identified by the first task and second task information identified by the second task. The first task identification is used for the adaptation category between the pre-training data and the target knowledge domain, and the second task identification is used for the adaptation identification between the pre-training data and the model to be trained. The fine-tuning and filtering module 50 can be specifically used for:

[0141] Based on the fine ranking model, the pre-training data is used to perform first task identification by combining first task information and second task identification by combining second task information to obtain first indicator data and second indicator data. The first indicator data is used to indicate the probability that the pre-training data belongs to the target knowledge domain, and the second indicator data is used to indicate the probability that the pre-training data is adapted to the model to be trained.

[0142] By integrating the data from the first indicator and the second indicator, the refined ranking evaluation results are obtained.

[0143] In some embodiments, the apparatus further includes a data source filtering module for:

[0144] After refining and filtering the intermediate dataset to obtain the target pre-training set, the data source to which each pre-training data in the target pre-training set belongs is obtained;

[0145] Based on the ranking evaluation results of each pre-training data, the distribution statistics of the data sources are performed to obtain the data source evaluation results of each data source.

[0146] Based on the data source evaluation results, the target data source is selected from each data source corresponding to the target pre-training set;

[0147] Retrieve updated open-source pre-trained data from the target data source to serve as the pre-trained data to be recalled.

[0148] In some embodiments, the acquisition module is specifically used for:

[0149] Obtain the initial dataset;

[0150] The recall model is used to recall data in the target knowledge domain from the initial dataset to obtain the initial recall dataset. The recall model is obtained by constraining the initial recall model to identify pre-trained data in the target knowledge domain based on a preset ratio of positive training samples and negative training samples. The positive training samples are data that belong to the target knowledge domain and meet the data quality conditions, while the negative training samples are obtained by randomly sampling the initial dataset.

[0151] In some embodiments, the acquisition module may be further specifically used for:

[0152] Retrieve data from multiple open-source web pages;

[0153] The system performs deduplication on multiple open-source web page datasets and generates an initial dataset based on the deduplicated dataset.

[0154] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0155] This application provides a computer device including a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement a pre-training data processing method as provided in the above method embodiments.

[0156] Figure 9 A schematic diagram of the hardware structure of an apparatus for implementing a pre-training data processing method provided in an embodiment of this application is shown. The apparatus may constitute or include the device or system provided in the embodiment of this application. Figure 9 As shown, device 10 may include one or more processors 1002 (shown as 1002a, 1002b, ..., 1002n in the figure) 1002 (processor 1002 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 9The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, device 10 may also include a... Figure 9 The more or fewer components shown, or having the same Figure 9 The different configurations shown.

[0157] It should be noted that the aforementioned one or more processors 1002 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be wholly or partially integrated into any other element within device 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0158] The memory 1004 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in this embodiment. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, thereby implementing the aforementioned method for processing pre-trained data. The memory 1004 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1004 may further include memory remotely located relative to the processor 1002, and these remote memories can be connected to the device 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0159] The transmission device 1006 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of device 10. In one example, the transmission device 1006 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 1006 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0160] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows a user to interact with the user interface of device 10 (or a mobile device).

[0161] This application also provides a computer-readable storage medium, which can be disposed in a server to store at least one instruction or at least one program related to implementing a pre-training data processing method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the pre-training data processing method provided in the above method embodiment.

[0162] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0163] This invention also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a pre-training data processing method provided in the various optional embodiments described above.

[0164] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0165] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device, equipment, and storage medium embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0166] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0167] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for processing pre-training data, characterized in that, The method includes: Obtain an initial recall dataset for the target knowledge domain, wherein the initial recall dataset includes multiple pre-trained data recalled for the target knowledge domain; Based on the first evaluation model, the pre-trained data of the initial recruitment dataset is used to predict data quality and obtain quality evaluation results. Based on the second evaluation model, the pre-training data of the initial recall dataset is used to predict the domain relevance, and the domain relevance prediction result is obtained. The relevance prediction result is used to indicate the relevance between the pre-training data and the target knowledge domain. The first evaluation model and the second evaluation model are constructed based on a large language model. Based on the quality evaluation results and the correlation prediction results, the initial dataset is filtered to obtain an intermediate dataset. The intermediate dataset is then sorted and filtered to obtain the target pre-training set.

2. The method according to claim 1, characterized in that, The first evaluation model predicts the data quality of the pre-trained data of the initial recruitment dataset based on discrete tiered evaluation. The second evaluation model predicts the domain relevance of the pre-trained data of the initial recruitment dataset based on discrete tiered evaluation.

3. The method according to claim 1, characterized in that, The data quality prediction based on the pre-training data of the initial recruitment dataset using the first evaluation model, and the resulting quality evaluation results, include: Obtain a first instruction text, which is used to provide guidance information required by the first evaluation model to perform the data quality prediction, including multiple data quality evaluation items; The first instruction text and the pre-trained data of the initial recall dataset are input into the first evaluation model to perform quality predictions on the pre-trained data for each of the multiple data quality evaluation items, and the obtained first prediction scores are fused to obtain the quality evaluation result.

4. The method according to claim 1, characterized in that, The domain relevance prediction results obtained by performing domain relevance prediction on the pre-trained data of the initial recruitment dataset based on the second evaluation model include: Obtain a second instruction text, which provides guidance information required by the second evaluation model to perform the correlation prediction, including multiple correlation evaluation items; The second instruction text and the pre-trained data of the initial recall dataset are input into the second evaluation model to predict the relevance of each of the multiple relevance evaluation items on the pre-trained data, and the obtained second prediction scores are fused to obtain the relevance prediction result.

5. The method according to any one of claims 1-4, characterized in that, The process of refining and filtering the intermediate dataset to obtain the target pre-training set includes: Obtain a third instruction text, which is used to provide guidance information for evaluating the performance of the fine-ranking model, the fine-ranking model being built based on a large language model; The third instruction text and the pre-trained data of the intermediate dataset are input into the fine ranking model for data evaluation to obtain fine ranking evaluation results. The fine ranking evaluation results are used to indicate the suitability of the pre-trained data with the target knowledge domain and the model to be trained. The model to be trained is a large language model of the target knowledge domain. The intermediate dataset is filtered based on the ranking evaluation results to obtain the target pre-training set.

6. The method according to claim 5, characterized in that, The third instruction text includes first task information of the first task identification and second task information of the second task identification. The first task identification is used for the adaptation category of the pre-training data and the target knowledge domain, and the second task identification is used for the adaptation identification of the pre-training data and the model to be trained. The step of inputting the third instruction text and the pre-trained data of the intermediate dataset into the fine-ranking model for data evaluation, and obtaining the fine-ranking evaluation result includes: Based on the fine ranking model, the pre-training data is used to perform first task identification combined with the first task information and second task identification combined with the second task information to obtain first indicator data and second indicator data. The first indicator data is used to indicate the probability that the pre-training data belongs to the target knowledge domain, and the second indicator data is used to indicate the probability that the pre-training data is adapted to the model to be trained. The ranking evaluation result is obtained by integrating the first indicator data and the second indicator data.

7. The method according to claim 5, characterized in that, After performing fine sorting and filtering on the intermediate dataset to obtain the target pre-training set, the method further includes: Obtain the data source to which each pre-training data in the target pre-training set belongs; Based on the ranking evaluation results of each of the pre-trained data, the distribution statistics of the data sources are performed to obtain the data source evaluation results of each of the data sources. Based on the evaluation results of the data sources, target data sources are selected from each data source corresponding to the target pre-training set; Updated open-source pre-trained data is obtained from the target data source to serve as pre-trained data for recall.

8. The method according to any one of claims 1-4, characterized in that, The initial dataset for obtaining the target knowledge domain includes: Obtain the initial dataset; The initial recall dataset is obtained by recalling the data to be recalled in the target knowledge domain based on the recall model. The recall model is obtained by constraining the initial recall model to identify the pre-trained data in the target knowledge domain based on a preset ratio of positive training samples and negative training samples. The positive training samples are data that belong to the target knowledge domain and meet the data quality conditions, and the negative training samples are obtained by randomly sampling the initial dataset.

9. The method according to claim 7, characterized in that, The process of obtaining the initial dataset includes: Retrieve data from multiple open-source web pages; The multiple open-source web page data are deduplicated to remove duplicate data, and the initial dataset is generated based on the deduplicated open-source web page data.

10. A pre-training data processing apparatus, characterized in that, The device includes: Acquisition module: used to acquire the initial recall dataset for the target knowledge domain, the initial recall dataset including multiple pre-trained data recalled for the target knowledge domain; Quality evaluation module: used to predict the data quality of the pre-trained data of the initial dataset based on the first evaluation model, and obtain the quality evaluation result; The relevance prediction module is used to predict the domain relevance of the pre-training data of the initial dataset based on the second evaluation model, and obtain the domain relevance prediction result. The relevance prediction result is used to indicate the relevance of the pre-training data to the target knowledge domain. The first evaluation model and the second evaluation model are constructed based on a large language model. Data filtering module: used to filter the initial dataset based on the quality evaluation results and the correlation prediction results to obtain an intermediate dataset; Fine-ranking and filtering module: used to fine-rank and filter the intermediate dataset to obtain the target pre-training set.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the method for processing pre-trained data as described in any one of claims 1 to 9.

12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or at least one program being loaded and executed by the processor to implement the method for processing pre-trained data as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the method for processing pre-trained data as described in any one of claims 1 to 9.