Large model pre-training data processing method and device, electronic equipment and storage medium
By identifying and optimizing the weak content in the pre-training data of the large model, and using the baseline and enhanced answer difference judgment of the target large model, the quality problem of pre-training data is solved and the performance of the large model is improved.
Patent Information
- Application Number
- CN202510667835.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, quality issues with pre-training data can lead to performance degradation in large models, and it is necessary to accurately identify and optimize problematic data to improve the performance of large models.
By identifying the weak content in the pre-training data of the target large model, an exploration query is generated. The difference between the output baseline and enhanced answers of the target large model is used to determine whether the weak content is problematic data, and corresponding optimization is performed.
Accurately identify and optimize question data in pre-training data, improving the answer accuracy and performance of large models.
Smart Images

Figure CN120804692A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large model, and particularly relates to a large model pre-training data processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] The quality of pre-training data used for training a large model has an important influence on the performance of the large model, and if there is problem data in the pre-training data, the performance of the large model will be reduced. Therefore, in order to improve the performance of the large model, it is necessary to accurately identify the problem data in the pre-training data so as to optimize the problem data subsequently. SUMMARY
[0003] The present application provides a large model pre-training data processing method and device, electronic equipment and storage medium to solve the technical problem of how to identify problem data in large model pre-training data.
[0004] The present application provides a large model pre-training data processing method, comprising:
[0005] Identifying weak content in pre-training data of a target large model;
[0006] Generating a probe query corresponding to the weak content;
[0007] Inputting the probe query into the target large model to obtain a baseline answer output by the target large model;
[0008] Retrieving a first data segment most relevant to the probe query in the pre-training data;
[0009] Inputting the probe query and the first data segment into the target large model to obtain an enhanced answer output by the target large model;
[0010] Judging whether the weak content is problem data according to the difference between the enhanced answer and the baseline answer.
[0011] According to the large model pre-training data processing method provided by the present application, the weak content in the pre-training data of the target large model is identified, comprising:
[0012] Clustering the pre-training data to obtain an outlying cluster in a clustering result;
[0013] Determining the perplexity of a preset language large model on the pre-training data to obtain a second data segment in the pre-training data with the perplexity greater than a first preset threshold;
[0014] Taking the intersection of the outlying cluster and the second data segment to obtain the weak content.
[0015] The method comprises the following steps: clustering the pre-training data according to the target large model, and determining a clustering theme keyword according to a performance defect of the target large model.
[0016] The method comprises the following steps: clustering the pre-training data according to the target large model, and determining a clustering theme keyword according to a performance defect of the target large model.
[0017] The method comprises the following steps: clustering the pre-training data according to the target large model, and determining a clustering theme keyword according to a performance defect of the target large model.
[0018] The method comprises the following steps: clustering the pre-training data according to the target large model, and determining a clustering theme keyword according to a performance defect of the target large model.
[0019] If the difference is greater than a second preset threshold corresponding to the difference, the weak content is determined as problem data, otherwise the weak content is determined as non-problem data.
[0020] The difference comprises ROUGE-L and / or BERTScore.
[0021] The method comprises the following steps: clustering the pre-training data according to the target large model, and determining a clustering theme keyword according to a performance defect of the target large model.
[0022] If the weak content is problem data, the pre-training data is optimized according to the weak content.
[0023] The method comprises the following steps: clustering the pre-training data according to the target large model, and determining a clustering theme keyword according to a performance defect of the target large model.
[0024] If the weak content comprises random codes, a regular expression library is established to filter irregular character combinations in the pre-training data.
[0025] If the weak content lacks the latest guidelines related to target content, the latest guidelines are added to the pre-training data.
[0026] If the weak content comprises table data, a tool is deployed to extract the table data when the target large model is trained using the pre-training data.
[0027] The application further provides a large model pre-training data processing device, comprising:
[0028] The recognition module is configured to recognize weak content in pre-training data of a target large model.
[0029] The generation module is configured to generate an exploration query corresponding to the weak content.
[0030] an input module configured to input the probe query into the target large model to obtain a baseline answer output by the target large model;
[0031] a retrieval module configured to retrieve a first data segment most relevant to the probe query from the pre-training data;
[0032] The input module is further configured to input the probe query and the first data segment into the target large model to obtain an enhanced answer output by the target large model;
[0033] a judgment module configured to judge whether the weak content is problem data according to a difference between the enhanced answer and the baseline answer.
[0034] The present application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the large model pre-training data processing method according to any one of the above when executing the program.
[0035] The present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the large model pre-training data processing method according to any one of the above.
[0036] The present application also provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the large model pre-training data processing method according to any one of the above.
[0037] The large model pre-training data processing method, device, electronic device, and storage medium provided by the present application can accurately identify problem data in large model pre-training data by initially identifying weak content in large model pre-training data, generating a probe query corresponding to the weak content, and judging whether the weak content is problem data according to a difference between an enhanced answer and a baseline answer of the probe query by a target large model. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0039] Figure 1 is a flowchart of the large model pre-training data processing method provided by the present application.
[0040] Figure 2 is a flowchart of step S1 provided by the present application.
[0041] Figure 3 is a schematic diagram of target large model baseline answer and enhanced answer accuracy provided by the present application;
[0042] Figure 4 is a structural schematic diagram of a large model pre-training data processing device provided by the present application.
[0043] Figure 5 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.
[0045] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement “comprises a” does not exclude the presence of another identical element in the process, method, article or device comprising the element. The terms “upper”, “lower” and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. Unless otherwise explicitly specified and limited, the terms “mount”, “connect”, “connect” should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0046] The terms "first", "second", and the like in the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in a "or" relationship.
[0047] The following will be described in conjunction with Figures 1-5 The large model pre-training data processing method, device, electronic equipment and storage medium provided by the present disclosure are described.
[0048] As Figure 1 The large model pre-training data processing method provided by the present disclosure includes steps S1-S6.
[0049] Step S1, identify weak content in the pre-training data of the target large model.
[0050] The target large model can be a large model in each field, for example, a large model in the medical field, which can answer questions related to diseases raised by medical staff.
[0051] The weak content in the pre-training data is data that is likely to have problems, that is, data that will reduce the accuracy of the target large model in answering questions, including structurally abnormal, sparse or scattered data clusters. The weak content can be determined manually according to experience, or can be identified according to an algorithm. At this time, only weak content with a high probability of having problems can be obtained, but the weak content is not necessarily problem data and needs to be further verified.
[0052] Step S2, generate a probe query corresponding to the weak content.
[0053] The probe query is a question input to the target large model subsequently. For example, if the weak content is related to calculating the dose of a drug, the probe query can be a question related to calculating the dose of the drug; if the weak content is related to a rare disease, the probe query can be a question related to the rare disease; if the weak content includes multi-modal data, the probe query can be a question related to multi-modal data processing.
[0054] Step S3, input the probe query to the target large model to obtain a baseline answer output by the target large model.
[0055] The baseline answer of the target large model is an answer obtained by inputting the probe query into the target large model. It can be understood that, due to the use of weak content in training the target large model, if the weak content is problem data, the accuracy of the target large model in answering questions related to the weak content will be relatively low; if the weak content is non-problem data, the accuracy of the target large model in answering questions related to the weak content will be relatively high.
[0056] Step S4, retrieving the first data segment most relevant to the probe query in the pre-training data.
[0057] Specifically, a sentence embedding model such as sentence-transformers / all-mpnet-base-v2 in the Hugging Face model library can be used to construct a sentence embedding index of the pre-training data, so that texts with similar semantics in the pre-training data are close in the vector space, and texts with different semantics are far apart, and then the semantic retrieval of the probe query is performed to obtain the first data segment.
[0058] Step S5, inputting the probe query and the first data segment into the target large model to obtain an enhanced answer output by the target large model.
[0059] The enhanced answer of the target large model is an answer obtained by inputting the probe query and the first data segment into the target large model. The enhanced answer is obtained based on the RAG (Retrieval-augmented Generation) technology. The accuracy of the enhanced answer of the large model is generally higher than that of the baseline answer.
[0060] Step S6, judging whether the weak content is problem data according to the difference between the enhanced answer and the baseline answer.
[0061] It can be understood that, if the weak content is problem data, the accuracy of the baseline answer of the target large model will be low; and the first data segment is the data most relevant to the weak content, which is high-quality data. Inputting the first data segment into the target large model can greatly improve the accuracy of the target large model in answering questions, that is, the difference between the enhanced answer and the baseline answer will be large. If the weak content is non-problem data, the accuracy of the baseline answer of the target large model will be high, and even if the first data segment is input into the target large model, the improvement of the accuracy of the target large model in answering questions will be relatively limited, that is, the difference between the enhanced answer and the baseline answer will be small.
[0062] Therefore, if the difference between the enhanced answer and the baseline answer is large, it can be considered that the weak content is problem data; if the difference between the enhanced answer and the baseline answer is small, it can be considered that the weak content is non-problem data.
[0063] From the above, the large model pre-training data processing method of the application first identifies weak content in the large model pre-training data, generates a probe query corresponding to the weak content, and then determines whether the weak content is problem data according to the difference between the enhanced answer and the baseline answer of the probe query of the target large model. The weak content can accurately identify the problem data in the large model pre-training data.
[0064] In order to accurately obtain the weak content in the pre-training data, in some embodiments, as shown in the figure, Figure 2 As shown in the figure, step S1 of the application can further include:
[0065] Step S11, clustering the pre-training data to obtain an outlying cluster in the clustering result;
[0066] Step S12, determining the perplexity of the pre-set language large model on the pre-training data to obtain a second data segment in the pre-training data with a perplexity greater than a first preset threshold;
[0067] Step S13, taking the intersection of the outlying cluster and the second data segment to obtain the weak content.
[0068] The outlying cluster in the clustering result includes a data cluster with abnormal structure, sparsity or dispersion. Training the target large model using the outlying cluster will reduce the accuracy of the target large model in answering questions. The pre-training data can be clustered according to clustering theme keywords, such as the dose of a certain drug or the name of a certain rare disease.
[0069] Perplexity (PPL) is used to quantify the uncertainty of the language large model on the text. The higher the perplexity of the pre-training data, the poorer the quality of the pre-training data. The pre-set language large model can be a trained GPT-3.5 model, and the first preset threshold can be 150.
[0070] Since the outlying cluster is poor quality data, the second data segment is also poor quality data, and the intersection of the two has a high probability of being problem data.
[0071] In this way, the weak content is identified from the dual perspectives of clustering and calculating perplexity, which can accurately identify the weak content in the pre-training data and improve the accuracy of identifying problem data.
[0072] Considering that the weak content is data that will most likely reduce the accuracy of the target large model in answering, if the weak content is deduced according to the performance defects of the target large model, the reliability of the weak content will be higher.
[0073] Therefore, in step S11, clustering the pre-training data can further include:
[0074] determine the clustering topic keyword according to the performance defect of the target large model;
[0075] cluster the pre-training data according to the clustering topic keyword.
[0076] For example, the performance defect of the target large model can include that the accuracy of the answer to a certain rare disease is too low, and the clustering topic keyword can include the name of the rare disease; the performance defect of the target large model can include that the error rate of calculating the dose of a certain drug is too high, and the clustering topic keyword can include the name of the drug dose.
[0077] In this way, the clustering topic keyword is determined according to the performance defect of the target large model, and the pre-training data is clustered according to the clustering topic keyword, which is equivalent to deducing the weak content according to the performance defect of the target large model. The reliability of the weak content can be improved, and the accuracy of identifying the problem data can be improved.
[0078] In some embodiments, step S6 can further include:
[0079] If the difference between the enhanced answer and the baseline answer of the target large model is greater than the second preset threshold corresponding to the difference, the weak content is determined as problem data, otherwise the weak content is determined as non-problem data.
[0080] If the difference between the enhanced answer and the baseline answer of the target large model is greater than the second preset threshold corresponding to the difference, the weak content is determined as problem data, otherwise the weak content is determined as non-problem data.
[0081] In this way, whether the weak content is problem data can be determined according to the size relationship between the difference between the enhanced answer and the baseline answer and the second preset threshold.
[0082] In some embodiments, the difference between the enhanced answer and the baseline answer of the target large model can include ROUGE-L and / or BERTScore. ROUGE-L and BERTScore are two commonly used automatic evaluation indicators, mainly used to measure the similarity between generated text (baseline answer) and reference text (enhanced answer). ROUGE-L is used to measure the longest matching word sequence between generated text and reference text, focusing on recall, that is, how much information in the reference text is covered by the generated text. BERTScore is a deep semantic representation using a pre-trained language model (such as BERT) to calculate the similarity between generated text and reference text in the vector space.
[0083] If the difference includes ROUGE-L, the second preset threshold corresponding to ROUGE-L can be 0.3; if the difference includes BERTScore, the second preset threshold corresponding to BERTScore can be 0.15.
[0084] In some embodiments, after step S6, the large model pre-training data processing method of the present application can further include:
[0085] If the weak content is problem data, the pre-training data is optimized according to the weak content.
[0086] After optimizing the pre-training data according to the weak content, subsequent training of the target large model using the optimized pre-training data can target the performance defects of the target large model for targeted optimization, thereby improving the performance of the target large model.
[0087] In some embodiments, optimizing the pre-training data according to the weak content can include at least one of the following:
[0088] If the weak content includes garbled code, a regular expression library is established to filter irregular character combinations in the pre-training data;
[0089] If the weak content lacks the latest guidelines related to the target content, the latest guidelines are added to the pre-training data;
[0090] If the weak content includes table data, a tool is deployed to extract the table data when training the target large model using the pre-training data.
[0091] Establishing a regular expression library to filter irregular character combinations in the pre-training data can avoid garbled code in the pre-training data, improve the quality of the pre-training data, and help improve the performance of the target large model.
[0092] If the weak content lacks the latest guidelines related to the target content, the latest guidelines are added to the pre-training data, and subsequent training of the target large model using the pre-training data can improve the accuracy of the target large model in answering questions related to the target content. For example, the target content can be a rare disease, and the latest guidelines related to the rare disease are added to the pre-training data, and subsequent training of the target large model using the pre-training data can improve the accuracy of the target large model in answering questions related to the rare disease.
[0093] Weak content includes tabular data, which can be considered as the target large model cannot correctly process tabular data. When training the target large model using pre-training data, deploying tools can extract tabular data, which can make the target large model correctly process tabular data, and thus improve the performance of the target large model. For example, the deployed tool can be TabNet, which is a deep learning architecture specially designed for structured tabular data (Tabular Data), which can accurately extract tabular data in pre-training data.
[0094] The target large model is a certain medical field large model in the present application:
[0095] In the clinical question and answer task of the medical field large model, the performance defects of the medical field large model include: the question and answer accuracy of rare diseases (such as Fabry disease) is less than 40%; the logic is confused when processing multi-modal medical reports (text + table); the drug dosage (such as vancomycin dosage) calculation error rate is abnormally high;
[0096] Based on the performance defects of the medical field large model, the clustering topic keywords "Fabry disease" and "vancomycin dosage" are obtained;
[0097] First, 10% of the 5TB pre-training data is extracted for stratified sampling (30% of medical journals, 45% of electronic medical records, and 25% of popular science articles), and a Sentence-BERT model is used to generate 768-dimensional text embedding of the extracted pre-training data;
[0098] Then, the HDBSCAN algorithm is used for hierarchical clustering based on the previously determined clustering topic keywords, with the following parameter settings: min_cluster_size=50, min_samples=3, 17 outlier clusters (32,000 data) are found in the clustering results, including:
[0099] No. 12 outlier cluster: contains a large number of scan version PDF conversion garbled codes (such as "Di abetes Mell!tus Typ€2");
[0100] No. 15 outlier cluster: mixed English, French and German cross-language case reports (accounting for <0.3%);
[0101] No. 23 outlier cluster: radiology reports with mixed tables and texts (CT / MRI structure abnormalities);
[0102] Then the perplexity PPL of the extracted pre-training data is calculated using the trained GPT-3.5 model, and the second data segment with PPL>150 (4.7% of the total data) is obtained, and typical high PPL samples are, for example: "The standard dose of adult intravenous vancomycin is 15-20 mg / kg q8-12h (adjusted according to creatinine clearance)", PPL=287;
[0103] Take the intersection of outliers and the second data segment (0.8% of the total) as weak data;
[0104] Then based on the clustering topic keywords (such as Fabry disease, vancomycin dose) and the medical term library, the probe query of weak data is automatically generated, the query type of the probe query includes drug dose, rare disease and multi-modal processing, and 20 trap questions designed by clinical physicians (such as whether ibuprofen and warfarin can be used simultaneously?) are manually supplemented;
[0105] Use Faiss to build IVF65536_HNSW32 index of pre-training data, retrieve the most relevant data segments to the probe query, and take the top 5 data segments in the retrieval results as the first data segment, filter the retrieval results with cosine similarity <0.7;
[0106] Get the baseline answer accuracy, enhanced answer accuracy of the probe query corresponding to the drug dose, rare disease and multi-modal processing as shown in Figure 3
[0107] Audit the weak content and find that the weak content has the following problems:
[0108] There is data noise: 23% of the data in No. 12 outlier cluster has OCR (Optical Character Recognition) recognition error;
[0109] Information is sparse: there are only 127 data related to Fabry disease, and there is no latest treatment guideline for Fabry disease;
[0110] There are structural defects: 87% of the radiology reports do not correctly parse the table structure.
[0111] As can be seen, the main problem of pre-training data is the chaotic format of original data.
[0112] Based on the problems existing in the weak content, the following optimization measures are proposed:
[0113] De-noising of pre-training data: establish regular expression library to filter irregular character combinations;
[0114] Enhance the pre-training data: supplement the latest guidelines of the clinical decision system;
[0115] Structural adjustment is made to the pre-training data: deploy TabNet for table data extraction.
[0116] After training the medical field large model using the optimized pre-training data, the accuracy of the medical field large model in calculating drug dosage is improved by 31.2pp, the F1-score for answering rare diseases is increased from 0.47 to 0.69, and the error rate for processing multi-modal data is reduced by 58%.
[0117] As shown in Figure 4 The large model pre-training data processing device provided by the application comprises:
[0118] The identification module is configured to identify weak content in pre-training data of a target large model;
[0119] The generation module is configured to generate a probe query corresponding to the weak content;
[0120] The input module is configured to input the probe query into the target large model to obtain a baseline answer output by the target large model;
[0121] The retrieval module is configured to retrieve a first data segment most relevant to the probe query from the pre-training data;
[0122] The input module is further configured to input the probe query and the first data segment into the target large model to obtain an enhanced answer output by the target large model;
[0123] The judgment module is configured to determine whether the weak content is problem data according to the difference between the enhanced answer and the baseline answer.
[0124] It should be noted that the large model pre-training data processing device provided by the application can execute the large model pre-training data processing method of any of the above embodiments when it is actually run, and therefore this embodiment will not be described here.
[0125] In some embodiments, the identification module can be further configured to:
[0126] Cluster the pre-training data to obtain an outlying cluster in the clustering result;
[0127] Determine the perplexity of the pre-set language large model on the pre-training data to obtain a second data segment in the pre-training data with a perplexity greater than a first preset threshold;
[0128] Take the intersection of the outlying cluster and the second data segment to obtain the weak content.
[0129] In some embodiments, the identification module can be further configured to:
[0130] Determine a clustering theme keyword according to the performance defects of the target large model;
[0131] The pre-training data is clustered according to the clustered topic keywords.
[0132] In some embodiments, the determining module can be further configured to:
[0133] If the difference is greater than a second preset threshold corresponding to the difference, the weak content is determined as problem data, otherwise the weak content is determined as non-problem data.
[0134] In some embodiments, the difference between the enhanced answer and the baseline answer can include ROUGE-L and / or BERTScore.
[0135] In some embodiments, the large model pre-training data processing apparatus can further include:
[0136] The optimization module is configured to optimize the pre-training data according to the weak content if the weak content is problem data.
[0137] In some embodiments, the optimization module can be further configured to:
[0138] If the weak content includes garbled code, a regular expression library is established to filter irregular character combinations in the pre-training data.
[0139] If the weak content lacks the latest guidelines related to the target content, the latest guidelines are added to the pre-training data.
[0140] If the weak content includes table data, a tool is deployed to extract the table data when the target large model is trained using the pre-training data.
[0141] Figure 5 is a structural schematic diagram of an electronic device provided by the present application, as Figure 5 shown, the electronic device can include a processor, a communications interface, a memory and a communications bus, wherein the processor, the communications interface and the memory complete mutual communication through the communications bus. The processor can call logical instructions in the memory to execute a large model pre-training data processing method.
[0142] In addition, the logic instructions in the above-mentioned memory can be realized in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0143] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer can execute the large model pre-training data processing method provided by the above-mentioned embodiments.
[0144] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the large model pre-training data processing method provided by the above-mentioned embodiments.
[0145] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.
[0146] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions essentially or the parts that contribute to the prior art can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.
[0147] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features therein can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A large model pre-training data processing method, characterized in that: include: Identify weak points in the pre-training data of the target large model; generating a probing query corresponding to the weak content; Inputting the exploration query into the target macro model to obtain a baseline answer output by the target macro model; Retrieving a first data segment most relevant to the probe query from the pre-training data; Inputting the exploration query and the first data segment into the target macro model to obtain an enhanced answer output by the target macro model; Whether the weak content is problematic data is determined based on the difference between the enhanced answer and the baseline answer.
2. The large model pre-training data processing method according to claim 1, characterized in that: The weaknesses in the pre-training data of the target large model include: Clustering the pre-training data to obtain outlier clusters in the clustering results; Determining the perplexity of the pre-trained data by a preset large language model, and obtaining a second data segment in the pre-trained data, wherein the perplexity is greater than a first preset threshold; An intersection of the outlier cluster and the second data segment is obtained to obtain the weak content.
3. The large model pre-training data processing method according to claim 2, characterized in that: Clustering the pre-training data includes: Determining clustering topic keywords based on performance defects of the target large model; The pre-training data is clustered according to the clustering theme keywords.
4. The large model pre-training data processing method according to claim 1, characterized in that: The determining whether the weak content is problematic data based on the difference between the enhanced answer and the baseline answer includes: If the difference is greater than a second preset threshold corresponding to the difference, the weak content is determined to be problematic data; otherwise, the weak content is determined to be non-problematic data.
5. The large model pre-training data processing method according to claim 4, characterized in that: The variability includes ROUGE-L and / or BERTScore.
6. The large model pre-training data processing method according to claim 1, characterized in that: After determining whether the weak content is problematic data based on the difference between the enhanced answer and the baseline answer, the method further includes: If the weak content is problematic data, the pre-training data is optimized according to the weak content.
7. The large model pre-training data processing method according to claim 6, characterized in that: The optimizing the pre-training data according to the weak content includes at least one of the following: If the weak content includes garbled characters, a regular expression library is established to filter out irregular character combinations in the pre-trained data; If the weak content lacks the latest guideline related to the target content, adding the latest guideline to the pre-training data; If the weak content includes tabular data, the deployment tool extracts the tabular data when training the target large model using the pre-training data.
8. A large model pre-training data processing device, characterized in that: include: The recognition module is used to identify weak content in the pre-training data of the target large model; A generation module, configured to generate a search query corresponding to the weak content; An input module, configured to input the exploration query into the target macro model and obtain a baseline answer output by the target macro model; a retrieval module, configured to retrieve a first data segment most relevant to the probe query from the pre-training data; The input module is further configured to input the exploration query and the first data segment into the target macro model to obtain an enhanced answer output by the target macro model; A judgment module is used to judge whether the weak content is problematic data based on the difference between the enhanced answer and the baseline answer.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the large model pre-training data processing method as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the large model pre-training data processing method as described in any one of claims 1 to 7 is implemented.