Method and apparatus for acquiring training samples and training large-scale model optimization.

By identifying and constructing training samples for queries that large-scale models struggle with, the method optimizes training to improve reasoning capabilities and accuracy in complex tasks, addressing inefficiencies in current models.

JP7844786B2Active Publication Date: 2026-04-14BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current large-scale models face challenges in accurately processing certain queries, leading to inefficiencies in reasoning tasks, particularly in complex scenarios, necessitating improved training methods to enhance their inference capabilities.

Method used

A method and apparatus for obtaining training samples by identifying queries that large-scale models cannot process correctly, constructing corresponding training samples, and performing optimization training to address these weaknesses, utilizing a data-driven approach that includes query mining, problem discovery, and sample construction.

Benefits of technology

The solution enhances the reasoning ability of large-scale models by accurately identifying and addressing their weaknesses, improving the quality of training samples, and optimizing the training process to enhance accuracy and efficiency in handling complex queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007844786000001
    Figure 0007844786000001
  • Figure 0007844786000002
    Figure 0007844786000002
  • Figure 0007844786000003
    Figure 0007844786000003
Patent Text Reader

Abstract

To provide a training sample acquisition and large scale model optimization training method and a device which improve a reasoning ability of a large scale model, etc.SOLUTION: A training sample acquisition method used for performing optimization training to a large-scale model comprises: a step 101 of using queries collected from a predetermined data source used as input of the large-scale model as candidate queries in response to the determination that the queries agree with an optimization trigger condition; a step 102 of sorting target queries which cannot be correctly processed by the large-scale model from the candidate queries; and a step 103 of constructing corresponding training samples, respectively on the basis of the respective target queries.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , ,

[0001] This disclosure relates to the field of artificial intelligence technology, and particularly to methods and apparatuses for obtaining training samples and optimizing training of large-scale models in fields such as large-scale models, deep learning, and natural language processing.

Background Art

[0002] A large-scale model is a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of natural language text. The emergence of large-scale models has fundamentally changed the way of human-machine interaction and can reconstruct the entire computing ecosystem.

Summary of the Invention

Problems to be Solved by the Invention

[0003] This disclosure provides a method and an apparatus for obtaining training samples and optimizing training of a large-scale model.

Means for Solving the Problems

[0004] The method for obtaining training samples includes: responding to being determined to meet the optimization trigger condition, taking a query collected from a predetermined data source that can be used as an input to the large-scale model as a candidate query; selecting a target query from the candidate queries, where the target query is a query that the large-scale model cannot correctly process; constructing corresponding training samples based on each target query, where the training samples are used for performing optimization training on the large-scale model.

[0005] The method for optimizing training of a large-scale model is: A step of obtaining training samples, wherein the training samples are corresponding training samples constructed based on each target query, the target queries are queries that a large-scale model selected from candidate queries cannot process correctly, and the candidate queries are queries collected from a predetermined data source that can be used as input to the large-scale model. The step includes performing optimization training on the large-scale model using the training samples.

[0006] The training sample acquisition device includes a query mining module, a problem discovery module, and a sample construction module. The query mining module is used to select candidate queries collected from a predetermined data source that can be used as input to a large-scale model, in response to a determination that the optimization trigger conditions are met. The problem discovery module is used to select a target query from the candidate queries, and the target query is a query that the large-scale model cannot process correctly. The sample building module is used to build corresponding training samples based on each target query, and these training samples are used to perform optimization training on the large-scale model.

[0007] The large-scale model optimization training system includes a sample acquisition module and a model training module. The sample acquisition module is used to acquire training samples, the training samples being corresponding training samples constructed based on each target query, the target queries being queries that the large-scale model cannot process correctly, selected from candidate queries, and the candidate queries being queries collected from predetermined data sources that can be used as input to the large-scale model. The model training module is used to perform optimization training on the large-scale model using the training samples.

[0008] Electronic devices are At least one processor, The memory includes, which is communicated with at least one processor, The memory stores instructions that can be executed by the at least one processor, and when an instruction is executed by the at least one processor, the at least one processor executes the method.

[0009] A non-temporary, computer-readable storage medium on which computer instructions are stored allows the computer to execute the method described above.

[0010] A computer program product includes a computer program / instruction, which, when executed by a processor, realizes the method described above.

[0011] It should be understood that the information described herein is not intended to identify any key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure can be readily understood through the following specification. [Brief explanation of the drawing]

[0012] The drawings are provided to better understand this application and do not limit it. [Figure 1] This is a flowchart of an embodiment of the training sample acquisition method described herein. [Figure 2] This is a flowchart of an embodiment of the large-scale model optimization training method described herein. [Figure 3] This is a schematic diagram of the overall implementation process of the large-scale model optimization training method described herein. [Figure 4]It is a schematic structural diagram of the configuration of Example 400 of the training sample acquisition device of the present disclosure. [Figure 5] It is a schematic structural diagram of the configuration of Example 500 of the large-scale model optimization training device of the present disclosure. [Figure 6] A schematic block diagram of an electronic device 600 for implementing an embodiment of the present disclosure is shown.

Mode for Carrying Out the Invention

[0013] Hereinafter, exemplary embodiments of the present application will be described based on the drawings. For ease of understanding, various details of the embodiments of the present application are included, and they should be regarded as merely illustrative. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] In addition, the term "and / or" in this specification only explains the relevant relationship of the relevant object, indicating that three types of relationships are possible. For example, A and / or B can represent three cases: only A exists, A and B exist simultaneously, or only B exists. It should be understood that the symbol " / " generally represents that the relevant objects before and after are in an "or" relationship.

[0015] FIG. 1 is a flowchart of an embodiment of the training sample acquisition method of the present disclosure. As shown in FIG. 1, it includes the following specific implementation manners.

[0016] In step 101, in response to being determined to meet the optimization trigger condition, a query collected from a predetermined data source that can be used as an input to the large-scale model is used as a candidate query.

[0017] In step 102, target queries are selected from candidate queries, and the target queries are queries that cannot be correctly processed by the large-scale model.

[0018] In step 103, corresponding training samples are respectively constructed based on each target query, and the training samples are used for performing optimization training on the large-scale model.

[0019] Using the solutions described in the embodiments of the above method, candidate queries can be collected, target queries that cannot be correctly processed by the large-scale model can be mined from them, and furthermore, corresponding training samples can be accurately constructed respectively based on each target query. Thereby, the quality of the constructed training samples can be improved, and furthermore, when performing optimization training on the large-scale model using the subsequently constructed training samples, accurate optimization can be performed on the weaknesses in the large-scale model inference ability, and the optimization effect and the like can be improved.

[0020] Specifically, what the optimization trigger conditions are can be determined according to actual needs. For example, it can refer to periodically executing the method of the embodiment shown in FIG. 1 after a predetermined period, or it can also refer to receiving a trigger instruction sent by a user, etc.

[0021] In step 101, when it is determined that the optimization trigger conditions are met, queries collected from a predetermined data source that can be used as the input of the large-scale model can be used as candidate queries. Candidate queries can be obtained by query mining (or called inference demand mining, etc.). Specifically, what data sources the predetermined data source includes can be determined according to actual needs. For example, it can include product logs, public evaluation sets, open data sets, and the Internet, etc.

[0022] In step 102, a target query can be selected from the acquired candidate queries. The target query is a query that the large-scale model cannot process correctly, and the large-scale model refers to a large-scale model that requires optimization training according to the method described herein. The large-scale model may also refer to the original large-scale model acquired through pre-training + supervised fine-tuning (SFT) training, or it may refer to the large-scale model after one or more optimization training sessions have been performed on the original large-scale model.

[0023] Preferably, when selecting a target query from candidate queries, the following processing can be performed on each candidate query: the candidate query is used as input to a large-scale model, a response corresponding to the candidate query generated by the large-scale model is obtained, and in response to the determination that the response is an error response that does not match the candidate query, the candidate query is set as the selected target query.

[0024] In other words, for each candidate query, a corresponding response can be generated using a large-scale model, and if the generated response is determined to be an error response, that candidate query can be determined as a target query. In short, the capabilities of the large-scale model can be used to support effectiveness evaluation and to discover weaknesses in the current reasoning ability.

[0025] Preferably, for any given candidate query, the corresponding response is input into a pre-trained evaluation model, and an evaluation result can be obtained to determine whether the response output by the evaluation model matches the candidate query.

[0026] For example, for a candidate query a, a large-scale model can be used to generate a response a corresponding to candidate query a. Then, candidate query a and response a can be input into an evaluation model to obtain an evaluation result showing whether or not the response a output by the evaluation model matches candidate query a. If the evaluation result shows that response a does not match candidate query a, candidate query a can be selected as the target query.

[0027] The evaluation model may be one that has been trained using a large number of pre-built training samples. Accordingly, the evaluation results can be obtained quickly using the evaluation model, and because the evaluation model has been trained using a large number of training samples, the accuracy of the output evaluation results is ensured.

[0028] In practical applications, for example, assuming there are 1000 candidate queries, the number of target queries could be 300, and each target query can be used to construct a set of error questions.

[0029] Furthermore, in step 103, you can build a corresponding training sample for each target query in the error set.

[0030] Preferably, for any target query, the following processing can be performed: the processing determines the problem type corresponding to the target query, determines the sample type corresponding to the problem type, and constructs a training sample of the sample type based on the target query.

[0031] The above process allows for the accurate construction of corresponding training samples for different target queries, thereby improving the optimization training effect of subsequent models.

[0032] Preferably, the problem types may include data overwrite problems, model capability problems, and data quality problems. Accordingly, a method for determining the problem type corresponding to any target query may include searching a training sample set, where the training sample set includes training samples used when performing SFT training on a large model (which may include training samples from SFT training each time); determining that the problem type corresponding to the target query is a data overwrite problem in response to the detection of no matching training samples for the target query; and determining whether the problem type corresponding to the target query is a model capability problem or a data quality problem using an in-context learning (ICL) method in response to the detection of training samples for the target query.

[0033] For any target query, the training sample set can first be searched to determine whether or not there are training samples in the training sample set that match the target query. Among these, training samples that match the target query can refer to training samples that are exactly the same as the target query, or they can refer to training samples whose similarity to the target query is greater than a predetermined threshold. The specific value of the predetermined threshold can be determined according to the actual needs. If it is determined that there are no training samples in the training sample set that match the target query, the problem type corresponding to the target query can be determined to be a data overwrite problem. Conversely, the ICL method can be used to further determine whether or not the problem type corresponding to the target query is a model capability problem. If it is not a model capability problem, the problem type corresponding to the target query can be determined to be a data quality problem.

[0034] Preferably, the sample type includes a Class 1 training sample and a Class 2 training sample, wherein the Class 1 training sample is a training sample for pre-training and the Class 2 training sample is a training sample for SFT training. Accordingly, a method for determining the sample type corresponding to the problem type of any target query may include the steps of: determining that the sample type is a Class 1 training sample in response to the determination that the problem type is a model capability problem; and determining that the sample type is a Class 2 training sample in response to the determination that the problem type is a data overwrite problem or a data quality problem.

[0035] In other words, for model capability issues, Type 1 training samples can be constructed and used for pre-training large-scale models, and for data quality issues or data overwrite issues, Type 2 training samples can be constructed and used for SFT training of large-scale models.

[0036] The above process makes it possible to discover multidimensional problems such as data overwrite problems, data quality problems, and model capability problems, and to accurately construct corresponding training samples for different types of problems, thereby improving the optimization training effect of subsequent models.

[0037] There are no limitations on how training samples are constructed for each target query. For example, for any target query, a training sample can be constructed using the following method, which involves searching the internet for responses corresponding to the target query, and further obtaining queries similar to the target query from the internet, obtaining corresponding responses, and then constructing a training sample based on the obtained content.

[0038] Based on the training sample built in Step 103, optimization training can be performed on a large-scale model.

[0039] Preferably, a large-scale model can be pre-trained using a Class 1 training sample to obtain a first model, then a second model can be obtained by performing SFT training on the first model using a Class 2 training sample, and an actual reasoning application can be performed using the second model. Alternatively, a large-scale model can be pre-trained using a Class 1 training sample and a large-scale model can be performed by performing SFT training on the large-scale model using a Class 2 training sample, and after the SFT training is completed and a third model is obtained, an actual reasoning application can be performed using the third model in response to the fact that the pre-training is not yet complete, and a first model can be obtained in response to the completion of the pre-training, a second model can be obtained by performing SFT training on the first model using a Class 2 training sample, and an actual reasoning application can be performed by replacing the third model with the second model.

[0040] In other words, for a large-scale model to be optimized, first, pre-training can be performed using Category 1 training samples. After pre-training is complete, then SFT training can be performed on the large-scale model using Category 2 training samples. Furthermore, actual reasoning applications can be performed on the large-scale model after SFT training. However, in actual applications, the pre-training time is usually long. If the actual reasoning application is performed on the large-scale model after waiting for pre-training and SFT training to be completed, the large-scale model may become unusable for a long period of time, affecting the normal processing of business operations. To overcome this problem, pre-training can be performed on the large-scale model using Category 1 training samples, and then SFT training can be performed on the large-scale model using Category 2 training samples. Since the SFT training time is relatively short, actual reasoning applications can be performed using the large-scale model after SFT training to avoid affecting the normal processing of business operations for a long period of time. Then, after pre-training is complete, SFT training can be performed on the large-scale model after pre-training using Category 2 training samples to obtain the large-scale model necessary for the current optimization training, and the large-scale model after SFT training can be replaced for actual reasoning applications.

[0041] In practical applications, when training a large-scale model using Category 1 or Category 2 training samples, other training samples can be combined. For example, pre-training can be performed by combining training samples used in past pre-training sessions to improve training effectiveness.

[0042] Preferably, before selecting a target query from the candidate queries described in step 102, the query type of each candidate query can be determined, i.e., a demand classification can be performed. Accordingly, after selecting a target query from the candidate queries, the number of target queries belonging to each different query type can be statistically determined, and the processing capacity of the large-scale model for different query types can be determined based on the statistical results.

[0043] The specific types of queries included in the aforementioned query types can be determined according to the actual needs. For example, they could include common sense reasoning, mathematical calculations, question and answer, and code generation.

[0044] There are no limitations on how the query type of each candidate query is determined; for example, it can be determined using a pre-trained classification model. Accordingly, by statistically analyzing the number of target queries belonging to different query types, the processing power of the large model for different query types can be determined, that is, it is possible to determine which / what kind of query types the large model is underperforming against, and to discover weaknesses in the reasoning ability of the large model.

[0045] Preferably, after performing optimization training on the large-scale model using the training samples described in step 103, the large-scale model after optimization training can be further tested using test queries corresponding to different query types, and the processing capability of the large-scale model for different query types after optimization training can be determined based on the test results.

[0046] For example, a test query can be used as input to a large-scale model after optimization training, and the corresponding response output by the large-scale model can be obtained. Furthermore, the test query and the corresponding response can be input to an evaluation model, and an evaluation result can be obtained from whether the test query output by the evaluation model matches the corresponding response. Furthermore, the evaluation result can be used as a test result, and by combining each test result, the processing ability of the large-scale model for different query types after optimization training can be determined. In addition, it is possible to determine in what aspects the large-scale model has improved after optimization training, and to verify the effectiveness of optimization training.

[0047] Figure 2 is a flowchart of an embodiment of the large-scale model optimization training method described herein. As shown in Figure 2, it includes the following specific implementation methods.

[0048] Step 201 involves obtaining training samples, each of which is a corresponding training sample constructed based on each target query, where the target query is a query selected from candidate queries that the large-scale model cannot process correctly, and the candidate queries are queries collected from a predetermined data source that can be used as input for the large-scale model.

[0049] In step 202, optimization training is performed on the large-scale model using the aforementioned training samples.

[0050] Because current large-scale models still have certain shortcomings when it comes to handling reasoning tasks, improving the reasoning capabilities of large-scale models, especially their ability to handle complex reasoning, is an urgent issue that needs to be addressed.

[0051] The solution described in the embodiment of the above method provides a data-driven large-scale model self-evolution method, enabling the optimization training of large-scale models through operations such as query mining, problem discovery, training sample construction, and model training. This improves the reasoning ability of large-scale models, and accordingly, using large-scale models for actual reasoning applications can improve the accuracy of reasoning results.

[0052] The training sample obtained in step 201 may be a training sample constructed by the method of the embodiment shown in Figure 1.

[0053] Preferably, the training samples may include a Class 1 training sample and a Class 2 training sample, and accordingly, step 202 may include a method of performing optimization training on a large model using the training samples, which may include pre-training the large model using the Class 1 training sample to obtain a first model, performing SFT training on the first model using the Class 2 training sample to obtain a second model, and performing actual reasoning applications using the second model, or pre-training the large model using the Class 1 training sample and performing SFT training on the large model using the Class 2 training sample, and after the SFT training is complete and a third model is obtained, performing actual reasoning applications using the third model in response to the pre-training not yet being completed, and in response to the pre-training being completed, obtaining the first model, performing SFT training on the first model using the Class 2 training sample to obtain a second model, and replacing the third model with the second model for actual reasoning applications.

[0054] In other words, for a large-scale model to be optimized, first, pre-training can be performed using Category 1 training samples. After pre-training is complete, then SFT training can be performed on the large-scale model using Category 2 training samples. Furthermore, actual reasoning applications can be performed on the large-scale model after SFT training. However, in actual applications, the pre-training time is usually long. If the actual reasoning application is performed on the large-scale model after waiting for pre-training and SFT training to be completed, the large-scale model may become unusable for a long period of time, affecting the normal processing of business operations. To overcome this problem, pre-training can be performed on the large-scale model using Category 1 training samples, and then SFT training can be performed on the large-scale model using Category 2 training samples. Since the SFT training time is relatively short, actual reasoning applications can be performed using the large-scale model after SFT training to avoid affecting the normal processing of business operations for a long period of time. Then, after pre-training is complete, SFT training can be performed on the large-scale model after pre-training using Category 2 training samples to obtain the large-scale model necessary for the current optimization training, and the large-scale model after SFT training can be replaced for actual reasoning applications.

[0055] Preferably, after performing optimization training on the large-scale model using the training samples described in step 202, the large-scale model after optimization training can be further tested using test queries corresponding to different query types, and the processing capability of the large-scale model for different query types after optimization training can be determined based on the test results.

[0056] For example, a test query can be used as input to a large-scale model after optimization training, and the corresponding response output by the large-scale model can be obtained. Furthermore, the test query and the corresponding response can be input to an evaluation model, and an evaluation result can be obtained from whether the test query output by the evaluation model matches the corresponding response. Furthermore, the evaluation result can be used as a test result, and by combining each test result, the processing ability of the large-scale model for different query types after optimization training can be determined. In addition, it is possible to determine in what aspects the large-scale model has improved after optimization training, and to verify the effectiveness of optimization training.

[0057] Combining the above explanation, Figure 3 is a schematic diagram of the overall implementation process of the large-scale model optimization training method of this disclosure. As shown in Figure 3, demand classification refers to determining the query type for each candidate query, instruction search refers to searching the training sample set for each target query to determine the problem type for each target query, and the large-scale model after optimization training can be used to generate a new product log in the data source and can be used for selecting target queries for the next optimization training, etc. The specific implementation of each step shown in Figure 3 can be found in the related explanations above and are omitted here.

[0058] Furthermore, the queries in the solutions described in this disclosure are typically in text format, and such queries may be problems, i.e., text-format problems to be input into a large-scale model. The solutions described in this disclosure can collect candidate queries from a given data source that can be used as input to a large-scale model in text format, preferably such candidate queries should include as many different query types as possible, such as common sense reasoning, mathematical calculations, question-and-answer, and code generation queries, then target queries that the large-scale model cannot process correctly can be selected from the candidate text-format queries, and a corresponding training sample can be constructed based on each target text-format query, and the large-scale model can be optimized and trained using the constructed training samples, so that the optimized and trained large-scale model can provide more accurate responses to the input text-format queries.

[0059] For the sake of brevity, the embodiments of each of the methods described above are all described as a combination of operations. However, those skilled in the art should recognize that the disclosure is not limited by the order of operations described, as some steps may be performed in other orders or simultaneously. Furthermore, all embodiments described herein are preferred embodiments, and the operations and modules described are not necessarily essential to the disclosure. Also, some embodiments lack detailed sections, and relevant descriptions in other embodiments can be referenced.

[0060] The above is a description of embodiments of the method, and the following will further describe the solutions described in this disclosure with reference to embodiments of the apparatus.

[0061] Figure 4 is a schematic diagram of the configuration of Embodiment 400 of the training sample acquisition apparatus of the present disclosure. As shown in Figure 4, it includes a query mining module 401, a problem discovery module 402, and a sample construction module 403.

[0062] The query mining module 401 is used to identify candidate queries collected from a given data source that can be used as input to a large-scale model, in response to a determination that the optimization trigger conditions are met.

[0063] The problem discovery module 402 selects a target query from candidate queries, and the target query is used because it is a query that the large-scale model cannot process correctly.

[0064] The sample construction module 403 is used to construct corresponding training samples based on each target query, and these training samples are used to perform optimization training on the large-scale model.

[0065] Using the solution described in the embodiment of the above-mentioned device, candidate queries can be collected, target queries that the large-scale model cannot process correctly can be mined from among them, and furthermore, corresponding training samples can be accurately constructed based on each target query, thereby improving the quality of the constructed training samples. Moreover, when performing optimization training on the large-scale model using the subsequently constructed training samples, it is possible to accurately optimize weaknesses in the large-scale model's reasoning ability, thereby improving the optimization effect.

[0066] If it is determined that the optimization trigger conditions are met, the query mining module 401 can select candidate queries collected from a given data source that can be used as input to a large-scale model. Candidate queries can be obtained through query mining (or inferential demand mining, etc.).

[0067] From the retrieved candidate queries, the problem detection module 402 can select target queries, which are queries that the large-scale model cannot process correctly.

[0068] Preferably, when the problem discovery module 402 selects a target query from candidate queries, it can perform the following processing on any candidate query: the processing involves taking the candidate query as input to a large-scale model, obtaining a response corresponding to the candidate query generated by the large-scale model, and, in response to determining that the response is an error response that does not match the candidate query, setting the candidate query as the selected target query.

[0069] Preferably, for any candidate query, the problem discovery module 402 can input the candidate query and its corresponding response into a pre-trained evaluation model and obtain an evaluation result showing whether the response output by the evaluation model matches the candidate query.

[0070] Furthermore, sample building module 403 builds corresponding training samples based on each target query.

[0071] Preferably, for any target query, the sample building module 403 can perform the following processing, which includes determining the problem type corresponding to the target query, determining the sample type corresponding to the problem type, and building a training sample of the sample type based on the target query.

[0072] Preferably, the problem types may include data overwrite problems, model capability problems, and data quality problems. Accordingly, the method by which the sample building module 403 determines the problem type corresponding to any target query may include searching a training sample set, the training sample set including training samples used when performing SFT training on a large model, determining that the problem type corresponding to the target query is a data overwrite problem in response to the detection of no training samples matching the target query, and determining, by the ICL method, whether the problem type corresponding to the target query is a model capability problem or a data quality problem in response to the detection of training samples matching the target query.

[0073] Preferably, the sample type includes a Class 1 training sample and a Class 2 training sample, wherein the Class 1 training sample is a training sample for pre-training, and the Class 2 training sample is a training sample for SFT training. Accordingly, for any target query, the sample construction module 403 may determine the sample type corresponding to the problem type of the target query by determining that the sample type is a Class 1 training sample in response to the problem type being determined to be a model capability problem, and by determining that the sample type is a Class 2 training sample in response to the problem type being determined to be a data overwrite problem or a data quality problem.

[0074] In other words, for model capability problems, we can build Type 1 training samples, and for data quality problems or data overwrite problems, we can build Type 2 training samples.

[0075] Preferably, the query mining module 401 can further determine the query type of each candidate query before selecting a target query from the candidate queries, and accordingly, the problem discovery module 402 can further statistically determine the number of target queries belonging to each different query type after selecting a target query from the candidate queries, and further determine the processing capacity of the large-scale model for different query types based on the statistical results.

[0076] Figure 5 is a schematic diagram of the configuration of Example 500 of the large-scale model optimization training apparatus of the present disclosure. As shown in Figure 5, it includes a sample acquisition module 501 and a model training module 502.

[0077] The sample acquisition module 501 is used to acquire training samples, the training samples being corresponding training samples constructed based on each target query, the target queries being queries that the large-scale model cannot process correctly selected from candidate queries, and the candidate queries being queries collected from predetermined data sources that can be used as input to the large-scale model.

[0078] The model training module 502 is used to perform optimization training on a large-scale model using the aforementioned training samples.

[0079] The solution described in the embodiment of the above-mentioned device provides a data-driven large-scale model self-evolution method, enabling the optimization training of large-scale models through operations such as query mining, problem discovery, training sample construction, and model training. This improves the reasoning ability of large-scale models, and accordingly, using large-scale models for actual reasoning applications can improve the accuracy of reasoning results.

[0080] Preferably, the training samples include a Class 1 training sample and a Class 2 training sample, and accordingly, the model training module 502 can pre-train a large-scale model using the Class 1 training sample to obtain a first model, perform SFT training on the first model using the Class 2 training sample to obtain a second model, and perform actual reasoning applications using the second model, or pre-train a large-scale model using the Class 1 training sample and perform SFT training on the large-scale model using the Class 2 training sample, respectively, and after the SFT training is completed and a third model is obtained, in response that pre-training is not yet complete, perform actual reasoning applications using the third model, and in response that pre-training is complete, obtain the first model, perform SFT training on the first model using the Class 2 training sample to obtain a second model, and replace the third model with the second model to perform actual reasoning applications.

[0081] Preferably, the model training module 502 can perform optimization training on a large-scale model using training samples, and then further test the optimized-trained large-scale model using test queries corresponding to different query types, and determine the processing capability of the optimized-trained large-scale model for different query types based on the test results.

[0082] The specific workflow of the embodiment of the apparatus shown in Figures 4 and 5 can be found in the related descriptions of the embodiments of the method described above, and is therefore omitted here.

[0083] In other words, the solutions described in this disclosure enable the self-evolution of large-scale model reasoning capabilities based on data-driven and closed-loop iteration, thereby improving the reasoning capabilities of large-scale models and further enhancing the accuracy of reasoning results.

[0084] The solutions described in this disclosure can be applied to the field of artificial intelligence technology, particularly to areas such as large-scale models, deep learning, and natural language processing. Artificial intelligence is the study of simulating certain human thought processes and intelligent actions (e.g., learning, reasoning, thinking, planning, etc.) on computers, and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing, while artificial intelligence software technologies mainly include several areas such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0085] The data in the embodiments described herein are not for any specific user and do not reflect the personal information of any specific user. In the proposed technologies described herein, all processing of relevant user personal information, including collection, storage, use, processing, transmission, provision, and disclosure, complies with applicable laws and regulations and does not violate public order and morality.

[0086] According to embodiments of the present disclosure, the present disclosure further provides electronic devices, readable storage media, and computer program products.

[0087] Figure 6 shows a schematic block diagram of an electronic device 600 for carrying out an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other appropriate computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure as required.

[0088] As shown in Figure 6, the device 600 includes a computing unit 601, which can perform various appropriate operations and processes based on computer programs stored in read-only memory (ROM) 602 or computer programs loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 can also store various programs and data necessary for the device 600 to operate. The computing unit 601, ROM 602, and RAM 603 are connected to each other via bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0089] Multiple components within the device 600 are connected to an I / O interface 605 and include an input unit 606 such as a keyboard and mouse, an output unit 607 such as various types of displays and speakers, a storage unit 608 such as disks and optical disks, and a communication unit 609 such as a network card, modem, and wireless communication transceiver. The communication unit 609 enables the device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunications networks.

[0090] The computing unit 601 is a general-purpose and / or dedicated processing component with various processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, a computing unit that runs various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the methods described herein. For example, in some embodiments, the methods described herein can be implemented as a computer software program tangibly contained in a machine-readable medium such as a storage unit 608. In some embodiments, part or all of the computer program is loaded and / or installed into the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the methods described herein can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the methods described herein via any other suitable method (e.g., by firmware).

[0091] Various implementations of the systems and technologies described herein can be implemented as digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), load-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include being implemented as one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be an application-specific or general-purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, at least one input device, and at least one output device.

[0092] Program code for carrying out the methods of this disclosure can be written using any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations defined in the flowcharts and / or block diagrams are performed. The program code may run entirely on a machine, partially on a machine, partially on a machine as a standalone software package, partially on a remote machine, or entirely on a remote machine or server.

[0093] In the context of this disclosure, machine-readable media may be tangible media that can contain or store programs used in conjunction with instruction execution systems, devices, or equipment. Machine-readable media may be machine-readable signal media or machine-readable storage media. Machine-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0094] To provide user interaction, the systems and techniques described herein can be implemented on a computer, which may include a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball), and the user may provide input to the computer via the keyboard and pointing device. Other types of devices may also be used to provide user interaction, for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from the user may be received in any form (including acoustic input, voice input, and haptic input).

[0095] The systems and technologies described herein can be implemented in a computing system including backend components (e.g., a data server), a computing system including middleware components (e.g., an application server), a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser, the user interacting with the implementation of the systems and technologies described herein through the graphical user interface or web browser), or in a computing system including any combination of such backend components, middleware components, and frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0096] A computer system can include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers that have a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server that incorporates blockchain technology.

[0097] It should be understood that steps can be rearranged, added, or deleted using the various forms of flows shown above. For example, each step described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the technical proposal disclosed herein can achieve the desired result.

[0098] The specific implementation methods described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art can make various modifications, combinations, subcombinations, and substitutions based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure must be within the scope of protection of this disclosure.

Claims

1. A method for acquiring training samples performed by a computer, In response to the determination that the optimization trigger conditions are met, the steps include selecting candidate queries from a predetermined data source that can be used as input to a large-scale model, and A step of selecting a target query from the candidate queries, wherein the target query is a query that the large-scale model cannot process correctly. The steps include: constructing a corresponding training sample based on each target query, wherein the training sample is used to perform optimization training on the large-scale model; The step of selecting a target query from the aforementioned candidate queries is: For each candidate query, the following steps are performed: The process takes the candidate query as input to the large-scale model, obtains a response corresponding to the candidate query generated by the large-scale model, and in response to the determination that the response is an error response that does not match the candidate query, sets the candidate query as a selected target query. The step of building the corresponding training sample based on each of the aforementioned target queries is as follows: For any target query, the following steps are included: The process involves determining the problem type corresponding to the target query, determining the sample type corresponding to the problem type, and constructing a training sample of the sample type based on the target query. The aforementioned problem types include data overwrite problems, model capability problems, and data quality problems. The step of determining the problem type corresponding to the aforementioned target query is: A step of searching for a training sample set, wherein the training sample set includes training samples used when performing supervised fine-tuning training on the large model, In response to the fact that no training samples matching the target query were found, the step of determining that the problem type corresponding to the target query is the data overwrite problem, The process includes: in response to the retrieval of training samples matching the target query, determining by a context-learning method whether the problem type corresponding to the target query is the model capability problem or the data quality problem; How to obtain training samples.

2. The step of determining that the response is an error response that does not match the candidate query is: The process includes inputting the candidate query and the response into a pre-trained evaluation model, and obtaining an evaluation result of whether the response output by the evaluation model matches the candidate query. The method for obtaining training samples according to claim 1.

3. The aforementioned sample type includes training samples of type 1 and training samples of type 2, wherein the training samples of type 1 are training samples for pre-training, and the training samples of type 2 are training samples for supervised fine-tuning training. The step of determining the sample type corresponding to the aforementioned problem type is: In response to the determination that the problem type is the model capability problem, the step of determining that the sample type is the training sample of the first class, The step of determining that the sample type is a training sample of the second class, in response to the determination that the problem type is the data overwrite problem or the data quality problem, The method for obtaining training samples according to claim 1.

4. Before the step of selecting a target query from the aforementioned candidate queries, there is a step of determining the query type of each candidate query, The process further includes, after selecting target queries from the candidate queries, statistically analyzing the number of target queries belonging to different query types, and determining the processing capacity of the large-scale model for different query types based on the statistical results. A method for obtaining training samples according to any one of claims 1 to 3.

5. A method for training a large-scale model optimization performed by a computer, A step of obtaining training samples, wherein the training samples are corresponding training samples constructed based on each target query, each target query is a query selected from candidate queries that a large-scale model cannot process correctly, and the candidate queries are queries collected from predetermined data sources that can be used as input to the large-scale model. The step includes performing optimization training on the large-scale model using the training samples, Selecting each target query from the aforementioned candidate queries is For each candidate query, the following steps are performed: The process takes the candidate query as input to the large-scale model, obtains a response corresponding to the candidate query generated by the large-scale model, and in response to the determination that the response is an error response that does not match the candidate query, sets the candidate query as a selected target query. The step of building the corresponding training sample based on each of the aforementioned target queries is as follows: For any target query, the following steps are included: The process involves determining the problem type corresponding to the target query, determining the sample type corresponding to the problem type, and constructing a training sample of the sample type based on the target query. The aforementioned problem types include data overwrite problems, model capability problems, and data quality problems. The step of determining the problem type corresponding to the aforementioned target query is: A step of searching for a training sample set, wherein the training sample set includes training samples used when performing supervised fine-tuning training on the large model, In response to the fact that no training samples matching the target query were found, the step of determining that the problem type corresponding to the target query is the data overwrite problem, The process includes: in response to the retrieval of training samples matching the target query, determining by a context-learning method whether the problem type corresponding to the target query is the model capability problem or the data quality problem; Large-scale model optimization training methods.

6. The aforementioned training sample includes a Class 1 training sample and a Class 2 training sample. The step of performing optimization training on the large-scale model using the training samples includes the steps of: pre-training the large-scale model using the training samples of the first type to obtain a first model; performing supervised fine-tuning training on the first model using the training samples of the second type to obtain a second model; performing actual reasoning applications using the second model; or pre-training the large-scale model using the training samples of the first type and supervised fine-tuning training on the large-scale model using the training samples of the second type, and after the supervised fine-tuning training is completed and a third model is obtained, performing actual reasoning applications using the third model in response to the fact that the pre-training is not yet complete; and in response to the completion of the pre-training, obtaining the first model, performing supervised fine-tuning training on the first model using the training samples of the second type to obtain a second model; and replacing the third model with the second model to perform actual reasoning applications. The large-scale model optimization training method according to claim 5.

7. The step of performing optimization training on the large-scale model using the training samples further includes testing the optimized-trained large-scale model using test queries corresponding to different query types, and determining the processing capability of the optimized-trained large-scale model for the different query types based on the test results. The large-scale model optimization training method according to claim 5.

8. A training sample acquisition device, At least one processor, The system comprises at least one processor and a memory that is communicated with by it, The memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the respective functions of the query mining module, the problem discovery module, and the sample building module are executed. The query mining module is used to select candidate queries, which are collected from predetermined data sources that can be used as input to a large-scale model, in response to the determination that the optimization trigger conditions are met. The problem discovery module is used to select a target query from the candidate queries determined by the query mining module, and the target query is a query that the large-scale model cannot process correctly. The sample building module is used to build corresponding training samples based on each target query selected by the problem discovery module, and the training samples are used to perform optimization training on the large-scale model. The problem discovery module performs the following processing on each arbitrary candidate query: the processing takes the candidate query as input to the large-scale model, obtains a response corresponding to the candidate query generated by the large-scale model, and in response to the determination that the response is an error response that does not match the candidate query, sets the candidate query as a selected target query. The sample building module performs the following processing for each arbitrary target query: the processing determines the problem type corresponding to the target query, determines the sample type corresponding to the problem type, and builds a training sample of the sample type based on the target query. The aforementioned problem types include data overwrite problems, model capability problems, and data quality problems. The sample building module searches for a training sample set for each arbitrary target query, the training sample set includes training samples used when performing supervised fine-tuning training on the large model, and in response to the absence of a training sample matching the target query, it determines that the problem type corresponding to the target query is the data overwrite problem, and in response to the presence of a training sample matching the target query, it determines, using a contextual learning method, whether the problem type corresponding to the target query is the model capability problem or the data quality problem. Training sample acquisition device.

9. The problem detection module inputs the candidate query and the response into a pre-trained evaluation model to determine that the response is an error response that does not match the candidate query, and obtains an evaluation result from the evaluation model to determine whether the response matches the candidate query. The training sample acquisition device according to claim 8.

10. The aforementioned sample type includes training samples of type 1 and training samples of type 2, wherein the training samples of type 1 are training samples for pre-training, and the training samples of type 2 are training samples for supervised fine-tuning training. The sample building module determines that the sample type is a training sample of the first type in response to the determination that the problem type is a model capability problem, and determines that the sample type is a training sample of the second type in response to the determination that the problem type is a data overwrite problem or a data quality problem. The training sample acquisition device according to claim 8.

11. The aforementioned query mining module is further used to determine the query type of each candidate query. The problem discovery module further statistically analyzes the number of target queries belonging to different query types and uses the statistical results to determine the processing capacity of the large-scale model for different query types. The training sample acquisition apparatus according to claim 8 or 9.

12. A large-scale model optimization training device, At least one processor, The system comprises at least one processor and a memory that is communicated with by it, The memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the functions of the sample acquisition module and the model training module are executed, respectively. The sample acquisition module is used to acquire training samples, the training samples being corresponding training samples constructed based on each target query, each target query being a query selected from candidate queries that the large-scale model cannot process correctly, and the candidate queries being queries collected from predetermined data sources that can be used as input to the large-scale model. The model training module is used to perform optimization training on the large-scale model using the training samples. The sample acquisition module performs the following processing on each arbitrary candidate query: the processing takes the candidate query as input to the large-scale model, obtains a response corresponding to the candidate query generated by the large-scale model, and in response to the determination that the response is an error response that does not match the candidate query, sets the candidate query as a selected target query. The sample acquisition module performs the following processing for each arbitrary target query: the processing determines the problem type corresponding to the target query, determines the sample type corresponding to the problem type, and constructs a training sample of the sample type based on the target query. The aforementioned problem types include data overwrite problems, model capability problems, and data quality problems. The sample acquisition module searches for a training sample set for each arbitrary target query, the training sample set includes training samples used when performing supervised fine-tuning training on the large model, and in response to the absence of a training sample matching the target query, it determines that the problem type corresponding to the target query is the data overwrite problem, and in response to the presence of a training sample matching the target query, it determines, using a contextual learning method, whether the problem type corresponding to the target query is the model capability problem or the data quality problem. A large-scale model optimization training system.

13. The aforementioned training sample includes a Class 1 training sample and a Class 2 training sample. The model training module pre-trains the large-scale model using the first type of training sample to obtain a first model, performs supervised fine-tuning training on the first model using the second type of training sample to obtain a second model, and performs actual reasoning applications using the second model; or, pre-trains the large-scale model using the first type of training sample and performs supervised fine-tuning training on the large-scale model using the second type of training sample, and after the supervised fine-tuning training is completed and a third model is obtained, in response that the pre-training is not yet complete, performs actual reasoning applications using the third model; in response that the pre-training is complete, obtains the first model, performs supervised fine-tuning training on the first model using the second type of training sample to obtain a second model, and replaces the third model with the second model to perform actual reasoning applications. The large-scale model optimization training device according to claim 12.

14. The model training module is further used to perform optimization training on the large-scale model using the training samples, then to test the optimized-trained large-scale model using test queries corresponding to different query types, and to determine the processing capability of the optimized-trained large-scale model for different query types based on the test results. The large-scale model optimization training apparatus according to claim 12 or 13.

15. It is an electronic device, At least one processor, The memory includes, which is communicated with at least one processor, The memory stores instructions that can be executed by the at least one processor, and when an instruction is executed by the at least one processor, the at least one processor performs the method according to any one of claims 1 to 3. electronic equipment.

16. Electronic device, At least one processor, The memory includes, which is communicated with at least one processor, The memory stores instructions that can be executed by the at least one processor, and when an instruction is executed by the at least one processor, the at least one processor performs the method according to any one of claims 5 to 7. electronic equipment.

17. A non-temporary, computer-readable storage medium in which computer instructions are stored, The computer instruction causes the computer to perform the method described in any one of claims 1 to 3. A non-temporary, computer-readable storage medium.

18. A non-temporary computer-readable storage medium in which computer instructions are stored, The computer instruction causes the computer to perform the method described in any one of claims 5 to 7. A non-temporary, computer-readable storage medium.

19. It is a computer program, When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is realized. Computer program.

20. A computer program, When the computer program is executed by a processor, the method according to any one of claims 5 to 7 is realized. Computer program.

Citation Information

Patent Citations

  • Generative large language model training method and model-based search method

    CN116226334A

  • Model training method and device suitable for large language model, equipment and medium

    CN116976424A

  • Neural network system, machine learning method, and program

    WO2019031305A1