Method and device for acquiring training data of fine-tuning large model
The inference chain and answers of mathematical problems are automatically generated through the large model, and used as training data to fine-tune the large model, solving the problem of low accuracy in mathematical problems calculation, and achieving the effect of reducing the calculation error rate and improving mathematical calculation ability.
Patent Information
- Application Number
- CN202510221434.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
When dealing with mathematical problems, the calculation results are not accurate, resulting in high calculation error rates and insufficient mathematical calculation ability.
Automatically generate inference chains and answers for mathematical problems through large models as training data, and are used to fine-tune the large model, thereby reducing the calculation error rate and improving mathematical computing power. Specific methods include obtaining mathematical problems, inputting large models to obtain solutions, generating executable code, executing code to verify the consistency of answers, and using consistent training data to fine-tune the large model.
By automatically generating high-accurate reasoning chains and answers, the calculation error rate of large language models for mathematical problems is reduced, the calculation ability of large models for mathematical problems is improved, and labor costs and time consumption are reduced.
Smart Images

Figure CN120069080A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of large model technology, and more particularly, to a method and apparatus for acquiring training data for fine-tuning a large model. Background Art
[0002] A large language model (LLM) is a large neural network model based on deep learning technology. It features hundreds of millions of parameters and is typically pre-trained on massive amounts of text data, enabling it to understand and generate natural language. Currently, LLMs are often used to solve problems involving data expressed in natural language. However, when used to solve mathematical problems, LLMs often suffer from low accuracy in mathematical calculations. Summary of the Invention
[0003] The embodiments in this specification aim to provide a method and apparatus for obtaining training data for fine-tuning a large model. The large model can automatically generate high-accuracy reasoning chains and answers to mathematical problems as training data for fine-tuning the large model, thereby reducing the calculation error rate of the large language model for mathematical problems, improving the calculation ability of the large model for mathematical problems, and solving the shortcomings of the existing technology.
[0004] According to a first aspect, a method for obtaining training data for fine-tuning a large model is provided, comprising:
[0005] Obtaining a first mathematical problem, inputting the first mathematical problem into a target macromodel, and obtaining a first solution corresponding to the first mathematical problem, the first solution including a reasoning chain for deriving an answer to the problem and the first answer;
[0006] Based on the first mathematical problem and the inference chain, a first executable code for determining an answer to the first mathematical problem is generated, and the first executable code is executed to obtain a second answer corresponding to the first mathematical problem; and whether the first answer and the second answer are consistent is determined. If the first answer and the second answer are consistent, the first mathematical problem is used as a training sample, and the inference chain and the first answer are used as training labels corresponding to the training sample to fine-tune the target large model.
[0007] In one possible implementation, executing the first executable code to obtain a second answer corresponding to the first mathematical problem includes:
[0008] Determine whether the first executable code matches the reasoning chain; if the first executable code matches the reasoning chain, execute the first executable code to obtain a second answer corresponding to the first mathematical problem.
[0009] In one possible implementation, the method further includes, if the first executable code does not match the reasoning chain, generating a second executable code corresponding to the first mathematical problem, and determining whether the second executable code matches the reasoning chain; and if the second executable code matches the reasoning chain, executing the second executable code to obtain a second answer corresponding to the second mathematical problem.
[0010] In one possible implementation, determining whether the first executable code matches the inference chain includes:
[0011] Determining a matching degree between the first executable code and the inference chain by using a preset model;
[0012] If the matching degree reaches a preset threshold, it is determined that the executable code and the reasoning chain match; if the matching degree does not reach the preset threshold, it is determined that the executable code and the reasoning chain do not match.
[0013] In one possible implementation, generating a first executable code for determining an answer to the first mathematical problem according to the first mathematical problem and the reasoning chain includes:
[0014] The prompt word including the first mathematical problem, the reasoning chain, and a prompt portion indicating that a corresponding executable code is generated based on the first mathematical problem and the reasoning chain is input into a preset model to obtain the first executable code.
[0015] In a possible implementation, the method further includes abandoning fine-tuning the target large model through the reasoning chain and the first answer if the first answer and the second answer are inconsistent.
[0016] In a possible implementation, obtaining a first mathematical problem includes:
[0017] randomly extracting a plurality of second math problems belonging to a target math domain from a pre-acquired math problem set;
[0018] A target prompt word is input into a preset model to obtain a first mathematical problem, wherein the target prompt word includes the multiple second mathematical problems and a prompt part indicating that a mathematical problem belonging to a target mathematical field but different from the multiple second mathematical problems is generated based on the multiple second mathematical problems.
[0019] In one possible implementation, inputting a first mathematical problem into a target macro model to obtain a first answer corresponding to the first mathematical problem includes:
[0020] Inputting the first mathematical problem into the target macro model multiple times to obtain multiple first solutions corresponding to the first mathematical problem;
[0021] Generating, based on the first mathematical problem and the chain of reasoning, a first executable code for determining an answer to the first mathematical problem, and executing the first executable code to obtain a second answer corresponding to the first mathematical problem, comprising:
[0022] For each first solution, a first executable code for determining the answer to the first mathematical problem is generated based on the first mathematical problem and the reasoning chain in the first solution. The first executable code is executed to obtain a second answer corresponding to the first mathematical problem.
[0023] According to a second aspect, a device for acquiring training data for fine-tuning a large model is provided, comprising:
[0024] an acquiring unit configured to acquire a first mathematical problem, input the first mathematical problem into a target macromodel, and obtain a first solution corresponding to the first mathematical problem, wherein the first solution includes a reasoning chain for deriving an answer to the problem and the first answer;
[0025] The determination unit is configured to generate a first executable code for determining an answer to the first mathematical problem based on the first mathematical problem and the reasoning chain, execute the first executable code to obtain a second answer corresponding to the first mathematical problem, and determine whether the first answer and the second answer are consistent. If the first answer and the second answer are consistent, use the first mathematical problem as a training sample, and the reasoning chain and the first answer as training labels corresponding to the training sample, for fine-tuning the target large model.
[0026] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in the first aspect.
[0027] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0028] By utilizing one or more of the methods, apparatuses, computing devices, and storage media in the above aspects, a large number of reasoning chains and answers for mathematical problems can be automatically generated through a large model, and more accurate reasoning chains and answers can be extracted from them as training data for fine-tuning the large model, thereby reducing the calculation error rate of the large language model for mathematical problems and improving the calculation ability of the large model for mathematical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 A schematic diagram showing a scheme for fine-tuning a large model;
[0031] Figure 2 A schematic diagram illustrating a method for obtaining training data for fine-tuning a large model according to an embodiment of this specification;
[0032] Figure 3 A flowchart illustrating a method for obtaining training data for fine-tuning a large model according to an embodiment of this specification is shown;
[0033] Figure 4 A schematic diagram illustrating obtaining a mathematical problem according to an embodiment of this specification;
[0034] Figure 5 A schematic diagram illustrating generating executable code according to an embodiment of this specification;
[0035] Figure 6 A schematic diagram illustrating a method for obtaining training data for fine-tuning a large model according to another embodiment of this specification.
[0036] Figure 7 A schematic diagram illustrating regenerating executable code according to an embodiment of this specification;
[0037] Figure 8 A structural diagram of a device for acquiring training data for fine-tuning a large model according to an embodiment of this specification is shown. DETAILED DESCRIPTION
[0038] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.
[0039] As mentioned above, a large language model (LLM) is a large neural network model based on deep learning technology. It is characterized by a parameter scale of hundreds of millions. It is usually pre-trained on massive text data and has the ability to understand and generate natural language. At present, large language models are often used to solve data problems expressed in natural language. However, since large language models essentially achieve the generation and understanding of natural language by learning the probability distribution of word sequences in large-scale text data languages, it is not based on numerical calculations that conform to mathematical calculation rules. Therefore, when large language models are used to solve mathematical problems, especially complex mathematical calculation problems, there is often a problem of incorrect mathematical calculation results.
[0040] Fine-tuning a large model involves further training a pre-trained general-purpose large language model (LLM) using domain-specific datasets to optimize the model's performance in that specific domain. Therefore, using a high-quality mathematical problem training set, a general-purpose large language model can be fine-tuned, thereby improving the large language model's ability to solve mathematical problems and reducing the error rate of the calculation results output by the fine-tuned large language model for mathematical problems. Figure 1 A schematic diagram showing a fine-tuning large model scheme is shown, Figure 1 As shown, mathematical problems can be constructed manually, and the expected calculation results for the mathematical problems can be manually labeled. Then, the large model can be fine-tuned using the manually generated mathematical problems and expected calculation results, thereby reducing the calculation error rate of the large language model for mathematical problems. However, the problems with this solution are: First, manually labeling high-quality mathematical training data requires a large number of labelers, which is very costly. Second, the labeling process is usually slow and consumes a lot of time and resources. Third, in actual production scenarios, the amount of training data required for fine-tuning large models often grows exponentially, and manual labeling often becomes a bottleneck for large-scale expansion of training data with high failure rates.
[0041] In order to solve the above technical problems, the embodiments of this specification provide a method for obtaining training data for fine-tuning a large model. Figure 2 A schematic diagram showing a method for obtaining training data for fine-tuning a large model according to an embodiment of this specification is shown. Figure 2 As shown, first, the acquired mathematical problem can be input into the macro model to obtain the solution. The solution includes the reasoning chain used to derive the solution and the solution. Then, executable code is generated based on the mathematical problem and the reasoning chain. The executable code is executed to determine whether the execution result is consistent with the solution. If the execution result is consistent with the solution, the mathematical problem and solution are used to fine-tune the target macro model.
[0042] The advantages of this method are: First, since the large model's solutions to mathematical problems, like its solutions to other problems, are essentially based on probabilistic statistics of data distribution, rather than numerical calculations that adhere to strict mathematical rules, the large model can output reasoning chains and answers to mathematical problems based on probabilistic statistics. However, the numerical calculations for each sub-step in the reasoning chain may not be correct, leading to errors in the final output answer. This method, however, allows the large model to generate reasoning chains and answers and then generate executable code based on the reasoning chains. Since the execution of computer executable code is numerical calculation that adheres to strict mathematical rules, highly accurate mathematical calculation results can be obtained by executing the executable code. Furthermore, the correctness of the answers output by the large model can be verified using these mathematical calculation results. Furthermore, based on the verification results, more accurate reasoning chains and answers can be identified from the large number of reasoning chains and answers generated by the large model. Therefore, this method allows the large model to generate a large number of reasoning chains and answers to mathematical problems and extract the more accurate ones for fine-tuning the large model, thereby reducing the error rate of the large language model's mathematical calculations and improving its computational capabilities. Second, this method can automatically generate training data for fine-tuning large models, thereby reducing labor costs and time consumption, and can quickly generate a large amount of training data to meet the needs of high-timeliness expansion of training data in production scenarios.
[0043] The detailed process of this method is further explained below. Figure 3 FIG. 1 is a flow chart showing a method for obtaining training data for fine-tuning a large model according to an embodiment of the present specification. Figure 3 Said method comprises at least the following steps:
[0044] Step S301: Obtain a first mathematical problem, input the first mathematical problem into a target macro model, and obtain a first solution corresponding to the first mathematical problem, wherein the first solution includes a reasoning chain for deriving an answer to the problem and a first answer;
[0045] Step S303: Generate a first executable code for determining an answer to the first mathematical problem based on the first mathematical problem and the reasoning chain; execute the first executable code to obtain a second answer corresponding to the first mathematical problem; determine whether the first answer and the second answer are consistent; if the first answer and the second answer are consistent, use the first mathematical problem as a training sample, and the reasoning chain and the first answer as training labels corresponding to the training sample, for fine-tuning the target large model.
[0046] First, in step S301, a first mathematical problem may be obtained, and the first mathematical problem may be input into a target large model to obtain a first answer corresponding to the first mathematical problem. In different embodiments, the first mathematical problem may be a different specific mathematical problem in different mathematical fields. In different embodiments, the first mathematical problem may be expressed in different forms, for example, it may be a mathematical problem described by natural language and / or a mathematical formula. In different embodiments, the target large model may be a large language model of different specific types, or may have a different neural network structure, and this specification does not limit this.
[0047] In different embodiments, the specific method of obtaining the first mathematical problem may be different. In one embodiment, a plurality of second mathematical problems belonging to the target mathematical field can be randomly extracted from a set of mathematical problems obtained in advance; the target prompt word is input into a preset model to obtain the first mathematical problem, and the target prompt word includes the plurality of second mathematical problems and a prompt part indicating that a mathematical problem belonging to the target mathematical field but different from the plurality of second mathematical problems is generated based on the plurality of second mathematical problems. In different specific embodiments, the target mathematical field may be a different mathematical field. In a specific embodiment, for example, it may be one or more of algebra, geometry, and calculus. In different specific embodiments, the preset model may be the same large language model as the target large model, or a different large language model from the target large model, or a pre-trained or optimized machine learning model for generating mathematical problems.
[0048] Figure 4 FIG. 1 shows a schematic diagram of obtaining a mathematical problem according to an embodiment of the present specification. Figure 2 As shown, for example, a plurality of math problems can be randomly extracted from a pre-acquired math problem set, such as algebraic problems, including, for example, "math problem 1, math problem...". Then, based on "math problem 1, math problem 2..." and a prompt part a indicating that a math problem belonging to an algebraic problem but different from "math problem 1, math problem 2..." is generated based on "math problem 1, math problem 2...", a prompt word is constructed, and the prompt word is input into the large model, for example, to obtain math problem 3. In different specific examples, the prompt word input into the large model may be different. In a specific example, for example, math problem 1 is "1+5=?" and math problem 2 is "4*5=?", the following prompt word can be constructed: "Generate a math problem that belongs to the same mathematical field as the following problems but is different from these problems, and these problems include '1+5=?', '4*5=?'", and the prompt word is input into the large model, for example, to obtain math problem 3 as "8+5=?".
[0049] The first solution obtained by inputting a first mathematical problem into the target macro model may include a chain of reasoning used to derive the answer to the problem and a first answer. The first answer is the answer to the first mathematical problem. Chain of Thought (CoT) is a technique for solving complex problems by explicitly displaying intermediate reasoning steps. Essentially, the macro model breaks down the problem-solving process into multiple logically related sub-steps to simulate the step-by-step thinking process of humans. In different specific embodiments, the chain of thought and first answer output by the target macro model may be different chains of reasoning and answers to the first mathematical problem. In one example, the mathematical problem input into the target macro model is "Z = X * Y X = 5 + 1 Y = 3 + 2, what is Z?" The answer output by the target macro model is, for example, "Z = 30," and the chain of thought output by the target macro model is, for example, "The calculation process is as follows: Step 1, based on X = 5 + 1, calculate X = 6; Step 2, based on Y = 3 + 2, calculate Y = 5; Step 3, based on Z = X * Y X = 6 Y = 5, calculate Z = 30."
[0050] Then, in step S303, first executable code for determining the answer to the first mathematical problem can be generated based on the first mathematical problem and the reasoning chain. The term "executable code" in this specification refers broadly to code that can be executed on a computer, i.e., code that can be directly parsed and executed by a computer system or runtime environment. For example, it can include one or more of compiled code, interpreted code, bytecode, and machine code.
[0051] The specific method for generating the executable code may vary in different embodiments. In one embodiment, a prompt word including the first mathematical problem, the reasoning chain, and a prompt portion indicating that the corresponding executable code is generated based on the first mathematical problem and the reasoning chain may be input into a preset model (e.g., a large model) to generate the first executable code. Figure 5 Schematic diagram of generating executable code according to an embodiment of this specification is shown. Figure 5 As shown, for example, the mathematical problem obtained in step S301 and the reasoning chain for the mathematical problem, as well as the prompt word of the prompt part b indicating that the corresponding executable code is generated based on the mathematical problem and the reasoning chain, can be input into the large model to obtain the executable code.
[0052] After obtaining the first executable code, the first executable code can be executed to obtain the second answer corresponding to the first mathematical problem. Furthermore, it can be determined whether the first answer and the second answer are consistent. If the first answer and the second answer are consistent, the first mathematical problem is used as a training sample, and the reasoning chain and the first answer are used as the training label corresponding to the training sample to fine-tune the target large model. In different embodiments, the specific method of fine-tuning the large model may be different according to the training sample and the training label, and this specification does not limit this. In one embodiment, all parameters of the target large model can be updated through the back propagation (BP) algorithm based on the above-mentioned training samples and training labels. In another embodiment, some parameters (such as low-rank parameters) or newly added parameters (such as newly added prompt embedding parameters or newly added inter-layer adapter parameters) of the target large model can also be updated through the back propagation algorithm based on the training sample and the training label.
[0053] Figure 6 A schematic diagram showing a method for obtaining training data for fine-tuning a large model according to another embodiment of this specification is shown. Figure 6 As shown, after executing the generated executable code, the first answer can be used as training data for fine-tuning the target large model based on whether the execution result of the executable code (i.e., the second answer) is consistent with the first answer in the first solution. If they are consistent, the first answer is used as training data for fine-tuning the target large model. Specifically, the first mathematical problem can be used as a training sample for fine-tuning, and the first answer can be used as the training label corresponding to the training sample.
[0054] In one embodiment, if the first answer and the second answer are inconsistent, fine-tuning the target large model through the reasoning chain and the first answer can be abandoned. Figure 6 In the example shown, if the first answer is inconsistent with the execution result of the executable code (i.e., the second answer), the first solution (including the reasoning chain and the first answer) is discarded and is not used as training data for fine-tuning the target large model.
[0055] In the aforementioned embodiment of generating executable code based on the first mathematical problem and the chain of reasoning using a large model, due to the instability of the content generated by the large model, in actual production scenarios, the calculation process of the code generated by the large model may be inconsistent with the chain of reasoning, thereby reducing the reliability of verifying the answer to the problem using this code. To improve the reliability of verifying the generated executable code, in one specific embodiment, it is possible to determine whether the first executable code and the chain of reasoning match. If the first executable code does match the chain of reasoning, the first executable code is executed to obtain the second answer corresponding to the first mathematical problem.
[0056] In different specific embodiments, the specific method of determining whether the first executable code and the reasoning chain match may be different. In a specific embodiment, the degree of match between the first executable code and the reasoning chain can be determined by a preset model; if the degree of match reaches a preset threshold, the executable code and the reasoning chain are determined to match; if the degree of match does not reach the preset threshold, the executable code and the reasoning chain are determined to not match. In different specific embodiments, the preset model is, for example, a large model, which is used to input the large model to determine the specific prompt words and preset thresholds of the degree of match between the first executable code and the reasoning chain. It can also be different specific prompt words and specific thresholds, and this specification does not limit this. In one example, for example, a prompt word containing the first executable code, the reasoning chain, and a prompt word indicating the prompt part for judging the degree of match between the first executable code and the reasoning chain can be input into the large model to obtain an executable code. In another example, the prompt word can also indicate the numerical range of the degree of match between the first executable code and the reasoning chain. In another specific embodiment, if the first executable code does not match the reasoning chain, a second executable code corresponding to the first mathematical problem may be generated to determine whether the second executable code matches the reasoning chain; if the second executable code matches the reasoning chain, the second executable code is executed to obtain a second answer corresponding to the second mathematical problem. Figure 7 FIG. 1 shows a schematic diagram of regenerating executable code according to an embodiment of the present specification. Figure 7 As shown, if it is determined that the first executable code does not match the reasoning chain, executable code corresponding to the first solution can be regenerated. Furthermore, it is re-determined whether the regenerated executable code matches the reasoning chain in the first solution. If so, the executable code is executed, and then, based on the execution result of the executable code and the first solution, it is determined whether the first solution should be used as training data for fine-tuning the target large model. If not, the above process of regenerating the executable code and determining whether it matches the reasoning chain in the first solution is repeated. Until the two match, it is determined whether the first solution should be used as training data for fine-tuning the target large model based on the execution result of the executable code that matches the reasoning chain in the first solution and the first solution.
[0057] In different embodiments, the preset model used to generate the mathematical problem in step S301 is similar to the preset model used to generate the executable code in step S303, and the preset model used to determine the matching degree between the executable code and the reasoning chain can be the same large language model as the target large model, or a different large language model from the target large model, or a pre-trained or optimized dedicated machine learning model. For example, in one example, the preset model used to generate the executable code can be a pre-trained code generation model. In different specific examples, the neural network structure of the pre-trained code generation model can be different, and this specification does not limit this.
[0058] Furthermore, in one embodiment, in step S301, the first mathematical problem may be input into the target macro model multiple times to obtain multiple first solutions corresponding to the first mathematical problem. Furthermore, in step S303, for each first solution, first executable code for determining the answer to the first mathematical problem may be generated based on the first mathematical problem and the reasoning chain within the first solution. The first executable code is executed to obtain a second solution corresponding to the first mathematical problem. Furthermore, it may be determined whether each first solution can be used as training data for fine-tuning the target. The specific determination method may be the same as the aforementioned determination of whether an individual first solution can be used as training data for fine-tuning the target, and will not be further described here.
[0059] Due to the instability of the content generated by the large model, the probability of a correct answer being present in multiple answers generated by the large model for the same math problem is significantly higher than the probability of a correct answer being present in a single answer generated by the large model for the same math problem. Therefore, in this embodiment, by sampling answers multiple times, the probability of obtaining high-quality answers can be increased. This, in turn, can further improve the quality of the training data used for fine-tuning the large model, enhancing the training effectiveness of fine-tuning the large model.
[0060] According to another aspect of the embodiment, a device for acquiring training data for fine-tuning a large model is also provided. Figure 8 FIG. 1 shows a structural diagram of a device for obtaining training data for fine-tuning a large model according to an embodiment of the present specification. Figure 8 As shown, the apparatus 800 includes:
[0061] An acquiring unit 81 is configured to acquire a first mathematical problem, input the first mathematical problem into a target macro model, and obtain a first solution corresponding to the first mathematical problem, wherein the first solution includes a reasoning chain for deriving an answer to the problem and the first solution;
[0062] The determination unit 82 is configured to generate a first executable code for determining an answer to the first mathematical problem based on the first mathematical problem and the reasoning chain, execute the first executable code to obtain a second answer corresponding to the first mathematical problem, and determine whether the first answer and the second answer are consistent. If the first answer and the second answer are consistent, use the first mathematical problem as a training sample, and the reasoning chain and the first answer as training labels corresponding to the training sample, for fine-tuning the target large model.
[0063] On another aspect, the present specification provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute any one of the above methods.
[0064] On another aspect, the present specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, any one of the above methods is implemented.
[0065] Yet another aspect of the present specification provides a computer program product, comprising a computer program / instruction, which implements any of the above methods when executed by a processor.
[0066] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0067] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0068] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0069] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flow charts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any particular order.
[0070] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0071] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0072] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0074] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0075] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0076] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0077] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0078] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0079] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.
[0080] The foregoing description is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. Those skilled in the art will appreciate that various modifications and variations of one or more embodiments of this specification are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification are intended to be included within the scope of the claims.
Claims
1. A method for obtaining training data for fine-tuning a large model, comprising: Obtaining a first mathematical problem, inputting the first mathematical problem into a target macro model, and obtaining a first answer corresponding to the first mathematical problem, wherein the first answer includes a reasoning chain for deriving an answer to the problem and a first answer; According to the first mathematical problem and the reasoning chain, a first executable code for determining the answer to the first mathematical problem is generated, and the first executable code is executed to obtain a second answer corresponding to the first mathematical problem; it is determined whether the first answer and the second answer are consistent, and if the first answer and the second answer are consistent, the first mathematical problem is used as a training sample, and the reasoning chain and the first answer are used as training labels corresponding to the training sample, for fine-tuning the target large model.
2. The method according to claim 1, wherein: Executing the first executable code to obtain a second answer corresponding to the first mathematical problem includes: Determine whether the first executable code matches the reasoning chain, and if the first executable code matches the reasoning chain, execute the first executable code to obtain a second answer corresponding to the first mathematical problem.
3. The method according to claim 2 further includes, if the first executable code does not match the reasoning chain, generating a second executable code corresponding to the first mathematical problem, and determining whether the second executable code and the reasoning chain match; if the second executable code matches the reasoning chain, executing the second executable code to obtain a second answer corresponding to the second mathematical problem.
4. The method according to claim 2, wherein: Determining whether the first executable code matches the inference chain includes: Determining the matching degree between the first executable code and the inference chain by using a preset model; If the matching degree reaches a preset threshold, it is determined that the executable code matches the reasoning chain; if the matching degree does not reach the preset threshold, it is determined that the executable code does not match the reasoning chain.
5. The method according to claim 1, wherein: According to the first mathematical problem and the chain of reasoning, generating a first executable code for determining an answer to the first mathematical problem, comprising: The prompt words including the first mathematical problem, the reasoning chain and the prompt part indicating that the corresponding executable code is generated based on the first mathematical problem and the reasoning chain are input into a preset model to obtain the first executable code.
6. The method according to claim 1 further comprises, if the first answer and the second answer are inconsistent, abandoning fine-tuning the target large model through the reasoning chain and the first answer.
7. The method according to claim 1, wherein: Get the first math problem, including randomly extracting a plurality of second math problems belonging to a target math field from a pre-acquired math problem set; The target prompt word is input into a preset model to obtain a first mathematical problem, wherein the target prompt word includes the plurality of second mathematical problems and a prompt part indicating that a mathematical problem belonging to a target mathematical field but different from the plurality of second mathematical problems is generated according to the plurality of second mathematical problems.
8. The method according to claim 1, wherein: Inputting a first mathematical problem into the target macro model to obtain a first answer corresponding to the first mathematical problem includes: Inputting the first mathematical problem into the target macro model multiple times to obtain multiple first answers corresponding to the first mathematical problem; Generating a first executable code for determining an answer to the first mathematical problem according to the first mathematical problem and the reasoning chain, and executing the first executable code to obtain a second answer corresponding to the first mathematical problem, including: For each first solution, a first executable code for determining the answer to the first mathematical problem is generated according to the first mathematical problem and the reasoning chain in the first solution, and the first executable code is executed to obtain a second answer corresponding to the first mathematical problem.
9. A device for acquiring training data for fine-tuning a large model, comprising: An acquisition unit is configured to acquire a first mathematical problem, input the first mathematical problem into a target macro model, and obtain a first answer corresponding to the first mathematical problem, wherein the first answer includes a reasoning chain for deriving an answer to the problem and a first answer; The determination unit is configured to generate a first executable code for determining an answer to the first mathematical problem based on the first mathematical problem and the reasoning chain, execute the first executable code to obtain a second answer corresponding to the first mathematical problem; determine whether the first answer and the second answer are consistent, and if the first answer and the second answer are consistent, use the first mathematical problem as a training sample, and the reasoning chain and the first answer as training labels corresponding to the training sample, for fine-tuning the target large model.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.
11. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 8 is implemented.