Sample generation method and device, model fine adjustment method and device, electronic equipment and medium
By generating multiple sample thinking chains and constructing sample weight charts, the thinking chain reasoning path of small models is optimized, and the problem of error accumulation in thinking chain reasoning is solved, and more efficient and accurate reasoning ability is achieved.
Patent Information
- Application Number
- CN202510713532.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-15
AI Technical Summary
Existing large models are prone to error accumulation in the process of thinking chain reasoning, resulting in insufficient inference accuracy and robustness, and existing training methods have data dependence and generalization problems.
By generating multiple sample thinking chains, building sample weight graphs, optimizing sample inference paths, and using large-model fine-tuning methods to improve the accuracy and robustness of thinking chains.
It improves the accuracy and robustness of small models in thinking chain reasoning, reduces the computing resource requirements, and enhances the model's adaptability in complex tasks.
Smart Images

Figure CN120494109A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence technology, particularly to the fields of large model technology and natural language processing technology. More specifically, the present disclosure provides a sample generation method, a model fine-tuning method, a question-and-answer information processing method, an apparatus, an electronic device, and a storage medium. Background Art
[0002] With the development of artificial intelligence technology, the application of large models is increasing. Large models can simulate the distributed reasoning process based on the Chain of Thought (CoT) to achieve efficient reasoning. Summary of the Invention
[0003] The present disclosure provides a sample generation method, a model fine-tuning method, a question-and-answer information processing method, an apparatus, a device, and a storage medium.
[0004] According to one aspect of the present disclosure, a sample generation method is provided, the method comprising: obtaining a plurality of sample thinking chains based on an initial thinking chain, wherein the sample thinking chains include a plurality of operations to be performed; obtaining a sample weight graph based on the plurality of sample thinking chains, wherein the sample weight graph includes a plurality of candidate nodes and at least one candidate edge, the candidate nodes represent operations to be performed, and the candidate weights for the candidate edges indicate the logical relationship between the two candidate nodes connected via the candidate edges; obtaining at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path based on the sample weight graph, wherein the sample reasoning path includes a plurality of sample nodes determined from the plurality of candidate nodes and at least one sample edge for connecting the plurality of sample nodes.
[0005] According to another aspect of the present disclosure, a model fine-tuning method is provided, the method comprising: inputting at least one sample inference path into a model to be fine-tuned to obtain at least one second optimized inference path; and fine-tuning the model to be fine-tuned according to the at least one first optimized inference path and the at least one second optimized inference path, wherein the at least one first optimized inference path corresponds one-to-one to the at least one sample inference path, and the at least one sample inference path and the first optimized inference path corresponding to the sample inference path are obtained according to the sample generation method provided by the present disclosure.
[0006] According to another aspect of the present disclosure, a method for processing question and answer information is provided, the method comprising: inputting question information into a target model to obtain a target thought chain for the question information; generating answer information for the question information based on the target thought chain, wherein the target model is fine-tuned according to the model fine-tuning method provided in the present disclosure.
[0007] According to another aspect of the present disclosure, a sample generation device is provided, which includes: a first acquisition module for obtaining multiple sample thinking chains based on an initial thinking chain, wherein the sample thinking chains include multiple operations to be executed; a second acquisition module for obtaining a sample weight graph based on the multiple sample thinking chains, wherein the sample weight graph includes multiple candidate nodes and at least one candidate edge, the candidate nodes represent operations to be executed, and the candidate weights for the candidate edges indicate the logical relationship between the two candidate nodes connected via the candidate edges; a third acquisition module for obtaining at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path based on the sample weight graph, wherein the sample reasoning path includes multiple sample nodes determined from multiple candidate nodes and at least one sample edge for connecting the multiple sample nodes.
[0008] According to another aspect of the present disclosure, a model fine-tuning device is provided, which includes: a fourth acquisition module, used to input at least one sample reasoning path into the model to be fine-tuned to obtain at least one second optimized reasoning path; a fine-tuning module, used to fine-tune the model to be fine-tuned according to at least one first optimized reasoning path and at least one second optimized reasoning path, wherein the at least one first optimized reasoning path corresponds one-to-one to the at least one sample reasoning path, and the at least one sample reasoning path and the first optimized reasoning path corresponding to the sample reasoning path are obtained according to the sample generation device provided by the present disclosure.
[0009] According to another aspect of the present disclosure, a question and answer information processing device is provided, which includes: a fifth acquisition module for inputting question information into a target model to obtain a target thinking chain for the question information; a generation module for generating answer information for the question information based on the target thinking chain, wherein the target model is fine-tuned according to the model fine-tuning device provided by the present disclosure.
[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.
[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.
[0012] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided according to the present disclosure when executed by a processor.
[0013] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0015] Figure 1 is a flow chart of a sample generation method according to one embodiment of the present disclosure;
[0016] Figure 2A is a schematic diagram of an initial thought chain according to one embodiment of the present disclosure;
[0017] Figure 2B is a schematic diagram of a first sample thought chain according to an embodiment of the present disclosure;
[0018] Figure 2C is a schematic diagram of a first sample thought chain according to one embodiment of the present disclosure;
[0019] Figure 2D is a schematic diagram of a third sample thought chain according to one embodiment of the present disclosure;
[0020] Figure 2E is a schematic diagram of a fourth sample thought chain according to one embodiment of the present disclosure;
[0021] Figure 3A is a schematic diagram of a sample inference graph according to one embodiment of the present disclosure;
[0022] Figure 3B is a schematic diagram of a sample weight graph according to an embodiment of the present disclosure;
[0023] Figure 4 is a schematic flow chart of a model fine-tuning method according to an embodiment of the present disclosure;
[0024] Figure 5 is a flowchart of a question-and-answer information processing method according to another embodiment of the present disclosure;
[0025] Figure 6 is a schematic block diagram of a sample generating apparatus according to an embodiment of the present disclosure;
[0026] Figure 7 is a schematic block diagram of a model fine-tuning device according to an embodiment of the present disclosure;
[0027] Figure 8 is a schematic block diagram of a question and answer information processing apparatus according to another embodiment of the present disclosure; and
[0028] Figure 9 This is a block diagram of an electronic device to which at least one of a sample generation method, a model fine-tuning method, and a question-and-answer information processing method can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] The large model can be a large language model (LLM), an image model, or a multimodal model. A multimodal model can perform inference or training based on one or more of multiple modalities, such as text, images, audio, and video.
[0031] Using prompts like "Let's think step by step" allows large models to reason efficiently through thought chains. Thought chains simulate the human step-by-step reasoning process, enabling models to better perform logical deduction, information integration, and multi-step reasoning on complex tasks. By explicitly unfolding the reasoning process, it reduces task complexity and minimizes error propagation during reasoning.
[0032] However, reasoning chains rely on long reasoning paths. An error in a single step in the chain can lead to a cumulative error propagation phenomenon: "One step wrong, all steps wrong, and ultimately wrong." An error in a single reasoning step in the chain could be caused by a single-step calculation error in the model, an incorrect reasoning sequence, or a logical break. Therefore, logical reasoning capabilities place great emphasis on the accuracy (single-step accuracy) and robustness (self-correction) of the model's generated reasoning chain.
[0033] To improve the accuracy of individual steps in a model's chain of reasoning, multiple strategies can be used to synthesize chains of reasoning to train or fine-tune the model. The first of these strategies could be to use a stronger teacher model to synthesize reasoning steps. The second strategy could be to train and iterate a stronger reward model to perform judgment and screening of intermediate steps to ensure their accuracy.
[0034] The first strategy favors the use of a stronger teacher model. This can include an industry-leading large language model, leveraging the teacher model's powerful generative capabilities to synthesize the intermediate steps of the thought chain. This strategy is essentially a process of capability data distillation. It assumes that as long as the teacher model is sufficiently powerful and reliable, the thought chain data generated by the teacher model will be largely accurate. The fine-tuned model gradually improves its own thought chain reasoning capabilities by imitating and learning from this high-quality thought chain data. However, this strategy suffers from data dependency and generalization issues. While the thought chain data generated by the teacher model is high-quality, it may be overly dependent on the teacher model's own knowledge structure and reasoning patterns. This can cause the fine-tuned model to overfit to the teacher model's style during learning and hinder generalization to a wider range of more diverse reasoning scenarios. Furthermore, no matter how powerful the teacher model is, it has its own knowledge boundaries and reasoning limitations. When faced with problems beyond the teacher model's capabilities, the thought chain data generated by the teacher model may contain errors or irrationalities. When learning from this data, the fine-tuned model may inadvertently inherit these errors, limiting its own reasoning capabilities.
[0035] The second strategy introduces an independent judgment model (i.e., a reward model) to assess the quality of each step in the chain of reasoning. This second strategy is essentially a post-verification measure. Its effectiveness is highly dependent on the accuracy and discriminative power of the reward model. When the reward model is sufficiently powerful, it can effectively identify and select high-quality reasoning steps, thereby ensuring the overall accuracy of the chain of reasoning. However, while the reward model aims to improve the accuracy of reasoning steps, it can also make misjudgments. If the reward model mistakenly judges a step as high-quality and retains it, this erroneous step can have a ripple effect throughout the chain of reasoning, causing errors to accumulate and amplify in subsequent steps. Furthermore, building an accurate and reliable reward model is challenging, requiring a large amount of labeled data and meticulous model tuning to ensure that the reward model accurately assesses the quality of reasoning steps. This not only increases the technical difficulty of implementation but also risks misjudgments due to data bias or model overfitting.
[0036] Therefore, in order to enable the model to generate an accurate chain of thought, the present disclosure provides a sample generation method and a model fine-tuning method. The sample generation method will be described below.
[0037] Figure 1 is a flow chart of a sample generation method according to an embodiment of the present disclosure.
[0038] like Figure 1 As shown, the method 100 may include operations S110 to S130.
[0039] In operation S110 , a plurality of sample thought chains are obtained according to the initial thought chain.
[0040] In the disclosed embodiment, the initial thought chain includes multiple initial operations. The initial operation can be a step in the initial thought chain. For example, the initial operation can be represented by an initial operation text.
[0041] In the disclosed embodiments, various processing can be performed on the initial thought chain to obtain a sample thought chain. The sample thought chain may include multiple pending operations. The pending operation may be a step in the sample thought chain. For example, the pending operation may be represented by a pending operation text.
[0042] In operation S120 , a sample weight map is obtained according to the plurality of sample thought chains.
[0043] In an embodiment of the present disclosure, a sample weight graph includes multiple candidate nodes and at least one candidate edge. A candidate node may represent an operation to be performed. The candidate weight for a candidate edge indicates the logical relationship between two candidate nodes connected via the candidate edge. For example, the logical relationship may indicate the probability of a candidate node being an upstream node, or may indicate the probability of another candidate node being a downstream node. The operation to be performed represented by the upstream node may be performed before the operation to be performed represented by the downstream node. For another example, a sample thinking chain may be represented by multiple candidate nodes. Candidate edges are established between multiple candidate nodes of different sample thinking chains, and the candidate weights of the candidate edges are determined to determine the sample weight graph.
[0044] In operation S130 , at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path are obtained according to the sample weight map.
[0045] In the disclosed embodiment, a sample inference path includes multiple sample nodes determined from multiple candidate nodes and at least one sample edge connecting the multiple sample nodes. For example, a candidate node may be randomly selected. Next, a candidate edge with a candidate weight within a randomly generated weight interval is selected from the multiple candidate edges connected to the selected candidate node. The two candidate nodes connected by the selected candidate edge may serve as two sample nodes, and the selected candidate edge may serve as a sample edge.
[0046] In the disclosed embodiments, various methods can be used to optimize the sample reasoning path. For example, a large model can be used to optimize the sample reasoning path to obtain a first optimized reasoning path. The large model can be the teacher model described above. For example, the large model can be the Wenxin large model.
[0047] Through the embodiments of the present disclosure, multiple sample thought chains can be generated based on the initial thought chain, which effectively improves the richness of the thought chain data. A sample weight graph is constructed according to the sample thought chain, and a sample reasoning path is determined from the sample weight graph. Thus, multiple different sample reasoning paths can be generated from an initial thought chain. The first optimized reasoning path corresponding to the sample reasoning path can be used as a label. Thus, the sample reasoning path and the first optimized reasoning path can be used to fine-tune and train smaller models, so that smaller-scale models can generate high-quality thought chains and achieve more efficient reasoning using fewer computing resources. It can be understood that compared with a large model that serves as a teacher model, a smaller model can be used as a student model.
[0048] It can be understood that the above describes the method of the present disclosure, and some ways of obtaining sample thought chains of the present disclosure will be further described below.
[0049] In some embodiments, in some implementations of the above-described operation S110, obtaining multiple sample thought chains based on the initial thought chain includes: performing at least one of multiple thought chain editing processes on the thought chain to be edited to obtain the sample thought chains. The thought chain to be edited is obtained based on the initial thought chain, and the thought chain to be edited includes multiple operations to be edited. The initial thought chain can be generated by a large model based on a problem. The problem can be a logic problem. For example, the logic problem can be a mathematical calculation problem, a decryption problem, etc.
[0050] For example, the initial thought chain can be edited multiple times to obtain a sample thought chain. During the first round of editing, the initial thought chain can be used as the thought chain to be edited. The thought chain to be edited in the first round can be edited to obtain the thought chain to be edited in the second round. After the editing end condition is met, a sample thought chain can be obtained. The editing end condition can be, for example, the completion of the predicted number of edits. The following will illustrate different thought chain processing methods in conjunction with an initial thought chain.
[0051] Figure 2A is a schematic diagram of an initial thought chain according to an embodiment of the present disclosure.
[0052] like Figure 2A As shown, question 20, "Given the sequence {a_n}, where a_n = frac{1}{f(n)}, where f(n) is the integer closest to sqrt{n}. If the sum of the first m terms of the sequence {a_n} is 20, what is m?", is provided to the large model, and the large model is asked to provide a chain of thought. It is understood that frac can represent a fraction (containing a numerator and denominator). sqrt represents the square root.
[0053] like Figure 2AAs shown, the initial thought chain cot20 may include multiple initial operations, including the following initial operations step201 to step2010.
[0054] In the initial operation step 201 , for each positive integer n, sqrt{n} is calculated.
[0055] In the initial operation step 202 , the integer f(n) closest to sqrt{n} is found. For example, when n=1, 2, sqrt{n} is closest to 1, so f(1)=f(2)=1.
[0056] In the initial operation step 203, when n=3, 4, 5, 6, sqrt{n} is closest to 2, so f(3)=f(4)=f(5)=f(6)=2.
[0057] In the initial operation step 204 , by analogy, it can be found that the value of f(n) increases as n increases, and each value of f(n) corresponds to a continuous range of n.
[0058] In the initial operation step 205, the terms of the sequence {a_n} are calculated based on the value of f(n). For example, when f(n) = 1, a_n = 1; when f(n) = 2, a_n = 1 / 2; and so on.
[0059] In the initial operation step 206 , starting from n=1, the items of the sequence {a_n} are accumulated until the accumulated sum reaches or exceeds 20.
[0060] In the initial operation step 207 , the number of items m when the cumulative sum reaches 20 is recorded, and the first 6 items are accumulated: 1+1+1 / 2+1 / 2+1 / 2+1 / 2=4.
[0061] In the initial operation step 208 , subsequent terms are calculated until the accumulated sum reaches 20.
[0062] In the initial operation step 209, through calculation, it can be found that when m=110, the sum of the first m terms of the sequence {a_n} is exactly 20.
[0063] In the initial operation step2010, the answer is m=110.
[0064] It is understandable that different large models may provide different thought chains, and the initial thought chain cot20 is only an example. The initial thought chain is described above, and multiple sample thought chains will be described below.
[0065] In some embodiments, multiple thought chain editing processes include operation deletion processes. The operation deletion process can delete at least one of the multiple operations to be edited in the thought chain to be edited. For example, the above-mentioned initial thought chain cot20 can be used as the thought chain to be edited. One or more initial operations in the initial thought chain cot20 can be deleted to obtain a first sample thought chain. Through the embodiment of the present disclosure, the operation deletion process is similar to the "cutting" process in gene editing, which realizes the "chain breaking process" to interrupt certain continuous reasoning steps in the initial thought chain to force the student model to reconnect these broken chains when training the student model. As a result, the model can not only learn how to process each reasoning operation independently, but also explore new connection methods and reasoning paths at the break, thereby enhancing its adaptability to complex, nonlinear reasoning tasks. The following will be combined with Figure 2B Further explain the first sample thinking chain.
[0066] Figure 2B is a schematic diagram of a first sample thought chain according to an embodiment of the present disclosure.
[0067] like Figure 2B As shown, the difference from the initial thinking chain cot20 is that the pending operation step217 of the first sample thinking chain cot21 is empty. In addition, the pending operation step211, the pending operation step212, the pending operation step213, the pending operation step214, the pending operation step215, the pending operation step216, the pending operation step218, the pending operation step219 and the pending operation step2110 of the first sample thinking chain are respectively the same as the initial operation step201, the initial operation step202, the initial operation step203, the initial operation step204, the initial operation step205, the initial operation step206, the initial operation step208, the initial operation step209 and the initial operation step2010, and are not repeated herein.
[0068] It can be understood that the above description of the present disclosure is based on the example of deleting an operation to be edited, but the present disclosure is not limited to this. Multiple operations to be executed can be deleted. For example, if the above-mentioned initial thinking chain cot20 is used as the thinking chain to be edited, multiple non-continuous initial operations can be determined from multiple initial operations as multiple operations to be deleted, and the multiple operations to be deleted can be deleted to obtain a sample thinking chain. Through the embodiment of the present disclosure, deleting non-continuous operations to be edited can effectively reduce the difficulty of the student model to reconstruct a complete thinking chain, effectively reduce training costs, and improve training efficiency.
[0069] It can be understood that the above describes the operation deletion processing in the present disclosure, and the following will describe some other thought chain editing processing methods in the present disclosure.
[0070] In some embodiments, the editing process of multiple thought chains also includes an operation sequence adjustment process. The operation sequence adjustment process can adjust the execution order of multiple operations to be edited. For example, the above-mentioned initial thought chain cot20 can be used as the thought chain to be edited. From the multiple initial operations of the initial thought chain cot20, one or more initial operations are selected, and the order of these selected initial operations is adjusted to obtain a second sample thought chain. Through the embodiment of the present disclosure, the reasoning steps in the thought chain are processed out of order, and their original logical order is disrupted, so that the student model learns to recognize and reorganize these disordered elements during the training process to form a new and reasonable reasoning sequence. In this way, the logical reconstruction ability of the model can be greatly improved, so that the model can still maintain the accuracy and coherence of reasoning when faced with inputs with missing information or disordered order. The following will be combined with Figure 2C Provide explanation.
[0071] Figure 2C is a schematic diagram of a first sample thought chain according to an embodiment of the present disclosure.
[0072] like Figure 2C As shown, the second sample thought chain cot22 includes multiple pending operations. The multiple pending operations include pending operations step_uk221, pending operations step222, pending operations step_uk223, pending operations step224, pending operations step_uk225, pending operations step226, pending operations step_uk227, pending operations step228, pending operations step229, and pending operations step2210. The pending operations step222, pending operations step224, pending operations step226, pending operations step228, pending operations step229, and pending operations step2210 are the same as the initial operations step202, initial operations step204, initial operations step206, initial operations step208, initial operations step209, and initial operations step2010 described above, and are not further described in this disclosure.
[0073] like Figure 2CAs shown, the operation to be executed step_uk221 is "[step unknow] When n=3,4,5,6, sqrt{n} is closest to 2, so f(3)=f(4)=f(5)=f(6)=2". The operation to be executed step_uk223 is "[step unknow] For each positive integer n, calculate sqrt{n}". The operation to be executed step_uk225 is "[step unknow] Record the number of terms m when the cumulative sum reaches 20, and add the first 6 terms: 1+1+1 / 2+1 / 2+1 / 2+1 / 2=4". The operation to be executed step_uk227 is "[step unknow] Calculate each term of the sequence {a_n} based on the value of f(n). For example, when f(n)=1, a_n=1; when f(n)=2, a_n=1 / 2; and so on". The pending operations step_uk221 , step_uk223 , step_uk225 , and step_uk227 have a sequence unknown symbol [step unknow] to instruct the model to redetermine the execution sequence of the pending operations step_uk221 , step_uk223 , step_uk225 , and step_uk227 .
[0074] It can be understood that the above describes the operation sequence adjustment process in the present disclosure, and the following will describe other thought chain editing processes in the present disclosure.
[0075] In some embodiments, multiple thought chain editing processes also include thought chain reorganization processes. Thought chain reorganization processes are used to generate at least one edited operation that is different from at least one of the multiple operations to be edited based on the thought chain to be edited. For example, the calculation method of the operation to be edited can be changed to obtain the edited operation. It can be understood that after executing the thought chain reorganization process, the calculation result of the new thought chain obtained is the same as the calculation result of the thought chain to be edited. Through the embodiment of the present disclosure, in the process of processing the thought chain, the reasoning elements or steps of the thought chain are reorganized to create a new thought chain, which can not only enrich the reasoning strategy library of the model, but also enable the model to flexibly switch between a variety of reasoning modes to adapt to the needs of different tasks. The following will be combined with Figure 2D Further explanation is provided.
[0076] Figure 2D is a schematic diagram of a third sample thought chain according to an embodiment of the present disclosure.
[0077] The initial thought chain cot20 is used as the thought chain to be edited. Based on the initial thought chain cot20, the large model is made to perform thought chain reorganization processing to obtain the third sample thought chain cot23.
[0078] like Figure 2D As shown, the multiple pending operations of the third sample thinking chain cot23 are different from the multiple initial operations of the initial thinking chain cot20. In the pending operation step231, for f(n)=1, there are 2 items, and the sum is 2×1 / 1=2. In the pending operation step232, for f(n)=2, there are 4 items, and the sum is 4×1 / 2=2. In the pending operation step233, for f(n)=3, there are 6 items, and the sum is 6×1 / 3=2. In the pending operation step234, for f(n)=4, there are 8 items, and the sum is 8×1 / 4=2. In the pending operation step235, for f(n)=5, there are 10 items, and the sum is 10×1 / 5=2. In the pending operation step236, for f(n)=6, there are 12 items, and the sum is 12×1 / 6=2. In pending operation step 237, we accumulate these sums and find that for each additional group of f(n) = j, the sum increases by 2. Therefore, in pending operation step 238, we need to accumulate 10 groups (i.e., j = 1 to j = 10) to reach a sum of 20. In pending operation step 239, the total number of items, m, is the sum of the number of items in each group, i.e., 2 + 4 + 6 + 8 + 10 + 12 + 14 + 16 + 18 + 20 = 110.
[0079] It can be understood that the above describes the thought chain reorganization process of the present disclosure, and the following will describe other thought chain editing processes of the present disclosure.
[0080] In some embodiments, multiple thought chain editing processes may also include operation insertion processes. The operation insertion process may insert at least one operation to be added into the thought chain to be edited. The operation to be added may be an operation irrelevant to the problem, an operation related to an error, or an irrelevant and wrong operation. Through the embodiments of the present disclosure, the operation to be added may not directly participate in or affect the final reasoning result, but it can simulate the uncertainty and interference of information in the real world, and it can also enable the model to gradually identify and filter these irrelevant information during the training process, while maintaining accurate tracking of the core reasoning path. The model can be given the ability to self-diagnose and correct errors, so that when the model encounters unexpected input or logical contradictions, it can automatically adjust the reasoning strategy to maintain the accuracy and consistency of the output. The following will be combined with Figure 2E Provide explanation.
[0081] Figure 2E is a schematic diagram of a fourth sample thought chain according to an embodiment of the present disclosure.
[0082] The initial thought chain cot20 can be used as the thought chain to be edited. The operations to be added, "When n = 7, 8, sqrt{n} is closest to 100, so f(7) = f(8) = 2," and "f(n) = 1 / 2 * sqrt{n}, the value of f(n) needs to be recalculated," are added to the initial thought chain cot20 to obtain the fourth sample thought chain cot24.
[0083] like Figure 2E As shown, the fourth sample thought chain cot24 includes multiple pending operations. The multiple pending operations include pending operations step244 and step247, which are consistent with the two pending operations to be added. Pending operations step241, step242, step243, step245, step246, step248, step249, step2410, step2411, and step2412 are the same as the initial operations step201, step202, step203, step204, step205, step206, step207, step208, step209, and step2010 described above, and are not further described in this disclosure.
[0084] It can be understood that some implementation methods of obtaining a sample reasoning path are described above, and some methods of obtaining a sample weight map will be described below.
[0085] In some embodiments, in some implementations of the above operation S120, obtaining a sample weight graph according to multiple sample thought chains includes: obtaining a sample reasoning graph according to multiple sample thought chains. The sample reasoning graph includes multiple candidate nodes and at least one candidate edge. Figure 3A and Figure 3B Further explanation is provided.
[0086] Figure 3A is a schematic diagram of a sample inference graph according to one embodiment of the present disclosure.
[0087] In some embodiments, multiple initial reasoning graphs may be determined based on multiple sample thought chains. Figure 3AAs shown, taking the above-mentioned first sample thinking chain cot21 and the third sample thinking chain cot23 as examples, multiple candidate nodes representing multiple operations to be executed of the first sample thinking chain cot21 can be determined. The multiple candidate nodes for representing multiple operations to be executed of the first sample thinking chain cot21 include candidate nodes N311 to candidate nodes N3110, which can respectively represent the above-mentioned operations to be executed step211 to step2110. Multiple candidate nodes for representing multiple operations to be executed of the third sample thinking chain cot23 can also be determined. The multiple candidate nodes for representing multiple operations to be executed of the third sample thinking chain cot23 include candidate nodes N331 to candidate nodes N339, which can respectively represent the above-mentioned operations to be executed step231 to step239.
[0088] Next, a candidate edge connecting two different candidate nodes can be established. As shown in Figure 3, a candidate edge can be established from candidate node N311 to candidate node N312, a candidate edge from candidate node N312 to candidate node N313, a candidate edge from candidate node N313 to candidate node N314, a candidate edge from candidate node N314 to candidate node N315, a candidate edge from candidate node N315 to candidate node N316, a candidate edge from candidate node N316 to candidate node N317, a candidate edge from candidate node N317 to candidate node N318, a candidate edge from candidate node N318 to candidate node N319, and a candidate edge from candidate node N319 to candidate node N3110. Thus, the initial reasoning graph corresponding to the first sample thought chain cot21 can be determined. A candidate edge can be established from candidate node N331 to candidate node N332, a candidate edge from candidate node N332 to candidate node N333, a candidate edge from candidate node N333 to candidate node N334, a candidate edge from candidate node N334 to candidate node N335, a candidate edge from candidate node N335 to candidate node N336, a candidate edge from candidate node N336 to candidate node N337, a candidate edge from candidate node N337 to candidate node N338, and a candidate edge from candidate node N338 to candidate node N339. Thus, the initial reasoning graph corresponding to the third sample thought chain cot23 can be determined.
[0089] Next, we can establish candidate edges between multiple candidate nodes of different initial reasoning graphs to obtain sample reasoning graphs. Figure 3AAs shown, a candidate edge can be established from candidate node N311 to candidate node N332, a candidate edge from candidate node N312 to candidate node N333, a candidate edge from candidate node N313 to candidate node N334, a candidate edge from candidate node N314 to candidate node N335, a candidate edge from candidate node N315 to candidate node N336, a candidate edge from candidate node N316 to candidate node N337, a candidate edge from candidate node N317 to candidate node N338, and a candidate edge from candidate node N318 to candidate node N339.
[0090] It will be understood that the above-described method for establishing candidate edges between candidate nodes in different initial reasoning graphs is merely an example, and the present disclosure is not limited thereto. In other embodiments, candidate edges between candidate nodes in different initial reasoning graphs may be randomly established. Alternatively, multiple candidate edges may be established from a candidate node in one initial reasoning graph to all candidate nodes in another reasoning graph to achieve full connectivity of the candidate nodes. Thus, two different candidate nodes in a sample reasoning graph are connected via one or more candidate edges. A candidate edge is a directed edge. A candidate edge points from a candidate node serving as an upstream node in two different candidate nodes to a candidate node serving as a downstream node in two different candidate nodes. For example, a candidate edge may be established from candidate node 233 to candidate node N312. Two candidate edges exist between candidate node 233 and candidate node N312. For a candidate edge from candidate node N312 to candidate node N333, candidate node N312 may serve as the upstream node, and candidate node N333 may serve as the downstream node. For the candidate edge pointing from candidate node N333 to candidate node N312, candidate node N312 can serve as a downstream node, and candidate node N333 can serve as an upstream node.
[0091] Understandably, for the sake of simplicity, Figure 3A The sample reasoning graph shown shows multiple candidate nodes corresponding to the first sample thought chain and the third sample thought chain. The sample reasoning graph may also include candidate nodes representing pending operations of other sample thought chains. Through the embodiment of the present disclosure, a sample reasoning graph is constructed based on the sample thought chains, which helps to generate richer and more diverse reasoning paths, can reduce the computing resources required to generate sample reasoning paths, and can also improve the generation efficiency of sample reasoning paths.
[0092] Next, in some embodiments of the above operation S120, at least one candidate weight for at least one candidate edge can be determined to obtain a sample weight graph. As mentioned above, the candidate weight indicates the logical relationship between two candidate nodes connected via the candidate edge. The logical relationship may include causal consistency. The candidate weight may include a causal consistency weight (CausalScore). For example, the causal consistency weight may represent the probability that the operation to be performed represented by the upstream node is the cause, or it may represent the probability that the operation to be performed represented by the downstream node is the result. The causal consistency weight may also be called a causal consistency evaluation value, which can be determined by a large model based on two operations to be performed represented by two candidate nodes. After determining the candidate weights of each of the multiple candidate edges, a sample weight graph can be obtained. The following will be combined with Figure 3B Provide explanation.
[0093] Figure 3B is a schematic diagram of a sample weight graph according to an embodiment of the present disclosure.
[0094] like Figure 3B As shown, Figure 3A Unlike the sample inference graph shown, the sample weight graph includes multiple candidate weights for multiple candidate edges. Taking candidate node N311 as an example, candidate edge e312 is a directed edge from candidate node N311 to candidate node N312, and candidate edge e313 is a directed edge from candidate node N311 to candidate node N332. The candidate weight for candidate edge e312 can be 0.7, and the candidate weight for candidate edge e313 can be 0.1.
[0095] It can be understood that the above describes some methods of determining a sample weight graph in the present disclosure, and the following describes some methods of determining a sample reasoning path.
[0096] In some embodiments, in some implementations of the above operation S130, obtaining at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path according to the sample weight map includes: determining at least one sample reasoning path according to at least one candidate weight.
[0097] In some embodiments, determining at least one sample reasoning path based on at least one candidate weight includes: determining a first-level sample node from a plurality of candidate nodes. For example, Figure 3B The first-level sample node is determined in the sample weight graph shown. In one example, a node can be randomly selected from candidate node N311 and candidate node N331 as the first-level sample node. In another example, a node can also be randomly selected from all candidate nodes as the first-level sample node. The following description will use candidate node N311 as the first-level sample node as an example.
[0098] In some embodiments, determining at least one sample inference path based on at least one candidate weight includes: determining a k+1th level sample node from at least one candidate downstream node of the kth level sample node based on the kth level sample node and at least one kth level candidate weight for at least one kth level candidate edge. Determining the kth level candidate edge pointing from the kth level sample node to the k+1th level sample node as the kth level sample edge. The kth level candidate edge points from the kth level sample node to a candidate downstream node of the kth level sample node in the sample weight graph. k is an integer greater than or equal to 1. Taking k=1 as an example, as described above, candidate node N311 is determined as a first-level sample node. Among the multiple candidate nodes, candidate nodes connected to candidate node N311 via candidate edges can be selected as candidate downstream nodes of candidate node N311. In this case, candidate edges e312 and e313 can each be selected as first-level candidate edges. A random weight can be randomly generated. If the random weight is 0.65, a random weight range can be determined based on the random weight. The maximum value of the random weight interval can be 0.75, and the minimum value can be 0.55. Based on the candidate weight interval, it can be determined that the candidate weight for the candidate edge e312 is in the random weight interval. Next, the candidate node N312 pointed to by the candidate edge e312 can be used as the second-level sample node. The candidate edge e312 can be used as the first-level sample edge. Through the embodiment of the present disclosure, random weighted walks can be implemented in the sample weight graph, and a large number of causal reasoning paths can be determined from the sample weight graph to fully train the model.
[0099] In some embodiments, determining at least one sample reasoning path based on at least one candidate weight further includes: in response to determining that the k+1th sample node does not have a candidate downstream node, determining a sample reasoning path based on k+1 sample nodes and k sample edges, the sample reasoning path including k+1 sample nodes and k sample edges, the k+1 sample node including the kth level sample node and the k+1th level sample node, and the k sample edges including the kth level sample edge. For example, if candidate node N311, candidate node N312, candidate node N313, candidate node N314, candidate node N315, candidate node N316, candidate node N317, candidate node N318, and candidate node N319 have been determined as sample nodes, candidate node N3110 can be determined as a level 10 candidate node based on candidate node N319 being a level 9 sample node. Figure 3BAs shown, candidate node N3110 does not have a candidate downstream node, and candidate node N311, candidate node N312, candidate node N313, candidate node N314, candidate node N315, candidate node N316, candidate node N317, candidate node N318, candidate node N319, and candidate node N3110 serve as the 1st level sample node to the 10th level sample node respectively. The candidate edges between these candidate nodes can be used as the 1st level sample edge to the 9th level sample edge, and a sample reasoning path can be obtained.
[0100] It can be understood that some methods for determining a sample reasoning path are described above, and some methods for determining a first optimized reasoning path will be described below.
[0101] In some embodiments, at least one sample reasoning path is adjusted using a large model to obtain at least one first optimized reasoning path corresponding to the at least one sample reasoning path.
[0102] For example, taking the above-mentioned sample reasoning path with candidate nodes N311, N312, N313, N314, N315, N316, N317, N318, N319, and N3110 as multiple sample nodes as an example, the large model can generate an operation represented by candidate node N317 to restore the deleted pending operation and obtain a first optimized reasoning path. Through the embodiments of the present disclosure, the large model can accurately determine the first optimized reasoning path corresponding to the sample reasoning path, so as to transfer the reasoning capability of the large model to the smaller model, thereby reducing the labor cost required for manual optimization and improving the optimization efficiency of the reasoning path.
[0103] For example, taking the sample reasoning path of multiple candidate nodes representing the operations to be executed step_uk221, the operations to be executed step222, the operations to be executed step_uk223, the operations to be executed step224, the operations to be executed step_uk225, the operations to be executed step226, the operations to be executed step_uk227, the operations to be executed step228, the operations to be executed step229 and the operations to be executed step2210 as multiple sample nodes, the large model can generate multiple reordered operations to form a reasonable reasoning path and obtain the first optimized reasoning path.
[0104] For example, taking the sample reasoning path of multiple candidate nodes representing operations to be executed step241, operations to be executed step242, operations to be executed step243, operations to be executed step244, operations to be executed step245, operations to be executed step246, operations to be executed step247, operations to be executed step248, operations to be executed step249, operations to be executed step2410, operations to be executed step2411 and operations to be executed step2412 as multiple sample nodes, the large model can delete one or more irrelevant or erroneous operations to form an accurate reasoning path and obtain the first optimized reasoning path.
[0105] It can be understood that the above describes the method for obtaining the first optimized reasoning path, and the sample generation method disclosed in the present invention will be further described below.
[0106] In some embodiments, the method 100 further includes determining a first target reasoning path based on the sample weight graph. For example, the optimal reasoning path can be determined from the sample weight graph using the large model as the first target reasoning path. Through the disclosed embodiments, a superior or optimal reasoning path can be found from a complex sample weight graph, thereby further transferring the reasoning capabilities of the large model to smaller models.
[0107] It can be understood that the above describes the sample generation method of the present disclosure, and the following describes the model fine-tuning method of the present disclosure.
[0108] Figure 4 is a schematic flowchart of a model fine-tuning method according to an embodiment of the present disclosure.
[0109] like Figure 4 As shown, method 400 may include operations S410 to S420.
[0110] In operation 410 , at least one sample inference path is input into a model to be fine-tuned to obtain at least one second optimized inference path.
[0111] In the embodiment of the present disclosure, the model to be fine-tuned can optimize the sample reasoning path. Compared with the large model that generates the sample reasoning path and the first optimized reasoning path, the model to be fine-tuned can be a student model with a smaller scale than the large model.
[0112] In operation S420 , the to-be-fine-tuned model is fine-tuned according to the at least one first optimized reasoning path and the at least one second optimized reasoning path.
[0113] In the embodiment of the present disclosure, the parameters of the to-be-fine-tuned model may be fine-tuned to reduce the difference between the first optimized reasoning path and the second optimized reasoning path corresponding to the same sample reasoning path.
[0114] In the embodiment of the present disclosure, at least one first optimized reasoning path corresponds one-to-one with at least one sample reasoning path. The at least one sample reasoning path and the first optimized reasoning path corresponding to the sample reasoning path are obtained according to method 100 .
[0115] Through the embodiments of the present disclosure, the model can be fine-tuned based on a large number of sample reasoning paths and the corresponding first optimized reasoning paths generated by the large model, so that the fine-tuned model can have a stronger and more accurate ability to generate thought chains, so as to perform efficient and accurate reasoning and efficiently solve logical problems.
[0116] It can be understood that the above describes the model fine-tuning method of the present disclosure, and the following will further describe the model fine-tuning method of the present disclosure.
[0117] In some embodiments, fine-tuning the model to be fine-tuned based on at least one first optimized reasoning path and at least one second optimized reasoning path includes: inputting the sample weight map into the model to be fine-tuned to obtain a second target reasoning path. For example, the model to be fine-tuned can be used to determine an optimal reasoning path from the sample weight map, which serves as the second target reasoning path. It is understood that due to the different structures and parameter counts of the model to be fine-tuned and the large model, there may be differences between the second target reasoning path and the first target reasoning path determined by the model to be fine-tuned and the large model, respectively.
[0118] In some embodiments, fine-tuning the model to be fine-tuned based on the at least one first optimized reasoning path and the at least one second optimized reasoning path includes fine-tuning the model to be fine-tuned based on a first target reasoning path, a second target reasoning path, the at least one first optimized reasoning path, and the at least one second optimized reasoning path. The first target reasoning path is determined based on the sample weight map.
[0119] For example, at least one structural loss is determined based on at least one first optimized reasoning path and at least one second optimized reasoning path. The structural loss can be determined based on the first optimized reasoning path and the second optimized reasoning path corresponding to the same sample reasoning path. In one example, the first optimized reasoning path may include multiple first nodes. The second optimized reasoning path may include multiple second nodes. Based on the multiple first operations represented by the multiple first nodes and the multiple second operations represented by the multiple second nodes, multiple first texts for representing the multiple first operations and multiple second texts for representing the multiple second operations can be obtained. Based on the multiple first texts and the multiple second texts, the edit distance between the first optimized reasoning path and the second optimized reasoning path can be determined. Based on the edit distance, the structural loss can be determined.
[0120] For example, causal loss can be determined based on the first target reasoning path and the second target reasoning path. In one example, a first causal consistency evaluation value for the first target reasoning path can be determined using a large model, and a second causal consistency evaluation value for the second target reasoning path can also be determined using the large model. The causal loss can be determined based on the difference between the first causal consistency evaluation value and the second causal consistency evaluation value.
[0121] For example, the model to be fine-tuned is fine-tuned based on the causal loss and at least one structural loss. In one example, a total loss can be determined based on a weighted sum of the causal loss and the structural loss. The model to be fine-tuned is fine-tuned based on the total loss. The weights for the causal loss and the structural loss can be preset.
[0122] It will be appreciated that the above description describes the model fine-tuning method disclosed herein. After fine-tuning the target model a preset number of times, or after at least one of the structural loss, causal loss, and total loss is less than or equal to a preset threshold, a target model can be obtained. The target model can process question-and-answer information. The question-and-answer information processing method disclosed herein will be described below.
[0123] Figure 5 is a flowchart of a question and answer information processing method according to another embodiment of the present disclosure.
[0124] like Figure 5 As shown, the method 500 may include operations S510 to S520.
[0125] In operation S510, question information is input into a target model to obtain a target thinking chain for the question information.
[0126] In an embodiment of the present disclosure, the target model is fine-tuned according to the method provided by the present disclosure. For example, the target model is fine-tuned according to the above method 400.
[0127] In the disclosed embodiment, the question information may be a question text provided by the user. The question text may be the text of a logic question. For example, the question text may be the text of a mathematical calculation question or the text of a decryption question.
[0128] In the embodiment of the present disclosure, the target thought chain includes multiple target operations, which may include target operation texts.
[0129] In operation S520, answer information for the question information is generated according to the target thought chain.
[0130] In the embodiment of the present disclosure, multiple target operations indicated by multiple target operation texts may be executed in sequence to obtain answer information.
[0131] It can be understood that the method of the present disclosure is described above, and the device of the present disclosure will be described below.
[0132] Figure 6 is a schematic block diagram of a sample generating apparatus according to an embodiment of the present disclosure.
[0133] like Figure 6 As shown, the apparatus 600 may include a first obtaining module 610 , a second obtaining module 620 , and a third obtaining module 630 .
[0134] The first obtaining module 610 is used to obtain multiple sample thought chains according to the initial thought chain. The sample thought chains include multiple operations to be executed.
[0135] The second obtaining module 620 is configured to obtain a sample weight graph based on the multiple sample thought chains. The sample weight graph includes multiple candidate nodes and at least one candidate edge. The candidate nodes represent operations to be performed, and the candidate weights for the candidate edges indicate the logical relationship between two candidate nodes connected by the candidate edge.
[0136] The third obtaining module 630 is configured to obtain, based on the sample weight graph, at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path. The sample reasoning path includes a plurality of sample nodes determined from the plurality of candidate nodes and at least one sample edge connecting the plurality of sample nodes.
[0137] In some embodiments, the first obtaining module includes an editing processing submodule configured to perform at least one of a plurality of thought chain editing processes on a thought chain to be edited to obtain a sample thought chain. The thought chain to be edited is obtained based on the initial thought chain, and the thought chain to be edited includes a plurality of operations to be edited.
[0138] In some embodiments, the multiple thought chain editing processes include an operation deletion process, an operation sequence adjustment process, a thought chain reorganization process, and an operation insertion process. The operation deletion process is used to delete at least one of the multiple operations to be edited in the thought chain to be edited. The operation sequence adjustment process is used to adjust the execution order of the multiple operations to be edited. The thought chain reorganization process is used to generate at least one edited operation that is different from at least one of the multiple operations to be edited based on the thought chain to be edited. The operation insertion process is used to insert at least one operation to be added into the thought chain to be edited.
[0139] In some embodiments, the second obtaining module includes: a first obtaining submodule for obtaining a sample reasoning graph based on multiple sample thought chains. The sample reasoning graph includes multiple candidate nodes and at least one candidate edge. A first determining submodule for determining at least one candidate weight for at least one candidate edge to obtain a sample weight graph.
[0140] In some embodiments, two different candidate nodes in a sample inference graph are connected via at least one candidate edge, where the candidate edge is a directed edge, and the candidate edge points from one candidate node serving as an upstream node to another candidate node serving as a downstream node. The logical relationship includes causal consistency, and the candidate weight includes a causal consistency weight.
[0141] In some embodiments, the third obtaining module includes: a second determining submodule configured to determine at least one sample inference path based on at least one candidate weight; and an adjusting submodule configured to adjust the at least one sample inference path using the large model to obtain at least one first optimized inference path corresponding to the at least one sample inference path.
[0142] In some embodiments, the second determination submodule includes: a first determination unit for determining a first-level sample node from a plurality of candidate nodes. A second determination unit for determining a k+1-th level sample node from at least one candidate downstream node of the k-th level sample node based on the k-th level sample node and at least one k-th level candidate weight for at least one k-th level candidate edge. The k-th level candidate edge points from the k-th level sample node to a candidate downstream node of the k-th level sample node in the sample weight graph, where k is an integer greater than or equal to 1. A third determination unit for determining the k-th level candidate edge pointing from the k-th level sample node to the k+1-th level sample node as the k-th level sample edge.
[0143] In some embodiments, the second determination submodule further includes: a fourth determination unit configured to, in response to determining that the k+1th sample node does not have a candidate downstream node, determine a sample reasoning path based on the k+1 sample nodes and the k sample edges. The sample reasoning path includes the k+1 sample nodes and the k sample edges, the k+1 sample nodes including the kth-level sample nodes and the k+1th-level sample nodes, and the k sample edges including the kth-level sample edges.
[0144] In some embodiments, the method further includes: a determination module configured to determine a first target reasoning path based on the sample weight map.
[0145] Figure 7 is a schematic block diagram of a model fine-tuning apparatus according to an embodiment of the present disclosure.
[0146] like Figure 7 As shown, the apparatus 700 may include a fourth obtaining module 710 and a fine-tuning module 720 .
[0147] The fourth obtaining module 710 is configured to input at least one sample reasoning path into the model to be fine-tuned to obtain at least one second optimized reasoning path.
[0148] The fine-tuning module 720 is configured to fine-tune the model to be fine-tuned according to at least one first optimized reasoning path and at least one second optimized reasoning path.
[0149] In an embodiment of the present disclosure, at least one first optimized reasoning path corresponds one-to-one to at least one sample reasoning path, and the at least one sample reasoning path and the first optimized reasoning path corresponding to the sample reasoning path are obtained according to the sample generation method provided by the present disclosure.
[0150] In some embodiments, the fine-tuning module includes a second obtaining submodule configured to input the sample weight map into the model to be fine-tuned to obtain a second target reasoning path. A fine-tuning submodule configured to fine-tune the model to be fine-tuned based on the first target reasoning path, the second target reasoning path, the at least one first optimized reasoning path, and the at least one second optimized reasoning path. The first target reasoning path is determined based on the sample weight map.
[0151] In some embodiments, the fine-tuning submodule includes: a fifth determination unit configured to determine at least one structural loss based on at least one first optimized reasoning path and at least one second optimized reasoning path; a sixth determination unit configured to determine a causal loss based on the first target reasoning path and the second target reasoning path; and a fine-tuning unit configured to fine-tune the model to be fine-tuned based on the causal loss and the at least one structural loss.
[0152] Figure 8 is a schematic block diagram of a question and answer information processing apparatus according to another embodiment of the present disclosure.
[0153] like Figure 8 As shown, the apparatus 800 may include a fifth obtaining module 810 and a generating module 820 .
[0154] The fifth obtaining module 810 is used to input the problem information into the target model to obtain a target thinking chain for the problem information.
[0155] The generation module 820 is used to generate answer information for the question information according to the target thinking chain.
[0156] In the embodiment of the present disclosure, the target model is fine-tuned according to the method provided by the present disclosure.
[0157] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0158] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0159] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0160] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. Computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.
[0161] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0162] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as at least one of the sample generation method, the model fine-tuning method, and the question-and-answer information processing method. For example, in some embodiments, at least one of the sample generation method, the model fine-tuning method, and the question-and-answer information processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of at least one of the sample generation method, model fine-tuning method, and question-and-answer information processing method described above may be performed. Alternatively, in other embodiments, computing unit 901 may be configured to perform at least one of the sample generation method, model fine-tuning method, and question-and-answer information processing method via any other suitable means (e.g., via firmware).
[0163] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0164] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0165] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0167] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0168] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0169] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0170] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A sample generation method, comprising: According to the initial thought chain, a plurality of sample thought chains are obtained, wherein the sample thought chains include a plurality of operations to be executed; Obtaining a sample weight graph based on the plurality of sample thought chains, wherein the sample weight graph includes a plurality of candidate nodes and at least one candidate edge, the candidate nodes representing the operations to be performed, and the candidate weights for the candidate edges indicating a logical relationship between two candidate nodes connected via the candidate edges; At least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path are obtained according to the sample weight graph, wherein the sample reasoning path includes a plurality of sample nodes determined from the plurality of candidate nodes and at least one sample edge for connecting the plurality of sample nodes.
2. The method according to claim 1, wherein The method of obtaining a plurality of sample thought chains based on the initial thought chain includes: The thought chain to be edited is subjected to at least one of a plurality of thought chain editing processes to obtain the sample thought chain, wherein the thought chain to be edited is obtained based on the initial thought chain, and the thought chain to be edited includes a plurality of operations to be edited.
3. The method according to claim 2, wherein: The plurality of thought chain editing processes include operation deletion processing, operation sequence adjustment processing, thought chain reorganization processing and operation insertion processing. The operation deletion process is used to delete at least one of the multiple operations to be edited in the thought chain to be edited; The operation sequence adjustment process is used to adjust the execution sequence of the plurality of operations to be edited; The thought chain reorganization process is used to generate at least one edited operation that is different from at least one of the plurality of operations to be edited according to the thought chain to be edited; The operation insertion process is used to insert at least one operation to be added into the thought chain to be edited.
4. The method according to claim 1, wherein The sample weight graph obtained according to the plurality of sample thought chains includes: Obtaining a sample reasoning graph according to the plurality of sample thought chains, wherein the sample reasoning graph includes the plurality of candidate nodes and at least one candidate edge; At least one candidate weight for at least one candidate edge is determined to obtain the sample weight graph.
5. The method according to claim 4, wherein Two different candidate nodes in the sample inference graph are connected via at least one candidate edge, and the candidate edge is a directed edge, and the candidate edge points from a candidate node as an upstream node in the two different candidate nodes to a candidate node as a downstream node in the two different candidate nodes. The logical relationship includes causal consistency, and the candidate weights include causal consistency weights.
6. The method according to claim 1, wherein Obtaining at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path according to the sample weight map includes: Determining at least one of the sample reasoning paths according to at least one of the candidate weights; At least one of the sample reasoning paths is adjusted using a large model to obtain at least one first optimized reasoning path corresponding to the at least one sample reasoning path.
7. The method according to claim 6, wherein: Determining at least one of the sample inference paths according to at least one of the candidate weights includes: Determining a first-level sample node from a plurality of candidate nodes; Determining, based on a k-th level sample node and at least one k-th level candidate weight for at least one k-th level candidate edge, a k+1-th level sample node from at least one candidate downstream node of the k-th level sample node, wherein the k-th level candidate edge points from the k-th level sample node to a candidate downstream node of the k-th level sample node in the sample weight graph, and k is an integer greater than or equal to 1; The k-th level candidate edge pointing from the k-th level sample node to the k+1-th level sample node is determined as the k-th level sample edge.
8. The method according to claim 7, wherein: Determining at least one of the sample reasoning paths according to at least one of the candidate weights further includes: In response to determining that the k+1th sample node does not have a candidate downstream node, determining the sample reasoning path based on the k+1 sample nodes and the k sample edges, the sample reasoning path including the k+1 sample nodes and the k sample edges, the k+1 sample node including the k-th level sample node and the k+1-th level sample node, and the k sample edges including the k-th level sample edge.
9. The method according to claim 1, further comprising: A first target reasoning path is determined according to the sample weight graph.
10. A model fine-tuning method, comprising: Inputting at least one sample reasoning path into the model to be fine-tuned to obtain at least one second optimized reasoning path; fine-tuning the to-be-fine-tuned model according to at least one first optimized inference path and at least one second optimized inference path, Among them, at least one of the first optimized reasoning paths corresponds one-to-one to at least one of the sample reasoning paths, and at least one of the sample reasoning paths and the first optimized reasoning path corresponding to the sample reasoning path are obtained according to the method according to any one of claims 1 to 9.
11. The method according to claim 10, wherein: Fine-tuning the to-be-fine-tuned model according to at least one first optimized reasoning path and at least one second optimized reasoning path comprises: Inputting the sample weight map into the model to be fine-tuned to obtain a second target reasoning path; Fine-tune the model to be fine-tuned according to a first target reasoning path, the second target reasoning path, at least one first optimized reasoning path, and at least one second optimized reasoning path, wherein the first target reasoning path is determined according to the sample weight map.
12. The method according to claim 11, wherein Fine-tuning the to-be-fine-tuned model according to the first target reasoning path, the second target reasoning path, at least one of the first optimized reasoning paths, and at least one of the second optimized reasoning paths includes: determining at least one structural loss based on at least one of the first optimized reasoning paths and at least one of the second optimized reasoning paths; determining a causal loss according to the first target reasoning path and the second target reasoning path; The to-be-fine-tuned model is fine-tuned according to the causal loss and at least one of the structural losses.
13. A method for processing question-answer information, comprising: Inputting problem information into a target model to obtain a target thinking chain for the problem information; Generate answer information for the question information based on the target thought chain, The target model is fine-tuned according to the method according to any one of claims 10 to 12.
14. A sample generation device, comprising: A first obtaining module is configured to obtain a plurality of sample thinking chains according to an initial thinking chain, wherein the sample thinking chains include a plurality of operations to be executed; a second obtaining module, configured to obtain a sample weight graph based on the plurality of sample thought chains, wherein the sample weight graph includes a plurality of candidate nodes and at least one candidate edge, the candidate nodes representing the operations to be performed, and the candidate weights for the candidate edges indicating a logical relationship between two candidate nodes connected via the candidate edges; A third acquisition module is used to obtain at least one sample reasoning path and a first optimized reasoning path corresponding to the sample reasoning path based on the sample weight graph, wherein the sample reasoning path includes multiple sample nodes determined from the multiple candidate nodes and at least one sample edge for connecting the multiple sample nodes.
15. A model fine-tuning device comprising: a fourth obtaining module, configured to input at least one sample reasoning path into the model to be fine-tuned to obtain at least one second optimized reasoning path; a fine-tuning module, configured to fine-tune the model to be fine-tuned according to at least one first optimized reasoning path and at least one second optimized reasoning path, At least one of the first optimized reasoning paths corresponds one-to-one to at least one of the sample reasoning paths, and the at least one sample reasoning path and the first optimized reasoning path corresponding to the sample reasoning path are obtained according to the apparatus according to claim 14.
16. A question-answer information processing device, comprising: a fifth obtaining module, configured to input the problem information into a target model and obtain a target thinking chain for the problem information; A generation module is used to generate answer information for the question information according to the target thinking chain, The target model is fine-tuned according to the apparatus according to claim 15.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13.
19. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.