Data processing method and device, computer equipment, medium and program product
By introducing diverse inference and occlusion path mechanisms in language model training, the problems of low inference accuracy and poor generalization ability of language models in the prior art are solved, and higher inference accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510099425.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the language model is trained based on chain reasoning prompts, and the inference results obtained are low in accuracy, poor generalization ability, and there is a problem of large model hallucination.
By obtaining the problem sample, using the first language model for diversified inference, multiple candidate paths are obtained, and then some of the inference steps of the candidate path are covered to obtain the occluded path. Then, based on the indication of the occluded path, the problem sample is reasoned through the second language model to obtain the path to be verified. Determining the same degree between the candidate path and the path to be verified, the result path is determined from multiple candidate paths for training the inference ability of the language model.
It improves the accuracy and generalization ability of the inference results of the language model, reduces the probability of large-scale model hallucination, and enhances the robustness and exploration space of the language model.
Smart Images

Figure CN120012933A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, computer equipment, medium and program product. Background Art
[0002] A language model is an artificial intelligence model that uses machine learning technology to understand and generate human language, and can perform tasks such as text analysis, sentiment analysis, language translation, and speech recognition.
[0003] In the related art, examples including multiple reasoning steps are generally presented to the language model through chain-of-thought (CoT) prompts, guiding the language model to show a similar reasoning process when answering questions.
[0004] However, the reasoning results obtained by training the language model based on chain reasoning prompts have low accuracy and poor generalization ability. Summary of the invention
[0005] In order to solve the above technical problems, the present application provides a data processing method, apparatus, computer equipment, medium and program product for improving the reasoning and generalization capabilities of language models.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] On the one hand, an embodiment of the present application provides a data processing method, the method comprising:
[0008] Get sample questions;
[0009] Reasoning the problem sample according to the first language model to obtain multiple candidate paths, where different candidate paths correspond to different reasoning methods for solving the problem sample, and the candidate paths include multiple reasoning steps with a reasoning order;
[0010] Covering part of the reasoning steps in each of the candidate paths to obtain covered paths corresponding to each of the candidate paths;
[0011] Based on the reasoning order indicated by each of the covered paths, the question sample is reasoned through the second language model to obtain paths to be verified corresponding to each of the covered paths, wherein the paths to be verified are paths obtained by reasoning the covered reasoning steps in the covered paths, and the first language model and the second language model are different language models;
[0012] Determine a result path from the multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path;
[0013] A training sample is determined according to the question sample, the multiple candidate paths and the result path, and the training sample is used to train the reasoning ability of the language model.
[0014] On the other hand, an embodiment of the present application provides a data processing device, the device comprising: an acquisition unit, an inference unit, a covering unit, a completion unit, a selection unit, and a determination unit;
[0015] The acquisition unit is used to acquire question samples;
[0016] The reasoning unit is used to reason the problem sample according to the first language model to obtain multiple candidate paths, where different candidate paths correspond to different reasoning methods for solving the problem sample, and the candidate paths include multiple reasoning steps with a reasoning order;
[0017] The covering unit is used to cover part of the reasoning steps in each of the candidate paths to obtain the covered paths corresponding to each of the candidate paths;
[0018] The completion unit is configured to perform reasoning on the question sample through a second language model based on the reasoning order indicated by each of the covered paths, to obtain paths to be verified corresponding to each of the covered paths, wherein the paths to be verified are paths obtained by reasoning on the covered reasoning steps in the covered paths, and the first language model and the second language model are different language models;
[0019] The selection unit is used to determine a result path from the multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path;
[0020] The determination unit is used to determine a training sample according to the question sample, the multiple candidate paths and the result path, and the training sample is used to train the reasoning ability of the language model.
[0021] On the other hand, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory:
[0022] The memory is used to store a computer program and transmit the computer program to the processor;
[0023] The processor is configured to execute the method described in the above aspects according to the instructions in the computer program.
[0024] On the other hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method described in the above aspects.
[0025] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the method described in the above aspects.
[0026] It can be seen from the above technical solution that the first language model is used to perform diversified reasoning on the problem sample to obtain multiple candidate paths, and some reasoning steps of the candidate path are covered to obtain the covered path. Based on the indication of the covered path, the second language model is used to perform diversified reasoning on the problem sample, so as to re-reason the covered reasoning steps in the covered path to obtain the path to be verified. According to the consistency between the candidate path and the corresponding path to be verified, the result path is obtained from the multiple candidate paths, and the accuracy of the reasoning result corresponding to the result path is high, so that the training sample is obtained according to the problem sample, the multiple candidate paths and the result path. Therefore, based on the multiple candidate paths in the training sample, the language model can learn a larger exploration space, and improve the generalization ability and robustness of the language model. Moreover, by introducing the first language model and the second language model with similar capabilities, the two can verify each other, and the capabilities complement each other. The second language model helps the first language model recognize the errors that the first language model cannot recognize, reduces the probability of the large model hallucination, and improves the accuracy of the result path, so that the trained language model can learn the result path with higher accuracy from multiple reasoning processes, and improves the accuracy of the subsequent reasoning results obtained based on the language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0028] Figure 1 A schematic diagram of an application scenario of a data processing method provided in an embodiment of the present application;
[0029] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;
[0030] Figure 3 A schematic diagram of a search tree provided in an embodiment of the present application;
[0031] Figure 4 A schematic diagram of adding a node to an initial search tree provided in an embodiment of the present application;
[0032] Figure 5 A schematic diagram of applying a plug-in to a business scenario provided in an embodiment of the present application;
[0033] Figure 6 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0034] Figure 7 A schematic diagram of the structure of a server provided in an embodiment of the present application;
[0035] Figure 8 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The embodiments of the present application are described below in conjunction with the accompanying drawings.
[0037] The process of training the language model based on the chain reasoning prompt to obtain the reasoning result can be broken down into: creating training data based on the problem sample, the training data includes examples of multiple reasoning steps, and training the language model again based on the training sample, so that the language model learns the reasoning ability based on the training data. For example, the problem sample is Zhang San has 5 tennis balls, he bought two more cans of tennis balls, each can has 3 tennis balls, how many tennis balls does Zhang San have now, the answer is 11. The answer to the training data created based on the problem sample is that Zhang San first had 5 tennis balls, 3 tennis balls in each can, 2 cans totaling 6 tennis balls, 5+6=11, Zhang San now has 11 tennis balls.
[0038] Therefore, by converting problem samples into training data with chained reasoning prompts, complex problems can be broken down into a series of coherent reasoning steps, guiding the language model to gradually delve into the core of the problem, thereby improving the efficiency and accuracy of solving complex reasoning tasks.
[0039] After analysis, it was found that in the process of reasoning about problem samples, the reasoning process is like a chain, that is, there is only one reasoning process, and the exploration space corresponding to this reasoning process is small.
[0040] However, there are different choices in each reasoning step, that is, each reasoning step may get a different reasoning process due to choosing a different reasoning direction, and different reasoning results will be obtained based on different reasoning processes, that is, the larger the exploration space corresponding to the reasoning process, the more accurate the reasoning results can be obtained. Taking the following Go task as an example, each reasoning step can choose a different position to play Go, and then different chess games will be obtained. Different chess games may get different reasoning results, such as losing or winning.
[0041] That is to say, if the training data is obtained only in the form of chain reasoning prompts, the language model trained with the training data can only perform reasoning based on the learned smaller exploration space and lacks comprehensive exploration of a wide range of possibilities. As a result, the inference results obtained by training the language model based on chain reasoning prompts are less accurate and have poor generalization ability. Moreover, due to the AI Hallucinations problem in the language model, that is, a language model cannot be aware of its own errors, resulting in lower accuracy in the inference results.
[0042] Based on this, the embodiment of the present application provides a data processing method, which expands the problem sample into multiple candidate paths, so that the language model can learn a larger exploration space based on the multiple candidate paths, thereby improving the generalization ability and robustness of the language model. Moreover, by introducing a first language model and a second language model with similar capabilities, the two can verify each other and complement each other, thereby reducing the probability of large model hallucinations and thus the accuracy of the inference results.
[0043] The data processing method provided in this application can be applied to computer devices with data processing capabilities, such as terminal devices and servers.
[0044] Among them, the terminal devices can specifically be desktop computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. The smart car-mounted devices can be car-mounted navigation terminals and car-mounted computers, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc., but are not limited to these.
[0045] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0046] In order to facilitate understanding of the data processing method provided in the embodiment of the present application, the application scenario of the data processing method is exemplarily introduced below by taking the execution subject of the data processing method as a server as an example.
[0047] See also Figure 1 , which is a schematic diagram of an application scenario of a data processing method provided in an embodiment of the present application. Figure 1As shown, the application scenario includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 can communicate through a communication network. The communication network uses standard communication technology and / or protocols, usually the Internet, but can also be any network, including but not limited to Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, dedicated network or any combination of virtual private networks. In some embodiments, customized or dedicated data communication technology can be used to replace or supplement the above data communication technology.
[0048] The terminal device 110 is installed with a client, which is used to provide a service for generating training samples. The terminal device 110 can expand the existing training data set, such as expanding each piece of data in the training data set as a problem sample to obtain a training sample with multiple reasoning steps. Specifically, after obtaining the problem sample, the terminal device 110 sends the problem sample to the server 120.
[0049] The server 120 is a server corresponding to the client, and is used to provide the client with a service of generating training samples. After receiving the question sample, the server 120 calls the first language model to perform diversified reasoning on the question sample to obtain multiple candidate paths. The multiple candidate paths can form a search tree. Figure 1 1 shows a search tree 1 taking four candidate paths as an example. For example, a candidate path may be node A-node B2-node C1. Each node in the search data corresponds to an inference step, and the relationship between the nodes is determined based on the inference order between the inference steps.
[0050] The server 120 masks some of the reasoning steps in the candidate path to obtain multiple masked paths, such as masking the last reasoning step of each candidate path. Figure 1 , node B1, node C1, node C2 and node B3 can be covered, thereby obtaining 4 covering paths, such as node A-node B2-covered node.
[0051] The server 120 calls a second language model with similar functions to the first language model, and performs diversified reasoning according to the reasoning order indicated by the covered path through the second language model, thereby inferring the covered nodes and then obtaining the path to be verified. Figure 1The second language model can first learn the reasoning process of node A and node B2, and then infer that the next covered node is C1, thus a path to be verified consisting of node A-node B2-node C1, and then 4 paths to be verified are obtained. The 4 paths to be verified can also get the search tree 2.
[0052] The server 120 determines a result path from the plurality of candidate paths according to the degree of similarity between the candidate paths and their corresponding to-be-verified paths, such as the one with the highest accuracy of the inference result obtained based on the result path. Figure 1 , the similarity of the four candidate paths is 50%, 100%, 66.7% and 50% respectively, and the resulting path is node A-node B2-node C1.
[0053] The server 120 obtains a training sample based on the question sample, multiple candidate paths and the result path, so as to train the reasoning ability of the language model based on the training sample. Figure 1 , a training sample can be obtained based on the problem sample, search tree 1 and the result path node A-node B2-node C1.
[0054] Therefore, based on the multiple candidate paths in the training sample, the language model can learn a larger exploration space, improving the generalization ability and robustness of the language model. Moreover, by introducing the first language model and the second language model with similar capabilities, the two can verify each other, and the second language model can help the first language model recognize errors that the first language model cannot recognize, reducing the probability of large model hallucinations and improving the accuracy of the result path, so that the trained language model can learn a more accurate result path from multiple reasoning processes, improving the accuracy of the subsequent reasoning results based on the language model.
[0055] The data processing method provided in the embodiment of the present application can be executed by the server. However, in other embodiments of the present application, the terminal device can also have similar functions as the server, thereby executing the data processing method provided in the embodiment of the present application, or the terminal device and the server can jointly execute the data processing method provided in the embodiment of the present application, which is not limited in the present embodiment.
[0056] It should be noted that the above application scenarios are only examples, and the data processing method provided in this embodiment can also be applied to other scenarios, which are not limited here.
[0057] A data processing method provided by the present application is described in detail below through a method embodiment.
[0058] See also Figure 2, which is a flow chart of a data processing method provided by an embodiment of the present application. For the sake of convenience, the following embodiment is still introduced by taking the execution subject of the data processing method as a server as an example. Figure 2 As shown, the data processing method includes S201-S206.
[0059] S201: Obtain question samples.
[0060] Question samples are samples used to train language models. The question samples include questions to be answered and answers. For example, the question to be answered is: There are now 1 virtual coin of bread, 2 virtual coins of bread, and 3 virtual coins of bread. Xiao Ming has 10 virtual coins. How many loaves of bread can Xiao Ming buy at most? Answer: 10 loaves of bread.
[0061] As a possible implementation method, an existing training data set can be obtained, and each data in the training data set can be used as a problem sample, so that each data can be expanded into a training sample, thereby expanding the training data set into a training sample set with reasoning steps.
[0062] All data collected in this application (such as question samples) are collected with the consent and authorization of the object to which the data belongs (such as users, institutions or enterprises), and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0063] S202: Reasoning the question sample according to the first language model to obtain multiple candidate paths.
[0064] A language model is an artificial intelligence model that uses machine learning technology to understand and generate human language. It has reasoning capabilities and can perform tasks such as text analysis, sentiment analysis, language translation, and speech recognition. The language model is trained based on a large amount of knowledge, which gives the language model a large amount of knowledge reserve. However, the language model is not very good at applying the learned knowledge, that is, it cannot fully stimulate the knowledge it has learned. In other words, since the exploration space learned by the language model is small, although it has reasoning capabilities, its reasoning capabilities are poor. For example, the language model has learned 80 points of knowledge, but the language model cannot integrate it, and it can only reflect 60 points of reasoning capabilities.
[0065] Based on this, in order to stimulate the language model to improve its application of learned knowledge, that is, to enhance the reasoning ability of the language model, the embodiment of the present application expands the problem sample into a training sample, so as to enable the language model to learn a larger exploration space through multiple candidate paths in the training sample. The process of obtaining multiple candidate paths is described below.
[0066] The first language model is a language model, and the first language model has reasoning ability, so the problem sample can be reasoned through the first language model to obtain multiple candidate paths. The embodiment of the present application does not specifically limit the first language model, such as a large language model (LLM), a small language model (SLM), etc.
[0067] Among them, reasoning refers to a form of thinking that derives a new judgment (conclusion) from one or several known judgments (knowledge). A candidate path is a reasoning method used to solve a problem sample, and different candidate paths correspond to different reasoning methods for solving a problem sample. The candidate path includes multiple reasoning steps, and there is a reasoning order between the multiple reasoning steps. The reasoning order is the order used for reasoning, that is, using the reasoning steps in the reasoning order can solve the problem sample.
[0068] The embodiments of the present application do not specifically limit the form of multiple candidate paths. As a possible implementation method, multiple candidate paths can be embodied in the form of a tree, that is, a search tree. Specifically, the problem sample is converted into a multi-step reasoning problem, each reasoning step includes a reasoning action and an answer obtained based on the reasoning action, and one reasoning step corresponds to a node in the search tree. The root node represents the problem sample, and the connection between nodes (i.e., edges) represents the reasoning action, so that the path from the root node to the leaf node corresponds to a candidate path. See formula (1):
[0069]
[0070] Among them, path t corresponds to a candidate path; x represents the root node; s i represents the i-th intermediate reasoning step (abbreviated as intermediate reasoning step), such as s1 is the first intermediate reasoning step; s d Represents a leaf node.
[0071] The embodiment of the present application does not specifically limit the method of obtaining multiple candidate paths through the first language model, which will be explained later based on B1-B4 or C1-C4 and will not be repeated here.
[0072] S203: Cover part of the reasoning steps in each candidate path to obtain the covered paths corresponding to each candidate path.
[0073] From the above, we can see that since language models generally have the problem of large model hallucination, the first language model may not be able to detect the problems it has in the reasoning process. Even if the exploration space is increased, it may still lead to lower accuracy of the reasoning results.
[0074] Based on this, the embodiment of the present application introduces the second language model as a control group of the first language model, which is equivalent to using the first language model as a generator and the second language model as a discriminator, and improving the accuracy of solving problem samples through the generator-discriminator model.
[0075] In order for the second language model to verify the first language model, first mask some of the reasoning steps in each candidate path to obtain the masked paths corresponding to each candidate path. Take the target candidate path among multiple candidate paths as an example. The target candidate path includes multiple reasoning steps. Some of the reasoning steps in the multiple reasoning steps can be masked to obtain the masked path corresponding to the target candidate path. Take each candidate path as the target candidate path to obtain the masked path corresponding to each candidate path.
[0076] The embodiments of the present application do not specifically limit the covering method, and three methods are used as examples for explanation below.
[0077] Method 1: Randomly cover any reasoning steps in the candidate path.
[0078] Taking the target candidate path as an example, the target candidate path can be expressed as reasoning step A-reasoning step B-reasoning step C. Reasoning step B can be masked, and the resulting masked path is reasoning step A-unknown reasoning step-reasoning step C, where the unknown reasoning step is the reasoning step waiting to be obtained through reasoning by the second language model.
[0079] Method 2: Cover the intermediate reasoning steps at fixed positions in the candidate path.
[0080] Among them, the fixed position can be set by numbers, proportions, etc. For example, the middle m reasoning steps in each candidate path are covered, where m is an integer greater than 0, to obtain the covered paths corresponding to each candidate path. For example, the third reasoning step in each candidate path can be covered. For another example, one-third of the reasoning steps in the middle of each candidate path can be covered. For example, if candidate path A includes 6 reasoning steps, the 3rd and 4th reasoning steps of candidate path A can be covered. If candidate path B includes 9 reasoning steps, the 4th to 6th reasoning steps of candidate path B can be covered, and so on.
[0081] Therefore, compared with method 1, the covering method of method 2 is more controllable, that is, the number and position of the reasoning steps covered by each candidate path are more similar, so that when comparing the degree of similarity between the covered path and the candidate path later, they are more comparable.
[0082] Method three: cover the last n reasoning steps in the candidate path.
[0083] The last n reasoning steps in each candidate path are covered, where n is an integer greater than 0, to obtain the covered paths corresponding to each candidate path. Taking n as 2 as an example, if candidate path A includes 6 reasoning steps, the 5th and 6th reasoning steps can be covered. If candidate path B includes 9 reasoning steps, the 8th and 9th reasoning steps of candidate path B can be covered, and so on.
[0084] Therefore, method three not only has the advantages of method two, but also covers the last n reasoning steps compared to method two in which the intermediate reasoning steps are covered, which is beneficial for the subsequent second language model to learn other reasoning steps in the candidate path (i.e., the remaining reasoning steps other than the n reasoning steps), and can supplement the reasoning steps after learning more reasoning capabilities, so that not only the subsequent supplemented reasoning steps are more accurate, but also more robust.
[0085] S204: Based on the reasoning order indicated by each covered path, the second language model is used to reason on the question sample to obtain the paths to be verified corresponding to each covered path.
[0086] The second language model is a language model, and the second language model has reasoning ability, so the second language model can be used to reason about the problem sample. Compared with the candidate path, although some reasoning steps in the covered path are covered, there is still a reasoning order between the uncovered reasoning steps and the covered reasoning steps. Therefore, in the process of reasoning about the problem sample through the second language model, reasoning is performed based on the reasoning order indicated by each covered path, and the covered reasoning steps in the covered path are completed, so as to obtain the path to be verified corresponding to the covered path, that is, the path to be verified is the path obtained by reasoning the covered reasoning steps in the covered path.
[0087] For example, if candidate path A includes 6 reasoning steps, the 5th and 6th reasoning steps can be masked to obtain masked path A. The second language model can first learn according to the first 4 reasoning steps to learn the reasoning method of the path, and then continue to reason on the 5th and 6th reasoning steps based on the learned reasoning method, so as to obtain a path to be verified including 6 reasoning steps. As a possible implementation method, after learning the reasoning method corresponding to the first 4 reasoning steps, the second language model continues to reason based on the reasoning method, and can infer 1 or 3 or more reasoning steps, that is, the number of reasoning steps included in the obtained path to be verified can be different from the number of reasoning steps included in the candidate path corresponding to the masked path, etc. This application does not make specific restrictions on this.
[0088] Although both the first language model and the second language model are language models, they are different language models, or in other words, they are language models trained based on incompletely overlapping training data. Based on this, the reasons that cause the first language model to have hallucination problems are generally different from the reasons that cause the second language model to have hallucination problems, so the first language model and the second language model can complement each other, such as the second language model can help the first language model realize its hallucination problems, and the first language model can help the second language model realize its hallucination problems.
[0089] S205: Determine a result path from multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path.
[0090] The path to be verified corresponding to the masked path is not only the result obtained by learning the masked path, but also the result obtained by reasoning based on the knowledge learned by the second language model. Therefore, by comparing the degree of similarity between the candidate path corresponding to the masked path and the path to be verified corresponding to the masked path, the difference between the first language model and the second language model can be obtained. The difference may be caused by the illusion problem of the two. Therefore, the higher the degree of similarity between the candidate path and the path to be verified, the greater the possibility that the accuracy is higher. Therefore, a result path with higher accuracy can be obtained from multiple candidate paths, that is, the result path is one of the multiple candidate paths, and the accuracy of the candidate path is higher.
[0091] The degree of similarity refers to the degree of similarity between the candidate path and the corresponding path to be verified, which can be evaluated from multiple dimensions such as the number of reasoning steps, the order of reasoning steps, whether the reasoning steps are the same, etc. This application does not make specific restrictions on this. For example, the candidate path is reasoning step A-reasoning step B-reasoning step C, and the corresponding path to be verified is reasoning step A-reasoning step B-reasoning step D, then the degree of similarity between the candidate path and its corresponding path to be verified is 66.67%.
[0092] The embodiments of the present application do not specifically limit the method of determining the result path. For example, the result path is determined from multiple candidate paths only based on the same degree, so that the result path is a candidate path with higher accuracy among the multiple candidate paths. For another example, the accuracy of each candidate path is determined, and the result path is determined based on the correctness and the same degree, that is, the result path is determined from multiple candidate paths based on the same degree between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path, and the correctness of the candidate path corresponding to the covered path. As a possible implementation method, please refer to formula (2):
[0093] t * = argmax t(Q(t)×C(t)) (2)
[0094] Among them, t * is the result path; Q(t) is the accuracy of the t-th candidate path; C(t) is the degree of similarity between the t-th candidate path and its corresponding covered path.
[0095] The accuracy of the candidate path is used to identify the accuracy of the answer obtained by solving the problem sample through the candidate path, for example, whether the category recognition based on the candidate path is accurate. For another example, taking the problem sample as a chess task, if the chess result obtained based on the candidate path is a win, the accuracy is 1, otherwise it is a loss, the accuracy is 0, which can be expressed as formula (3):
[0096]
[0097] Among them, Q(s d ,a d ) is a leaf node s d When performing reasoning action a d The accuracy rate obtained later.
[0098] Therefore, by selecting candidate paths with the same high degree and high accuracy as the result path, not only can the accuracy of the result path itself be guaranteed to be high, but also the candidate paths with low accuracy caused by the hallucination problem of the first language model can be eliminated from the candidate paths, thereby further improving the accuracy of the result path and then improving the effectiveness of subsequent training samples.
[0099] S206: Determine a training sample according to the problem sample, multiple candidate paths and the result path.
[0100] The training samples include problem samples, multiple candidate paths and result paths, wherein the problem samples in the training samples can be used as problems waiting to be solved, the multiple candidate paths are used to guide the language model to learn a larger exploration space, and the result path is used to guide the language model to correctly find a path with higher accuracy from multiple candidate paths, so that the reasoning ability of the language model can be trained based on the training samples and the reasoning ability of the language model can be improved. For example, the first language model can be trained again based on the obtained training samples to improve the reasoning ability of the first language model.
[0101] The embodiments of the present application do not specifically limit the method for determining the training samples, and three methods are used as examples for illustration below.
[0102] Method 1: directly use multiple candidate paths to obtain training samples.
[0103] Directly use all candidate paths, that is, all candidate paths, problem samples and result samples, to obtain training samples.
[0104] Method 2: Screen multiple candidate paths to obtain training samples.
[0105] A target candidate path is obtained by screening from multiple candidate paths, and the target candidate path is a candidate path whose degree of identity is greater than the same threshold among multiple candidate paths, that is, the target candidate path has a high degree of identity with its corresponding path to be verified, so that the accuracy of the target candidate path is high. Thus, the target candidate path is used as the target path, and a training sample is obtained based on the target path, the problem sample and the result sample. The embodiment of the present application does not specifically limit the size of the same threshold, and those skilled in the art can set it according to actual needs.
[0106] Therefore, compared with method 1, although the number of target paths constituting the training samples in method 2 is smaller, the accuracy of the target paths is higher, so that the language model subsequently trained based on the training samples can expand the exploration space while improving the accuracy of exploration.
[0107] Method three, obtaining training samples based on multiple candidate paths and multiple paths to be verified.
[0108] A target candidate path is obtained by screening from multiple candidate paths. The target candidate path is a candidate path with a degree of similarity greater than the same threshold among multiple candidate paths, that is, the target candidate path has a high degree of similarity with its corresponding path to be verified, so that the accuracy of the target candidate path is high. A target path to be verified is obtained by screening from multiple paths to be verified. The target path to be verified is a path to be verified with a degree of similarity greater than the same threshold among multiple paths to be verified, that is, the target path to be verified has a high degree of similarity with its corresponding candidate path, so that the accuracy of the target path to be verified is high. Thus, the target candidate path and the target path to be verified are used as the target path, and a training sample is obtained based on the target path, the problem sample and the result sample.
[0109] Therefore, method three not only has the advantages of method two, but also compared with method two, method three has a larger number of target paths that constitute the training samples, and the accuracy of the target paths is higher. Therefore, the language model subsequently trained based on the training samples can learn a larger exploration space, thereby achieving the goal of further expanding the exploration space while improving the accuracy of exploration.
[0110] As a possible implementation method, the language model can be trained again based on the training samples, thereby expanding the exploration space learned by the language model and improving the reasoning ability of the language model. The training of the language model using the training samples obtained in the above method 2 or method 3 is used as an example for explanation.
[0111] According to the reasoning order indicated by the target path in the training sample, the problem sample in the training sample is inferred through the language model, thereby obtaining multiple reasoning paths and the predicted optimal path among the multiple reasoning paths. Among them, the reasoning path is obtained by learning the multiple reasoning steps included in the target path by the language model. The reasoning path can be consistent with the target path or slightly different from the target path, and this application does not make specific restrictions on this.
[0112] The predicted optimal path is the inference path with higher accuracy among multiple inference paths, and the result path is the path with higher accuracy (such as the highest) among multiple target paths. Based on the difference between the predicted optimal path and the result path, the model parameters of the language model are adjusted so that the predicted optimal path becomes more and more like the result path, thereby obtaining an optimized language model.
[0113] As a result, the accuracy of the training samples is high, and the exploration space indicated by the target path is large. Therefore, training the language model again based on the training samples can stimulate the reasoning ability of the language model, making it more proficient in using existing knowledge, thereby improving the accuracy of the reasoning results. Taking the language model as the first language model as an example, the candidate path that is not used as the target path may be due to the low accuracy caused by the hallucination problem of the first language model. The first language model can be trained based on the training samples in the future, so that the first language model can be aware of the errors in its reasoning process and improve the accuracy of the first language model.
[0114] In addition, compared with large language models, small language models have fewer model parameters and weaker reasoning capabilities. By retraining the small language model with training samples, the reasoning capability of the small language model can be improved to a greater extent. That is, it is more suitable for small language models and can make up for the problem of insufficient self-evaluation ability of small language models.
[0115] It can be seen from the above technical solution that the first language model is used to perform diversified reasoning on the problem sample to obtain multiple candidate paths, and some reasoning steps of the candidate path are covered to obtain the covered path. Based on the indication of the covered path, the second language model is used to perform diversified reasoning on the problem sample, so as to re-reason the covered reasoning steps in the covered path to obtain the path to be verified. According to the consistency between the candidate path and the corresponding path to be verified, the result path is obtained from the multiple candidate paths, and the accuracy of the reasoning result corresponding to the result path is high, so that the training sample is obtained according to the problem sample, the multiple candidate paths and the result path. Therefore, based on the multiple candidate paths in the training sample, the language model can learn a larger exploration space, and improve the generalization ability and robustness of the language model. Moreover, by introducing the first language model and the second language model with similar capabilities, the two can verify each other, and the capabilities complement each other. The second language model helps the first language model recognize the errors that the first language model cannot recognize, reduces the probability of the large model hallucination, and improves the accuracy of the result path, so that the trained language model can learn the result path with higher accuracy from multiple reasoning processes, and improves the accuracy of the subsequent reasoning results obtained based on the language model. In addition, this method of automatically expanding problem samples can reduce the quality requirements for training samples and reduce resource overhead.
[0116] As can be seen from the above, in the related art, the language model is generally retrained based on a reasoning method, resulting in a smaller exploration space learned by the language model. Based on this, in order to expand the exploration space, the present application performs diversified reasoning on the question sample through the first language model to obtain multiple candidate paths.
[0117] The following two methods are used as examples for description. For method 1, please refer to B1-B4 for details, and for method 2, please refer to C1-C4 for details.
[0118] B1: Reason the question sample based on the first language model to obtain the candidate path to be determined.
[0119] The first language model has reasoning ability, so the problem sample can be reasoned through the first language model to obtain one or more candidate paths. A candidate path is a reasoning method for solving the problem sample, and different candidate paths correspond to different reasoning methods for solving the problem sample. The candidate path includes multiple reasoning steps, and there is a reasoning order between the multiple reasoning steps.
[0120] B2: Based on the reasoning steps, determine the synonymous steps.
[0121] The similarity between the synonymous step and the reasoning step is greater than the similarity threshold, that is, the similarity between the synonymous step and the reasoning step is high. The embodiment of the present application does not specifically limit the size of the similarity threshold, and those skilled in the art can set it according to actual needs.
[0122] The embodiments of the present application do not specifically limit the method of obtaining synonymous steps based on reasoning steps. For example, if the reasoning steps are described by text, the vocabulary in the reasoning steps can be expanded to synonyms, or the reasoning steps can be converted into synonymous sentences, etc., to obtain synonymous steps with a high degree of similarity to the reasoning steps.
[0123] The embodiment of the present application does not specifically limit which reasoning steps are selected from multiple reasoning steps to generate their corresponding synonymous steps. All or part of the reasoning steps may be generated.
[0124] B3: Obtain the search tree based on the pending candidate paths and synonymous steps.
[0125] As can be seen from the above, the search tree includes multiple nodes, each node corresponds to an inference step, and the multiple inference steps and synonymous steps of the inference steps included in the candidate path to be determined are all nodes in the search tree. Among them, the root node of the search tree is a node obtained based on the problem sample, that is, the node corresponds to the problem sample, and there is an inference order between the multiple inference steps. The relationship between the nodes is determined based on the inference order. For example, the inference order of the inference step corresponding to the parent node in the search tree is before the inference order of the inference step corresponding to the child node of the parent node, that is, the inference step corresponding to the parent node is an inference step before the inference step corresponding to its child node, and the synonymous step is a sibling node of the node corresponding to its corresponding inference step.
[0126] See also Figure 3 , which is a schematic diagram of a search tree provided in an embodiment of the present application. Figure 3 In the example, the question sample is inferred according to the first language model to obtain a candidate path to be determined, which can be represented as node A-node B1-node C1, and each node corresponds to an inference step. For the inference step corresponding to node B1, a synonymous step is generated to obtain nodes B2 and B3 corresponding to the synonymous step. For the inference step corresponding to node C1, a synonymous step is generated to obtain node C2 corresponding to the synonymous step, thereby obtaining a search tree.
[0127] As a possible implementation method, in order to enrich the search tree and expand the exploration space, it is possible to determine whether the path corresponding to each leaf node in the search tree has completed the solution to the problem sample. For example, the first language model and other language models can be used to continue to expand the leaf nodes that have not been solved to obtain the child nodes of the leaf nodes. For another example, the child nodes of the brother nodes are used as the child nodes of the leaf nodes that have not been solved, etc. This application does not make specific restrictions on this.
[0128] B4: Determine multiple candidate paths based on the search tree.
[0129] The candidate path is the path from the root node to the leaf node. For example, starting from the root node of the search tree, searching downward until a leaf node is found, thereby obtaining the path from the root node to the leaf node, i.e., the candidate path. In other words, among the multiple reasoning steps included in the candidate path, the first reasoning step is the root node, and the last reasoning step is the leaf node.
[0130] Therefore, by performing synonymous conversion on the reasoning steps, the synonymous steps corresponding to the reasoning steps are obtained, and then the exploration space, such as the search tree, is constructed together with the synonymous steps and the reasoning steps, which expands the size of the exploration space so as to obtain more candidate paths. Moreover, the exploration space is displayed in the form of a search tree, which is concise and clear, and is convenient and quick to obtain candidate paths later.
[0131] After introducing the first method for determining the candidate path, the second method for determining the candidate path is described below.
[0132] C1: Select the best existing path from the initial search tree through the first language model.
[0133] The initial search tree is a search tree that has not yet been constructed. The initial search tree includes one or more nodes, each node corresponds to an inference step, and the first node included in the initial search tree is the root node. The root node is a node obtained based on the problem sample, that is, the node corresponds to the problem sample.
[0134] As the reasoning continues, each time a new reasoning step is added, a node is added to the initial search tree. There is a reasoning order between multiple reasoning steps, and the relationship between nodes is determined based on the reasoning order. For example, the reasoning order of the reasoning step corresponding to the parent node in the initial search tree is before the reasoning order of the reasoning step corresponding to the child node of the parent node, that is, the reasoning step corresponding to the parent node is an reasoning step before the reasoning step corresponding to its child node.
[0135] The following is an explanation of the process of adding nodes to the initial search tree.
[0136] First, initialize the search tree, that is, by inputting the problem sample, determine the root node of the initial search tree, which can be expressed as s0. The obtained initial search tree can be expressed as formula (4):
[0137] T={s0}, s0=x (4)
[0138] Among them, T is the initial search tree, s0 is the root node, and x represents the root node.
[0139] Then, nodes are added to the initial search tree, that is, child nodes are added downward from the root node until a search tree is obtained. The following is an example in which the initial search tree includes multiple nodes and nodes are continuously added to the initial search tree.
[0140] The initial search tree includes multiple leaf nodes, and a leaf node is a node without child nodes. Multiple paths are obtained from the root node to each leaf node, and the existing optimal path is selected from the multiple paths. The existing optimal path is the path with the highest accuracy rate obtained for solving the problem sample among the multiple paths included in the initial search tree. The embodiment of the present application does not specifically limit the method of obtaining the existing optimal path. Figure 4 Let's take one method as an example.
[0141] See also Figure 4 , which is a schematic diagram of adding nodes to an initial search tree provided by an embodiment of the present application. Figure 4 The initial search tree in includes 11 nodes. A / B in each node means that the node has been visited B times and the correct number of times is A. For example, the root node has been visited 21 times (or simulated 21 times based on the root node), and the correct number of times is 12, that is, the correct rate is 57.14%. The following uses the four steps of selection, expansion, simulation and backtracking as examples to explain the process of adding nodes to the initial search tree.
[0142] (1) Selection. Search from the root node downwards, and select a "child node that is most worth searching" each time until the search reaches a leaf node. The path from the root node to the leaf node is the optimal path mentioned above, such as Figure 4 3 / 3 of the leaf nodes in .
[0143] The child node most worth continuing to search refers to the node that can find the optimal path. As a possible implementation method, the child node most worth continuing to search can be determined by the upper confidence bound formula, and the node with the largest evaluation value obtained based on the upper confidence bound formula is selected. The upper confidence bound formula is shown in formula (5):
[0144]
[0145] Among them, UCT(s,a) is the evaluation value of node s under reasoning action a; Q(s,a) is the number of correct answers of node s under reasoning action a; N(s,a) is the number of visits to the node under reasoning action a; N parent (s) is the number of visits to the parent node of node s; c is a balancing parameter. The larger the c is, the more inclined it is to perform a breadth search on the initial search tree; the smaller the c is, the more inclined it is to perform a depth search on the initial search tree.
[0146] C2: Expand the leaf nodes in the existing optimal path through the first language model to obtain the child nodes of the leaf nodes.
[0147] The leaf node with the optimal path is the node for continued search in the initial search tree, and expansion is performed based on the leaf node to obtain the child nodes of the leaf node.
[0148] (2) Expansion. If the leaf node of the optimal path is an unexpanded node, that is, the leaf node has no child nodes in the initial search tree, then the leaf node can be further expanded to obtain the child nodes of the leaf node. Figure 4 , first mark the leaf node as 0 / 0.
[0149] The embodiment of the present application does not specifically limit the expansion method. Five methods are used as examples for explanation below and will not be repeated here.
[0150] C3: Add the child nodes of the leaf nodes to the initial search tree through the first language model to obtain a search tree.
[0151] The child nodes of the leaf nodes are added to the initial search tree, so that new nodes are added to the initial search tree, and C1-C3 are repeated continuously until the search tree is obtained. In order to facilitate subsequent calculations, the values of each node can be updated after adding the node, which is explained in detail below.
[0152] (3) Simulation. Starting from the child nodes of the leaf nodes, use the fast rollout policy and other methods to solve the problem sample and obtain an inference result. As a possible implementation method, the problem can be solved multiple times to obtain the average of multiple inference results as the final inference result, or the problem can be solved only once to obtain the inference result. This application does not make specific restrictions on this. Figure 4 , assuming that the inference result based on the child node 0 / 0 of the leaf node is incorrect, the child node 0 / 0 of the leaf node can be updated to 0 / 1.
[0153] (4) Backtracking. Add the inference results obtained by simulation to all its parent nodes, so as to achieve the numerical update of each node, and then realize the initial search tree obtained by this expansion. Figure 4 , all parent nodes of the leaf node's child nodes plus 0 / 1.
[0154] C4: Obtain multiple candidate paths from the search tree through the first language model.
[0155] After repeatedly executing C1-C3, a search tree based on the initial search tree is obtained, and then a candidate path is obtained based on the search tree. The search tree includes multiple leaf nodes, and the path from the root node to the leaf node is a candidate path.
[0156] Therefore, after constructing the initial search tree based on the root node, by selecting the existing optimal path, the child nodes are expanded for the leaf nodes of the existing optimal path, such as by continuously adding child nodes to the initial search tree through selection-expansion-simulation-backtracking, so as to obtain the search tree, and finally obtain the candidate path based on the search tree. This method has a fast search speed, high efficiency in obtaining the search tree, and thus high efficiency in obtaining the candidate path, and the search through the search tree is convenient and fast, and the logic is clear. In addition, the complexity of the search can be controlled by setting the expansion method of each expansion, so as to speed up the search efficiency and improve the accuracy.
[0157] The embodiment of the present application does not specifically limit the expansion method, that is, does not limit the specific implementation method of C2. For example, from multiple preset expansion methods, one expansion method or multiple expansion methods are determined as the target expansion method, and according to the target expansion method, the leaf nodes in the existing optimal path are expanded through the first language model to obtain the child nodes of the leaf nodes.
[0158] Thus, by presetting multiple expansion methods, one expansion method is selected from the multiple expansion methods as the target expansion method, or some of the expansion methods are selected from the multiple expansion methods as the target expansion method, so that the first language model expands the leaf nodes in the optimal path through the target expansion method to obtain the child nodes of the leaf nodes. Thus, the flexibility of node expansion is achieved through multiple expansion methods, and the exploration space is further expanded.
[0159] The following uses the five preset expansion methods as examples to explain them respectively.
[0160] (1) Generate one-step reasoning.
[0161] The reasoning step corresponding to the leaf node in the existing optimal path is continued to be reasoned through the first language model to obtain the next reasoning step of the reasoning step corresponding to the leaf node in the existing optimal path, and the next reasoning step is determined as a child node of the leaf node in the existing optimal path.
[0162] That is, after determining the existing optimal path, continue to reason based on the leaf node of the existing optimal path to obtain the next reasoning step of the reasoning step corresponding to the leaf node, and determine the next reasoning step as the child node of the leaf node in the existing optimal path.
[0163] For example, if the leaf node is the i-th node in the existing optimal path in the initial search tree, it can be represented as s i , then the child nodes of the leaf node can be represented as s i+1 , that is, the i+1th node in the existing optimal path in the initial search tree. At this time, due to the increase in nodes in the initial search tree, the existing optimal path can be re-determined.
[0164] It should be noted that in the process of continuing reasoning, not only the last node in the existing optimal path is needed, but also other nodes in the existing path, so as to learn based on the multiple reasoning steps included in the existing optimal path to obtain the next reasoning step. Therefore, each reasoning needs to input all the previous reasoning processes.
[0165] Therefore, through one reasoning, the next reasoning step is continued for one reasoning step to obtain the next reasoning step, that is, the child node of the leaf node is obtained based on the leaf node, which is not only convenient and fast, but also has high accuracy.
[0166] (2) Generate remaining inferences.
[0167] The first language model is used to continue reasoning the reasoning steps corresponding to the leaf nodes in the existing optimal path, and the remaining reasoning steps after the reasoning steps corresponding to the leaf nodes in the existing optimal path are obtained, that is, all the remaining reasoning steps. The child nodes of the leaf nodes are determined according to the remaining reasoning steps and the reasoning order of the remaining reasoning steps, that is, the reasoning step with the first reasoning order in the remaining reasoning steps is determined as the child node of the leaf node. Further, the reasoning step with the second reasoning order in the remaining reasoning steps is determined as the child node of the child node of the leaf node. And so on, the existing optimal path is completed until the candidate path is obtained.
[0168] That is to say, after determining the existing optimal path, continue to reason based on the leaf nodes of the existing optimal path until the reasoning is completed and the remaining reasoning steps are obtained. Then, based on the remaining reasoning steps and the reasoning order between the remaining reasoning steps, the existing optimal path can be completed to obtain the candidate path.
[0169] Therefore, through one reasoning, all reasoning steps after one reasoning step are reasoned to obtain the remaining reasoning steps, that is, all nodes after the leaf node are obtained based on the leaf node. This is not only convenient and fast, but also compared with method (1), it only needs to input all the previous reasoning processes once, without multiple inputs, which reduces the amount of calculation and improves the calculation efficiency.
[0170] As a possible implementation method, in order to improve the accuracy while increasing the computational complexity, method (1) can be used in the first few inferences on a path to improve the accuracy so that the first language model can learn more of the reasoning logic of this method, and method (2) can be used in the subsequent inferences to improve the computational efficiency.
[0171] (3) Extract sub-questions and answer them.
[0172] The first language model is used to perform semantic understanding on the existing optimal path and the problem sample to obtain the first sub-problem sample, wherein the first sub-problem sample and the part of the problem solved by the existing optimal path constitute the problem sample. In other words, based on the semantic understanding of the problem sample and the semantic understanding of the existing optimal path, the degree to which the existing optimal path solves the problem sample can be clarified, and then the sub-problem, i.e., the first sub-problem sample, can be regenerated based on the degree of solution.
[0173] The first sub-problem sample is inferred through the first language model to obtain the first inference step for solving the first sub-problem sample for the first time, and the first inference step for solving the first sub-problem sample obtained for the first time is determined as the first child node of the leaf node. That is, after obtaining the first sub-problem sample with a more accurate description, the first inference is performed based on the first sub-problem sample to obtain an inference step, and the child node of the leaf node, that is, the first child node, is obtained based on the inference step.
[0174] Therefore, through semantic understanding, the processing progress of the existing optimal path for the problem sample can be clarified, so that the next sub-problem to be solved can be clarified again by generating the first sub-problem sample, thereby improving the accuracy of reasoning. In addition, the first reasoning step obtained through a reasoning is determined as the first child node of the leaf node to improve the accuracy of the first child node, thereby improving the accuracy of the search tree.
[0175] (4) Re-answer the sub-questions.
[0176] The first language model is used to perform semantic understanding on the existing optimal path and problem samples to obtain a first sub-problem sample, the first language model is used to perform a first reasoning on the first sub-problem sample to obtain a first reasoning step for solving the first sub-problem sample, the first language model is used to perform a second reasoning on the first sub-problem sample to obtain a first reasoning step for solving the first sub-problem sample.
[0177] If the first reasoning step for solving the subproblem sample obtained again is different from the first reasoning step for solving the subproblem sample obtained for the first time, that is, the reasoning steps obtained by the two reasonings are different, then the first reasoning step for solving the first subproblem sample obtained again is determined as a child node of the leaf node with the existing optimal path, that is, the second child node. In other words, if the first subproblem sample is reasoned twice through the first language model, the reasoning step obtained by the second reasoning is determined as the first child node.
[0178] Alternatively, if the first reasoning step obtained again for solving the sub-problem sample is different from the first reasoning step obtained for solving the sub-problem sample for the first time, that is, the reasoning steps obtained by the two reasonings are different, then the reasoning step obtained by the first reasoning is determined as the first child node of the leaf node, and the reasoning step obtained by the second reasoning is determined as the second child node of the leaf node. The first child node and the second child node are different child nodes, that is, the reasoning results obtained by the two reasonings are both determined as child nodes of the leaf node, and the two are brother nodes of each other.
[0179] Therefore, by reasoning on the first sub-problem sample twice, if the reasoning steps obtained by the two reasonings are different, not only can the reasoning steps obtained by the second reasoning be used as the child nodes of the leaf node, so as to deepen the first language model's understanding of the first sub-problem sample through multiple reasonings, thereby obtaining more accurate reasoning steps, but the reasoning steps obtained by the two reasonings can also be used as child nodes of the leaf node, thereby expanding the number of child nodes included in the search tree, so as to expand the exploration space learned by the language model.
[0180] (5) Regenerate the sub-questions and answer them.
[0181] After using the first language model to perform semantic understanding on the existing optimal path and problem samples to obtain the first sub-problem sample, the first language model is used again to perform semantic understanding on the existing optimal path and problem samples to obtain the second sub-problem sample. The second sub-problem sample and the part of the problem you solved on the existing optimal path constitute the problem sample, and the first sub-problem sample and the second sub-problem sample are different sub-problems.
[0182] The second sub-problem sample is inferred through the first language model to obtain a first inference step for solving the second sub-problem sample, and the first inference step for solving the second sub-problem sample is determined as the third child node of the leaf node. The third child node and the first child node are different nodes, and the search tree may include only the third child node, or may include the third child node and the first child node, which is not specifically limited in this application.
[0183] Therefore, through semantic understanding, it is possible to clearly understand the processing progress of the existing optimal path for the problem sample. After generating the first sub-problem sample, a sub-problem is generated again, that is, the second sub-problem sample. By generating sub-problems multiple times, the next sub-problem to be solved is clearly identified, thereby improving the accuracy of reasoning and the accuracy of the search tree.
[0184] In order to facilitate further understanding of the technical solution provided in the embodiments of the present application, the data processing method provided in the embodiments of the present application is introduced as a whole by taking the execution subject of the data processing method as a server as an example, see S1-S7 for details.
[0185] S1: Get question samples.
[0186] S2: Reasoning the question sample based on the first language model to obtain multiple candidate paths.
[0187] The reasoning process can refer to the aforementioned B1-B4 or C1-C4 method, and the following is explained using the C1-C4 method.
[0188] (1) Initialize the search tree. Obtain the root node of the search tree based on the problem sample, as shown in the above formula (4).
[0189] (2) Generate child nodes of the root node. Infer the question sample according to the first language model to obtain multiple child nodes of the root node, and then add the multiple child nodes to the root node to obtain an updated initial search tree.
[0190] (3) Selection. Search from the root node downwards, and each time you search, you can select a "child node that is most worth continuing to search" based on formula (5) until you reach a leaf node.
[0191] (4) Extension. If the leaf node is an unexpanded node, the leaf node can be further expanded to obtain the child nodes of the leaf node. The expansion method can be selected from the above methods (1) to (5), which will not be repeated here.
[0192] (5) Simulation. Starting from the child nodes of the leaf node, use the fast move strategy to simulate and solve the problem sample and obtain an inference result.
[0193] (6) Backtracking. Add the inference results obtained by simulation to all its parent nodes, so as to achieve the numerical update of each node, and then realize the initial search tree obtained by this expansion. The search tree can be obtained by continuously going through (3)-(6).
[0194] (7) Obtain candidate paths based on the search tree. Each leaf node from the root node to the search tree can be used as a candidate path, and each candidate path can be expressed as the above formula (1).
[0195] (8) Calculate the accuracy of each candidate path, such as whether the game is won or not. For details, see formula (3).
[0196] S3: Cover some reasoning steps in each candidate path to obtain the covered paths corresponding to each candidate path.
[0197] For example, the second half of the reasoning steps in each candidate path are randomly covered to obtain a covered path, which can be expressed as formula (6):
[0198]
[0199] Where t(i) is the covered path corresponding to the i-th candidate path; x represents the root node; s i-1 It indicates that the i-1th reasoning step in the middle (referred to as the middle reasoning step), that is, the di-th reasoning step of the candidate path corresponding to formula (6) is covered.
[0200] S4: Based on the reasoning order indicated by each covered path, the second language model is used to reason about the question sample to obtain the paths to be verified corresponding to each covered path.
[0201] Based on the reasoning order indicated by each covered path, the second language model is used to reason the problem sample, and the subsequent reasoning steps corresponding to each covered path are obtained, and the path to be verified is obtained. For example, each covered path is input into the second language model respectively, so that the second language model outputs the subsequent reasoning steps of each covered path, and multiple paths to be verified are obtained.
[0202] The training data for training the first language model and the training data for training the second language model are not exactly the same, so the reasoning ability of the first language model and the second language model are not exactly the same. The first language model can be regarded as a generator and the second language model as a discriminator, so that the first language model can be supplemented by the second language model and the reasoning ability of the first language model can be improved through the generation and verification mechanism.
[0203] S5: Determine a result path from multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path.
[0204] Determine the consistency between the path to be verified and its corresponding candidate path, that is, the degree of similarity. Based on the degree of similarity between the two and the accuracy of the candidate path, select the candidate path with the highest score from multiple candidate paths as the result path. For details, see formula (2).
[0205] S6: Determine the training samples according to the problem samples, target path and result path.
[0206] The target path includes a target candidate path and a target path to be verified. The target candidate path is a candidate path with a same degree greater than a same threshold among multiple candidate paths, and the target path to be verified is a path to be verified with a same degree greater than the same threshold among multiple paths to be verified.
[0207] S7: According to the reasoning order indicated by the target path in the training sample, the question sample in the training sample is reasoned by the first language model to obtain multiple reasoning paths and a predicted optimal path among the multiple reasoning paths.
[0208] The first language model is trained again based on the training samples, so that based on the target path in the training samples, not only the exploration space learned by the first language model can be expanded, but also the effectiveness of the exploration space can be ensured.
[0209] S8: According to the difference between the predicted optimal path and the result path, the model parameters of the first language model are adjusted to obtain an optimized language model.
[0210] In addition, S1-S8 can be packaged into a reusable plug-in and provided to downstream business departments. Through efficient interface design, the plug-in can be quickly integrated into different business scenarios to meet business needs. Figure 5 , which is a schematic diagram of a plug-in application provided by an embodiment of the present application to a business scenario. Figure 5 In the example, the plug-in is applied in three business scenarios, and three business lines are generated for the three business scenarios.
[0211] Moreover, corresponding experiments were conducted on the news business, that is, the training data of the original language model was expanded based on the data processing method provided in the embodiment of the present application, and the original language model was trained again, which resulted in improvements in different application scenarios of the news business. The experimental results are shown in Table 1.
[0212] Table 1
[0213]
[0214] Therefore, by introducing the dual-model collaboration mechanism of generator and discriminator, the reasoning process is decoupled into two major steps: generating paths obtained by diversified reasoning and mutual verification. The generator uses the selection-expansion-simulation-backtracking method to generate a search tree and construct multiple candidate paths, thereby expanding the exploration space; the discriminator completes some reasoning steps to verify consistency and screen out higher quality result paths. This not only effectively improves the accuracy of the reasoning results, but also solves the problem of insufficient self-evaluation ability of small language models.
[0215] By introducing five diversified reasoning methods, including one-step reasoning, problem decomposition, and rephrasing, the coverage of the search space has been significantly expanded. The reasoning process of the language model is more similar to human reasoning behavior, which significantly improves the robustness and generalization of reasoning, and can achieve a significant improvement in reasoning ability without relying on labeled data or external supervision models.
[0216] With respect to the data processing method described above, the present application also provides a corresponding data processing device so that the above data processing method can be applied and implemented in practice.
[0217] See also Figure 6 , which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. Figure 6 As shown, the data processing device 600 includes: an acquisition unit 601, an inference unit 602, a covering unit 603, a completion unit 604, a selection unit 605 and a determination unit 606;
[0218] The acquisition unit 601 is used to acquire question samples;
[0219] The reasoning unit 602 is used to reason the question sample according to the first language model to obtain multiple candidate paths, where different candidate paths correspond to different reasoning methods for solving the question sample, and the candidate paths include multiple reasoning steps with a reasoning order;
[0220] The covering unit 603 is used to cover part of the reasoning steps in each candidate path to obtain a covered path corresponding to each candidate path;
[0221] The completion unit 604 is used to perform reasoning on the question sample through the second language model based on the reasoning order indicated by each of the covered paths, to obtain paths to be verified corresponding to each of the covered paths, wherein the paths to be verified are paths obtained by reasoning the covered reasoning steps in the covered paths, and the first language model and the second language model are different language models;
[0222] The selection unit 605 is used to determine a result path from the multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path;
[0223] The determining unit 606 is used to determine a training sample according to the question sample, the multiple candidate paths and the result path, and the training sample is used to train the reasoning ability of the language model.
[0224] It can be seen from the above technical solution that an embodiment of the present application provides a data processing device, including an acquisition unit, an inference unit, a masking unit, a completion unit, a selection unit and a determination unit. The problem sample is subjected to diversified inference through the first language model to obtain multiple candidate paths, and some of the reasoning steps of the candidate paths are masked to obtain the masked path. Based on the indication of the masked path, the problem sample is subjected to diversified inference through the second language model, so as to re-reason the masked reasoning steps in the masked path to obtain the path to be verified. According to the consistency between the candidate path and the corresponding path to be verified, a result path is obtained from the multiple candidate paths, and the accuracy of the reasoning result corresponding to the result path is relatively high, so that a training sample is obtained according to the problem sample, the multiple candidate paths and the result path.
[0225] Therefore, based on the multiple candidate paths in the training sample, the language model can learn a larger exploration space, improving the generalization ability and robustness of the language model. Moreover, by introducing the first language model and the second language model with similar capabilities, the two can verify each other and complement each other. The second language model helps the first language model recognize errors that the first language model cannot recognize, reducing the probability of large model hallucinations and improving the accuracy of the result path, so that the trained language model can learn a more accurate result path from multiple reasoning processes, improving the accuracy of the subsequent reasoning results based on the language model.
[0226] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0227] Reasoning the question sample according to the first language model to obtain a pending candidate path, wherein the pending candidate path includes a plurality of reasoning steps having a reasoning order;
[0228] Determine a synonymous step according to the reasoning step, wherein the similarity between the synonymous step and the reasoning step is greater than a similarity threshold;
[0229] A search tree is obtained according to the pending candidate path and the synonymous step, wherein the multiple reasoning steps included in the pending candidate path are nodes in the search tree, the root node of the search tree is a node obtained based on the problem sample, the parent node of the search tree is the previous reasoning step of the child node of the parent node, and the synonymous step is a sibling node of the reasoning step corresponding to the synonymous step;
[0230] According to the search tree, the multiple candidate paths are determined, where the candidate paths are paths corresponding to the root node to the leaf nodes.
[0231] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0232] Selecting an existing optimal path from an initial search tree through the first language model, wherein each node of the initial search tree corresponds to one of the reasoning steps, the root node of the initial search tree is a node obtained based on the problem sample, the parent node of the initial search tree is a previous reasoning step of a child node of the parent node, and the existing optimal path is a path with the highest accuracy obtained by solving the problem sample among multiple paths included in the initial search tree;
[0233] Expanding the leaf nodes in the existing optimal path by using the first language model to obtain child nodes of the leaf nodes;
[0234] Adding the child nodes of the leaf nodes to the initial search tree by using the first language model to obtain the search tree;
[0235] The multiple candidate paths are obtained from the search tree by using the first language model.
[0236] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0237] The first language model is used to continue reasoning the reasoning step corresponding to the leaf node in the existing optimal path to obtain the next reasoning step of the reasoning step corresponding to the leaf node in the existing optimal path, and the next reasoning step is determined as a child node of the leaf node in the existing optimal path.
[0238] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0239] The first language model is used to continue reasoning the reasoning steps corresponding to the leaf nodes in the existing optimal path to obtain the remaining reasoning steps after the reasoning steps corresponding to the leaf nodes in the existing optimal path, and the child nodes of the leaf nodes are determined according to the remaining reasoning steps and the reasoning order of the remaining reasoning steps.
[0240] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0241] Performing semantic understanding on the existing optimal path and the question sample by using the first language model to obtain a first sub-question sample;
[0242] Reasoning the first sub-problem sample by using the first language model to obtain a first reasoning step for solving the first sub-problem sample;
[0243] The first reasoning step obtained for solving the first sub-problem sample for the first time is determined as the first child node of the leaf node.
[0244] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0245] Reasoning the first sub-problem sample again by using the first language model to obtain a first reasoning step for solving the first sub-problem sample again;
[0246] If the first reasoning step obtained again for solving the sub-problem sample is different from the first reasoning step obtained for solving the sub-problem sample for the first time, the first reasoning step obtained again for solving the sub-problem sample is used to determine the second child node of the leaf node, and the first child node and the second child node are different child nodes; or
[0247] The first inference step for solving the sub-problem sample is used to replace the first child node of the leaf node.
[0248] As a possible implementation manner, the reasoning unit 602 is specifically configured to:
[0249] Performing semantic understanding on the existing optimal path and the problem sample by using the first language model to obtain a second sub-problem sample, where the second sub-problem sample and the first sub-problem sample are different sub-problem samples;
[0250] Reasoning the second sub-problem sample by using the first language model to obtain a first reasoning step for solving the second sub-problem sample;
[0251] The first reasoning step for solving the second sub-problem sample is determined as a third child node of the leaf node, and the third child node is a different child node from the first child node.
[0252] As a possible implementation manner, the determining unit 606 is specifically configured to:
[0253] Acquire a target path, wherein the target path includes the target candidate path, or the target path includes the target candidate path and a target path to be verified, wherein the target candidate path is a candidate path among the multiple candidate paths whose degree of similarity is greater than a same threshold, and the target path to be verified is a path to be verified among the multiple paths to be verified whose degree of similarity is greater than the same threshold;
[0254] The training sample is determined according to the problem sample, the target path and the result path.
[0255] As a possible implementation manner, the device further includes a training unit, which is used to:
[0256] According to the reasoning order indicated by the target path in the training sample, reasoning the problem sample in the training sample by using the language model to obtain multiple reasoning paths and a predicted optimal path among the multiple reasoning paths;
[0257] According to the difference between the predicted optimal path and the result path, the model parameters of the language model are adjusted to obtain an optimized language model.
[0258] As a possible implementation, the covering unit 603 is specifically used for:
[0259] The last n reasoning steps or the middle m reasoning steps in each of the candidate paths are covered to obtain covered paths corresponding to each of the candidate paths, where n is an integer greater than 0 and m is an integer greater than 0.
[0260] The present application also provides a computer device, which may be a server or a terminal device. The following will introduce the computer device provided by the present application from the perspective of hardware entity. Figure 7 The following is a schematic diagram of the server structure. Figure 8 Shown is a schematic diagram of the structure of the terminal equipment.
[0261] See also Figure 7 , which is a schematic diagram of a server structure provided in an embodiment of the present application. The server 1400 may have relatively large differences due to different configurations or performances, and may include one or more processors 1422, such as central processing units (CPU), memory 1432, one or more application programs 1442 or storage media 1430 (such as one or more massive storage devices) for data 1444. Among them, the memory 1432 and the storage medium 1430 may be temporary storage or permanent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the processor 1422 may be configured to communicate with the storage medium 1430 to execute a series of instruction operations in the storage medium 1430 on the server 1400.
[0262] The server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.
[0263] The steps performed by the server in the above embodiment can be based on the Figure 7 The server structure shown.
[0264] The processor 1422 is used to perform the following steps:
[0265] Get sample questions;
[0266] Reasoning the problem sample according to the first language model to obtain multiple candidate paths, where different candidate paths correspond to different reasoning methods for solving the problem sample, and the candidate paths include multiple reasoning steps with a reasoning order;
[0267] Covering part of the reasoning steps in each of the candidate paths to obtain covered paths corresponding to each of the candidate paths;
[0268] Based on the reasoning order indicated by each of the covered paths, the question sample is reasoned through the second language model to obtain paths to be verified corresponding to each of the covered paths, wherein the paths to be verified are paths obtained by reasoning the covered reasoning steps in the covered paths, and the first language model and the second language model are different language models;
[0269] Determine a result path from the multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path;
[0270] A training sample is determined according to the question sample, the multiple candidate paths and the result path, and the training sample is used to train the reasoning ability of the language model.
[0271] Optionally, the processor 1422 may also execute the method steps of any specific implementation of the data processing method in the embodiments of the present application.
[0272] See also Figure 8 , which is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application. Taking the terminal device as a smart phone as an example, Figure 8 The block diagram of the partial structure of the smart phone is shown, and the smart phone includes: a radio frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a wireless fidelity (WiFi) module 1570, a processor 1580, and a power supply 1590. Those skilled in the art can understand that Figure 8 The structure of the smartphone shown in the figure does not constitute a limitation of the smartphone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0273] Combine the following Figure 8 A detailed introduction to the various components of a smartphone:
[0274] The RF circuit 1510 may be used for receiving and sending signals during information transmission or calls. In particular, after receiving the downlink information from the base station, it is sent to the processor 1580 for processing; in addition, the designed uplink data is sent to the base station.
[0275] The memory 1520 may be used to store software programs and modules. The processor 1580 implements various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 1520 .
[0276] The input unit 1530 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the smartphone. Specifically, the input unit 1530 may include a touch panel 1531 and other input devices 1532. The touch panel 1531, also known as a touch screen, can collect user touch operations on or near it and drive the corresponding connection device according to a pre-set program. In addition to the touch panel 1531, the input unit 1530 may also include other input devices 1532. Specifically, other input devices 1532 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, and the like.
[0277] The display unit 1540 may be used to display information input by the user or information provided to the user and various menus of the smartphone. The display unit 1540 may include a display panel 1541, and the display panel 1541 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0278] The smartphone may also include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. As for other sensors that may be configured in the smartphone, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., they will not be described in detail here.
[0279] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the smartphone. The audio circuit 1560 can transmit the received audio data to the speaker 1561 after converting the received audio data into an electrical signal, which is converted into a sound signal for output; on the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and converted into audio data, and then the audio data is output to the processor 1580 for processing, and then sent to another smartphone through the RF circuit 1510, or the audio data is output to the memory 1520 for further processing.
[0280] The processor 1580 is the control center of the smartphone, and uses various interfaces and lines to connect various parts of the entire smartphone, and executes various functions of the smartphone and processes data by running or executing software programs and / or modules stored in the memory 1520, and calling data stored in the memory 1520. Optionally, the processor 1580 may include one or more processing units.
[0281] The smart phone also includes a power supply 1590 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption management.
[0282] Although not shown, the smartphone may also include a camera, a Bluetooth module, etc., which will not be described in detail here.
[0283] In the embodiment of the present application, the memory 1520 included in the smart phone can store a computer program and transmit the computer program to the processor.
[0284] The processor 1580 included in the smart phone can execute the data processing method provided in the above embodiment according to the instructions in the computer program.
[0285] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program, wherein the computer program is used to execute the data processing method provided in the above embodiment.
[0286] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the data processing method provided in various optional implementations of the above aspects.
[0287] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the above program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the above storage medium can be at least one of the following media: read-only memory (English: Read-Only Memory, abbreviated: ROM), RAM, magnetic disk or optical disk, etc. Various media that can store computer programs.
[0288] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0289] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0290] It should be noted that each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, in which the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0291] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a technician familiar with the technical field within the technical scope disclosed in the present application should be included in the protection scope of the present application. Based on the implementation methods provided in the above aspects, the present application can also be further combined to provide more implementation methods. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A data processing method, characterized in that: The method comprises: Get sample questions; Reasoning the problem sample according to the first language model to obtain multiple candidate paths, where different candidate paths correspond to different reasoning methods for solving the problem sample, and the candidate paths include multiple reasoning steps with a reasoning order; Covering part of the reasoning steps in each of the candidate paths to obtain covered paths corresponding to each of the candidate paths; Based on the reasoning order indicated by each of the covered paths, the question sample is reasoned through the second language model to obtain paths to be verified corresponding to each of the covered paths, wherein the paths to be verified are paths obtained by reasoning the covered reasoning steps in the covered paths, and the first language model and the second language model are different language models; Determine a result path from the multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path; A training sample is determined according to the question sample, the multiple candidate paths and the result path, and the training sample is used to train the reasoning ability of the language model.
2. The method according to claim 1, characterized in that: The reasoning on the question sample according to the first language model to obtain multiple candidate paths includes: Reasoning the question sample according to the first language model to obtain a pending candidate path, wherein the pending candidate path includes a plurality of reasoning steps with a reasoning order; Determine a synonymous step according to the reasoning step, wherein the similarity between the synonymous step and the reasoning step is greater than a similarity threshold; A search tree is obtained according to the pending candidate path and the synonymous step, wherein the multiple reasoning steps included in the pending candidate path are nodes in the search tree, the root node of the search tree is a node obtained based on the problem sample, the parent node of the search tree is the previous reasoning step of the child node of the parent node, and the synonymous step is a sibling node of the reasoning step corresponding to the synonymous step; According to the search tree, the multiple candidate paths are determined, where the candidate paths are paths corresponding to the root node to the leaf node.
3. The method according to claim 1, characterized in that The reasoning on the question sample according to the first language model to obtain multiple candidate paths includes: Selecting an existing optimal path from an initial search tree through the first language model, wherein each node of the initial search tree corresponds to one of the reasoning steps, the root node of the initial search tree is a node obtained based on the problem sample, the parent node of the initial search tree is a previous reasoning step of a child node of the parent node, and the existing optimal path is a path with the highest accuracy obtained by solving the problem sample among multiple paths included in the initial search tree; Expanding the leaf nodes in the existing optimal path by using the first language model to obtain child nodes of the leaf nodes; Adding the child nodes of the leaf nodes to the initial search tree by using the first language model to obtain the search tree; The multiple candidate paths are obtained from the search tree by using the first language model.
4. The method according to claim 3, characterized in that The step of expanding the leaf nodes in the existing optimal path by using the first language model to obtain child nodes of the leaf nodes includes: The first language model is used to continue reasoning the reasoning step corresponding to the leaf node in the existing optimal path to obtain the next reasoning step of the reasoning step corresponding to the leaf node in the existing optimal path, and the next reasoning step is determined as a child node of the leaf node in the existing optimal path.
5. The method according to claim 3, characterized in that: The step of expanding the leaf nodes in the existing optimal path by using the first language model to obtain child nodes of the leaf nodes includes: The first language model is used to continue reasoning the reasoning steps corresponding to the leaf nodes in the existing optimal path to obtain the remaining reasoning steps after the reasoning steps corresponding to the leaf nodes in the existing optimal path, and the child nodes of the leaf nodes are determined according to the remaining reasoning steps and the reasoning order of the remaining reasoning steps.
6. The method according to claim 4, characterized in that The step of expanding the leaf nodes in the existing optimal path by using the first language model to obtain child nodes of the leaf nodes includes: Performing semantic understanding on the existing optimal path and the question sample by using the first language model to obtain a first sub-question sample; Reasoning the first sub-problem sample by using the first language model to obtain a first reasoning step for solving the first sub-problem sample; The first reasoning step obtained for solving the first sub-problem sample for the first time is determined as the first child node of the leaf node.
7. The method according to claim 6, characterized in that The method further comprises: Reasoning the first sub-problem sample again by using the first language model to obtain a first reasoning step for solving the first sub-problem sample again; If the first reasoning step obtained again for solving the sub-problem sample is different from the first reasoning step obtained for solving the sub-problem sample for the first time, the first reasoning step obtained again for solving the sub-problem sample is used to determine the second child node of the leaf node, and the first child node and the second child node are different child nodes; or The first inference step for solving the sub-problem sample is used to replace the first child node of the leaf node.
8. The method according to claim 6, characterized in that The method further comprises: Performing semantic understanding on the existing optimal path and the problem sample by using the first language model to obtain a second sub-problem sample, where the second sub-problem sample and the first sub-problem sample are different sub-problem samples; Reasoning the second sub-problem sample by using the first language model to obtain a first reasoning step for solving the second sub-problem sample; The first reasoning step for solving the second sub-problem sample is determined as a third child node of the leaf node, and the third child node is a different child node from the first child node.
9. The method according to claim 1, characterized in that: The determining of a training sample according to the problem sample, the multiple candidate paths, and the result path includes: Acquire a target path, wherein the target path includes the target candidate path, or the target path includes the target candidate path and a target path to be verified, wherein the target candidate path is a candidate path among the multiple candidate paths whose degree of similarity is greater than a same threshold, and the target path to be verified is a path to be verified among the multiple paths to be verified whose degree of similarity is greater than the same threshold; The training sample is determined according to the problem sample, the target path and the result path.
10. The method according to claim 9, characterized in that The method further comprises: According to the reasoning order indicated by the target path in the training sample, reasoning the problem sample in the training sample by using the language model to obtain multiple reasoning paths and a predicted optimal path among the multiple reasoning paths; According to the difference between the predicted optimal path and the result path, the model parameters of the language model are adjusted to obtain an optimized language model.
11. The method according to claim 1, characterized in that: The step of masking some of the reasoning steps in each of the candidate paths to obtain masked paths corresponding to each of the candidate paths includes: The last n reasoning steps or the middle m reasoning steps in each of the candidate paths are covered to obtain covered paths corresponding to each of the candidate paths, where n is an integer greater than 0 and m is an integer greater than 0.
12. A data processing device, characterized in that: The device comprises: an acquisition unit, an inference unit, a covering unit, a completion unit, a selection unit and a determination unit; The acquisition unit is used to acquire question samples; The reasoning unit is used to reason the problem sample according to the first language model to obtain multiple candidate paths, where different candidate paths correspond to different reasoning methods for solving the problem sample, and the candidate paths include multiple reasoning steps with a reasoning order; The covering unit is used to cover part of the reasoning steps in each of the candidate paths to obtain the covered paths corresponding to each of the candidate paths; The completion unit is configured to perform reasoning on the question sample through a second language model based on the reasoning order indicated by each of the covered paths, to obtain paths to be verified corresponding to each of the covered paths, wherein the paths to be verified are paths obtained by reasoning on the covered reasoning steps in the covered paths, and the first language model and the second language model are different language models; The selection unit is used to determine a result path from the multiple candidate paths according to the degree of similarity between the candidate path corresponding to the covered path and the path to be verified corresponding to the covered path; The determination unit is used to determine a training sample according to the question sample, the multiple candidate paths and the result path, and the training sample is used to train the reasoning ability of the language model.
13. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 11 according to the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the method is executed on a computer device, the computer device is enabled to execute the method according to any one of claims 1 to 11.
Citation Information
Cited By
Tax document verification method and device, electronic equipment, computer readable storage medium and computer program product
CN121457464A