Data Processing Method and Apparatus

The data processing method enhances LLMs' reasoning capabilities by expanding simple question-answer pairs into detailed reasoning paths, addressing the lack of effective training data for complex problems and improving logical accuracy.

CN119990333BActive Publication Date: 2025-07-15ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510461488.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-15
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

When existing large language models deal with problems that require long-link logical reasoning, they lack effective training data, especially the intermediate reasoning steps of mathematical logic questions and logical reasoning questions, which leads to insufficient model reasoning capabilities.

Method used

By determining the inference steps and answers corresponding to the initial problem, after performing correctness verification, target data containing detailed inference steps is generated, and multiple inference paths are generated using autoregression. Combined with the corrective model error correction steps, it is expanded into long inference link data for model training.

Benefits of technology

It realizes effective training data expansion of long inference ability of large language models, improves the inference accuracy and logical clarity of the model in complex problems, and enhances the model's reflection and reasoning ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990333B_ABST
    Figure CN119990333B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a data processing method and apparatus. The method includes: determining at least one problem reasoning step corresponding to an initial problem, and a reasoning answer corresponding to the initial problem obtained based on the at least one problem reasoning step; performing a correctness verification on the at least one problem reasoning step and the reasoning answer according to the initial answer corresponding to the initial problem to obtain a verification result; obtaining target data according to the verification result, the initial problem, and the at least one problem reasoning step, where the target data includes the initial problem, target reasoning steps, and a target answer corresponding to the initial problem; compared with a simple question-and-answer pair including the initial problem and the initial answer, the target data adds target reasoning steps, thereby realizing the extension of the long reasoning link for the simple question-and-answer pair, and providing effective training data for improving the model reasoning ability subsequently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence, and particularly to a data processing method. One or more embodiments of this specification also relate to a data processing device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art. Background Art

[0002] In recent years, technologies related to large language models (LLMs) have been continuously developing, and the capabilities of LLMs have also been continuously enhanced. However, for some problems that require strong dependence on logical reasoning and long-chain reasoning to obtain results, such as common mathematical reasoning and logical reasoning problems, there is still room for improvement.

[0003] However, when enhancing the model's reasoning ability by fine-tuning the model, there is often a lack of effective training data because many mathematical logic problems are multiple-choice questions, true or false questions, or directly give the final answer without the intermediate key reasoning steps, and cannot be directly used as training data for the model's reasoning ability. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a model training method, a data processing device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, a data processing method is provided, including:

[0006] Determine at least one problem reasoning step corresponding to the initial problem, and the reasoning answer corresponding to the initial problem obtained based on the at least one problem reasoning step;

[0007] According to the initial answer corresponding to the initial problem, perform a correctness verification on the at least one problem reasoning step and the reasoning answer to obtain a verification result;

[0008] According to the verification result, the initial problem, and the at least one problem reasoning step, obtain target data, where the target data includes the initial problem, the reasoning step, and the target answer corresponding to the initial problem.

[0009] According to the second aspect of the embodiments of this specification, a data processing device is provided, including:

[0010] The inference module is configured to determine at least one problem inference step corresponding to the initial problem, and an inference answer corresponding to the initial problem obtained based on the at least one problem inference step;

[0011] The verification module is configured to perform a correctness verification on the at least one problem inference step and the inference answer according to the initial answer corresponding to the initial problem, and obtain a verification result;

[0012] The obtaining module is configured to obtain target data according to the verification result, the initial problem, and the at least one problem inference step, where the target data includes the initial problem, the inference step, and the target answer corresponding to the initial problem.

[0013] According to a third aspect of the embodiments of the present specification, there is provided a model training method, including:

[0014] Determine the model to be adjusted;

[0015] Train the model to be adjusted with the target data to obtain a target model, where the target data is obtained by the above data processing method.

[0016] According to a fourth aspect of the embodiments of the present specification, there is provided a computing device, including:

[0017] A memory and a processor;

[0018] Wherein, the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above data processing method are implemented.

[0019] According to a fifth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above data processing method are implemented.

[0020] According to a sixth aspect of the embodiments of the present specification, there is provided a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above data processing method are implemented.

[0021] The data processing method provided by an embodiment of this specification can obtain inference path data including problem inference steps corresponding to an initial problem by determining at least one problem inference step corresponding to the initial problem and an inference answer corresponding to the initial problem obtained based on the at least one problem inference step. And when performing a correctness verification on the at least one problem inference step and the inference answer according to the initial answer corresponding to the initial problem, correct target data can be obtained through the obtained verification result, the initial problem, and the at least one problem inference step. Wherein, the target data includes the initial problem, the inference steps, and the target answer corresponding to the initial problem. Compared with a simple question-and-answer pair including the initial problem and the initial answer, the target data adds target inference steps, thereby realizing the extension of the long inference link for the simple question-and-answer pair, and providing effective training data for improving the model inference ability subsequently. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic diagram of a scenario of a data processing method provided by an embodiment of this specification;

[0023] Figure 2 is a flowchart of a data processing method provided by an embodiment of this specification;

[0024] Figure 3 is a schematic diagram of the chained data construction process of a data processing method provided by an embodiment of this specification;

[0025] Figure 4 is a schematic diagram of the tree-shaped data construction process of a data processing method provided by an embodiment of this specification;

[0026] Figure 5 is a schematic diagram of the structure of a data processing device provided by an embodiment of this specification;

[0027] Figure 6 is a block diagram of the structure of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.

[0029] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any or all possible combinations of one or more of the associated listed items.

[0030] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0031] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0032] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability, such as large language models (LLMs), multi-modal pre-training models, etc.

[0033] When large models are applied in practice, only a small number of samples are needed to fine-tune the pre-trained models for application in different tasks. Large models can be widely applied in fields such as natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0034] First, explain the noun terms related to one or more embodiments of this specification.

[0035] few-shots: Some example data used to provide reference for the question-answering model so that the question-answering model can obtain the final answer based on the reasoning steps.

[0036] shortcut sample: A shortcut sample, which in the embodiments of this specification refers to the data that directly infers the correct answer.

[0037] backtracking sample: A backtracking sample, which in the embodiments of this specification refers to the data where there are some errors in the reasoning process, but after the model's self-reflection and error correction, the correct answer is finally inferred.

[0038] Traditional methods of using fine-tuning models to enhance the model's reasoning ability involve a large amount of manual annotation work, which not only requires expensive human and time costs but also requires further proofreading of the consistency of manual annotation. In addition, since many mathematical problems, logical problems, etc. are multiple-choice questions, true or false questions, or directly give the final answer without the intermediate key reasoning steps, many existing logical question-and-answer data cannot be well utilized.

[0039] For example, a common question-and-answer data example is as follows:

[0040] Question: "There are 15 different circles. What is the maximum possible number of intersection points of the circles?"

[0041] Options: (A) 390 (B) 100 (C) 110 (D) 180 (E) 210. Answer: E.

[0042] Long reasoning data example:

[0043] Question: "There are 15 different circles. What is the maximum possible number of intersection points of the circles?"

[0044] Options: (A) 390 (B) 100 (C) 110 (D) 180 (E) 210.

[0045] Answer: Analysis: Each pair of distinct circles can have at most two intersection points. For 15 distinct circles, the maximum possible number of intersection points can be calculated using combinatorics.

[0046] 1. Calculate the number of pairs of circles: 15×14 / 2 = 105 pairs.

[0047] 2. Each pair of circles can have at most 2 intersection points, so the total number of intersection points is 105×2 = 210.

[0048] Therefore, the maximum possible number of intersection points is 210, so the answer is (E) 210.

[0049] Ordinary Q&A data example:

[0050] Question: "What is the estuary of the water body where the Bartram's Bridge is located?"

[0051] Answer: The Delaware River.

[0052] Long reasoning data example:

[0053] Question: "What is the estuary of the water body where the Bartram's Bridge is located?"

[0054] Answer: The estuary of the water body where the Bartram's Bridge is located, which is the estuary of Crum Creek, is the Delaware River in Edystone, Pennsylvania.

[0055] From the above examples, it can be seen that there is a lot of existing data, such as math problems, code problems, logic problems, etc. Many of them are simple answers or judgment multiple-choice questions, etc. The answers are too short, which may lead to the situation of the model guessing answers and is not conducive to understanding the complete reasoning process; for scenarios with high requirements for reasoning ability, the reasoning process and reasoning result are equally important. Therefore, the embodiments of this specification propose a data processing method for automatically expanding ordinary Q&A data into long reasoning link data and then fine-tuning the model to efficiently improve the long reasoning ability of the model.

[0056] To solve the above technical problems, in this specification, a data processing method is provided. This specification also relates to a model training method, a data processing device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.

[0057] See Figure 1 , Figure 1 which shows a schematic diagram of a scenario of a data processing method provided according to an embodiment of this specification.

[0058] Specifically, this data processing method is applied to a data processing system, which includes an edge device 102 and a server 104. The edge device 102 is used to send an initial question and an initial answer to the server 104. In practical applications, a user can input the initial question and the initial answer in the edge device 102 in the form of text or voice. If the voice form is adopted, the edge device 102 will also include corresponding voice processing parts, such as modules for voice parsing, voice-to-text conversion, voice synthesis, etc., which are used to convert the question input by the user through voice into text. This specification does not limit this.

[0059] An answer question model is trained in the server 104. When the server 104 receives the initial question sent by the edge device 102, at least one question reasoning step corresponding to the initial question is determined by using the answer question model, and an inference answer corresponding to the initial question obtained based on the at least one question reasoning step is determined. According to the initial answer corresponding to the initial question, the at least one question reasoning step and the inference answer are verified for correctness to obtain a verification result. According to the verification result, the initial question and the at least one question reasoning step, target data is obtained, where the target data includes the initial question, target reasoning steps, and the target answer corresponding to the initial question. The target data can be returned to the edge device 102. Or when a model training request is carried in the request of the edge device 102, the model to be adjusted can be trained by using the target data to obtain a target model, and then the call interface corresponding to the target model or the target model is returned to the edge device 102, so that the edge device 102 can call the target model through the call interface, or deploy the target model in the edge device 102.

[0060] The edge device 102 may include a browser, an APP (Application), or a web application such as an H5 (Hyper Text Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The edge device may be developed based on the software development kit (SDK) of the corresponding service provided by the server side, such as developed based on the real-time communication (RTC) SDK. The edge device may be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device may have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, data processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0061] The server 104 can be understood as a server that provides various services, including a physical server and a cloud server. For example, a server that provides communication services for multiple clients, or a server for background training that provides support for the models used on the client side, or a server that processes the data sent by the client, etc. It should be noted that the server 104 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server 104 can also be a server of a distributed system, or a server combined with a blockchain. The server 104 can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0062] It is worth noting that the data processing method provided in the embodiments of this specification can be executed by the server 104. In other embodiments of this specification, the question-and-answer model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions as the server 104, and thus execute the data processing method provided in the embodiments of this specification; in other embodiments, the data processing method provided in the embodiments of this specification can also be jointly executed by the edge device 102 and the server 104.

[0063] The data processing method provided by the embodiments of this specification can obtain the inference path data including problem inference steps corresponding to the initial problem by determining at least one problem inference step corresponding to the initial problem and the inference answer corresponding to the initial problem obtained based on the at least one problem inference step. And when verifying the correctness of the at least one problem inference step and the inference answer according to the initial answer corresponding to the initial problem, the correct target data can be obtained through the obtained verification result, the initial problem, and the at least one problem inference step. Wherein, the target data includes the initial problem, the inference steps, and the target answer corresponding to the initial problem; compared with the simple question-and-answer pair including the initial problem and the initial answer, the target data increases the target inference steps, thus realizing the extension of the long inference link for the simple question-and-answer pair, and providing effective training data for improving the model inference ability subsequently.

[0064] See Figure 2 , Figure 2 FIG. shows a flowchart of a data processing method provided by an embodiment of this specification, which specifically includes the following steps.

[0065] Step 202: Determine at least one problem inference step corresponding to the initial problem and the inference answer corresponding to the initial problem obtained based on the at least one problem inference step;

[0066] Wherein, the initial problem can be understood as the original problem proposed by the user. For example, the original problem can be a math problem, a logic problem, or a code problem, etc., which is not limited here. The problem inference step can be understood as a series of logical deduction or calculation steps required to solve the initial problem; each problem inference step is a key node in the inference process. For example, in a math problem, the problem inference step can be "first calculate A", "then calculate B", "finally get C". The inference answer can be understood as the final result obtained through the problem inference step. The inference answer is the solution to the initial problem. For example, the inference answer can be the final answer to the math problem, the conclusion of the logic problem, etc., which is not limited here.

[0067] For example, the initial problem is a math problem: "Xiaoli has 4 apples, eats 2, and buys 1 more. How many apples does she have now?"; at this time, the at least one problem inference step corresponding to the initial problem includes "Xiaoming originally had 4 apples", "ate 2, and there are 4 - 2 = 2 left", "bought 1 more, and now there are 2 + 1 = 3"; through the above problem inference steps, the inference answer corresponding to the initial problem obtained is 3 apples.

[0068] In fact, at least one question reasoning step corresponding to the initial question can be determined through manual annotation or a question-answering model, and an inference answer can be derived based on the at least one question reasoning step. To save labor costs and improve data processing efficiency, in one or more embodiments of this specification, by using a question-answering model and example data, detailed reasoning can be performed on the initial question or the initial question and the initial answer to obtain the question reasoning steps corresponding to the initial question, and then the inference answer can be derived based on the question reasoning steps. The specific implementation is as follows:

[0069] The determination of at least one question reasoning step corresponding to the initial question, and the inference answer corresponding to the initial question obtained based on the at least one question reasoning step, includes:

[0070] Determine the initial question and example data;

[0071] Input the initial question and the example data into the question-answering model, or input the initial question, the initial answer corresponding to the initial question, and the example data into the question-answering model, and use the question-answering model to obtain the at least one question reasoning step corresponding to the initial question and the inference answer corresponding to the initial question, where the inference answer is obtained based on the at least one question reasoning step.

[0072] Among them, the initial answer can be understood as a short answer corresponding to the initial question. This short answer can be the final answer directly given for the initial question (such as the option of a multiple-choice question or "true / false" of a judgment question). The initial answer may be incomplete or lack a detailed reasoning process.

[0073] Example data can be understood as reference data or examples related to the initial question. The example data contains a detailed reasoning process of the question. For example, the example data can include questions similar to the initial question and their solution reasoning processes. The example data can help the question-answering model better understand the question and generate question reasoning steps. The question-answering model can be understood as a large language model used to generate question reasoning steps and inference answers. The main task of the question-answering model is to generate detailed reasoning steps and answers based on the input questions and example data.

[0074] In practical applications, in addition to inputting the initial question and example data into the question-answering model to obtain question reasoning steps, the initial answer can also be input into the question-answering model. This initial answer can be used as the reference answer of the question-answering model to help the question-answering model better understand the goal of the question and generate corresponding question reasoning steps.

[0075] Specifically, the question-and-answer model will generate at least one question reasoning step based on the input initial question, example data (and possibly the initial answer). The question reasoning step is the key logical derivation process for solving the initial question. After generating at least one question reasoning step, the question-and-answer model will obtain the final reasoning answer based on these steps.

[0076] The data processing method provided by the embodiments of this specification generates detailed question reasoning steps through the question-and-answer model. The question-and-answer model can not only obtain the final reasoning answer but also display the complete reasoning process. This helps the question-and-answer model to be more logically clear and accurate when solving complex problems. By clarifying the question reasoning steps, the question-and-answer model needs to deduce step by step according to logic, thereby reducing the possibility of guessing, improving the accuracy of the reasoning answer, and providing data support for subsequent adjustments to other models.

[0077] In one or more embodiments of this specification, the question-and-answer model will gradually generate the question reasoning step corresponding to the current reasoning time step in an autoregressive manner, so as to obtain at least one question reasoning step corresponding to the initial question according to the question reasoning steps of each current reasoning time step from the initial question to the reasoning answer. The specific implementation is as follows:

[0078] Obtaining the at least one question reasoning step corresponding to the initial question by using the question-and-answer model includes:

[0079] In the question-and-answer model, obtain the question reasoning step corresponding to the current reasoning time step in turn in an autoregressive manner until the current reasoning time step is the end reasoning time step;

[0080] Obtain the at least one question reasoning step according to the question reasoning steps corresponding to the current reasoning time step obtained in turn.

[0081] Among them, autoregression can be understood as a way of generating text. When the model generates text, at the current time step, only one word or one segment is generated, and then at the next time step, the next word or segment is generated based on the content that has been generated. In the embodiments of this specification, at the current time step, a reasoning step is obtained, and then at the next time step, the next reasoning step is generated based on the reasoning step that has been generated.

[0082] The reasoning time step can be understood as each step or stage when the question-and-answer model generates the reasoning step. At each reasoning time step, a new reasoning step is generated based on the reasoning steps that have been generated; the current reasoning time step is the current reasoning time step being processed.

[0083] Specifically, in the Q&A model, the Q&A model generates question reasoning steps step by step (i.e., at sequentially determined current inference time steps) in an autoregressive manner. Each generated question reasoning step depends on the output of the previous inference time step. The generated question reasoning steps at each step are combined until a complete reasoning process (including at least one question reasoning step) is generated.

[0084] In the data processing method provided by the embodiments of this specification, through the autoregressive manner, the Q&A model can generate question reasoning steps step by step, ensuring that each step is based on the previously generated logical deduction. This way makes the reasoning process clearer and more coherent. And since each step depends on the output of the previous inference time step, the Q&A model can continuously adjust and correct during the generation process, reducing the possibility of error accumulation; by generating question reasoning steps step by step, it is easier to optimize and expand each question reasoning step. For example, when an error occurs in a certain reasoning step, backtracking and correction can be performed.

[0085] In one or more embodiments of this specification, in at least one initial reasoning step corresponding to the current inference time step, determine a question reasoning step, so that in the next inference time step marked as the current inference time step when reasoning, through the autoregressive manner, based on the initial question and the obtained question reasoning steps, obtain the corresponding at least one initial reasoning step. By repeatedly executing the above steps, sequentially obtain the question reasoning steps corresponding to the current inference time step. The specific implementation is as follows:

[0086] The obtaining of the question reasoning steps corresponding to the current inference time step in sequence through the autoregressive manner until the current inference time step is the end inference time step includes:

[0087] Through the autoregressive manner, obtain at least one initial reasoning step corresponding to the current inference time step, and determine the question reasoning step from the at least one initial reasoning step;

[0088] In the case of determining that there is a next inference time step for the current inference time step, mark the next inference time step as the current inference time step,

[0089] Continue to execute through the autoregressive manner, obtain at least one initial reasoning step corresponding to the current inference time step, and determine the question reasoning step from the at least one initial reasoning step until the current inference time step is the end inference time step.

[0090] Specifically, at the current inference time step, when the Q&A model generates inference steps, it usually generates multiple possible candidate inference steps (i.e., initial inference steps). When generating multiple candidate inference steps, each candidate inference step has a corresponding probability value, indicating the likelihood that the Q&A model believes this candidate inference step is correct. Therefore, the Q&A model can select the candidate inference step with a higher probability value as the problem inference step for the current inference time step based on the probability values corresponding to the candidate inference steps.

[0091] For example, continuing with the above example, at the current inference time step, the initial inference steps sampled by the Q&A model include: the first initial inference step is "Xiaoli initially had 4 apples", the second initial inference step is "Xiaoli has 4 apples", and the third initial inference step is "Xiaoli originally had 4 apples". The Q&A model will calculate the probability of each initial inference step. For example, the probability of the first initial inference step is 0.5, the probability of the second initial inference step is 0.2, and the probability of the third initial inference step is 0.3. Finally, it selects the first initial inference step with the highest probability as the problem inference step for the current inference time step.

[0092] If the complete inference process has not been generated at the current inference time step, that is, when there is a next inference time step, the Q&A model needs to continue generating subsequent inference steps. Specifically, it marks the next inference time step as the current inference time step and continues to generate at least one initial inference step corresponding to the current inference time step. In fact, in the new current inference time step, the Q&A model will continue to generate at least one initial inference step in an autoregressive manner and determine the problem inference step for the current inference time step from the at least one initial inference step. This process will repeat until the current inference time step is the end inference time step and the inference process ends.

[0093] In specific implementation, the end inference time step can be the inference time step for generating the inference answer or the inference time step when the set maximum number of inference time steps is reached. In practical applications, when the Q&A model gradually conducts problem inference, the inference steps corresponding to the end inference time step usually contain an end symbol. Therefore, when the problem inference step generated at the current inference time step contains an end symbol, this current inference time step is the end inference time step. Additionally, setting the maximum number of inference time steps can prevent the Q&A model from falling into an infinite loop when generating inference steps, ensure that the inference process can end within a reasonable time, and avoid unnecessary resource waste. Therefore, when the current inference time step reaches the maximum number of inference time steps, this current inference time step is determined as the end inference time step.

[0094] The data processing method provided by the embodiments of this specification can obtain, from the initial question to the end inference time step, at least one question inference step corresponding to the initial question, which is composed of the question inference steps corresponding to each current inference time step, by selecting a question inference step in at least one initial inference step corresponding to each current inference time step.

[0095] In one or more embodiments of this specification, when expanding the initial question and the initial answer to obtain feature data for enhancing the inference ability of the model, in order to obtain as much feature data as possible, diverse question inference steps need to be generated according to an initial question. The specific implementation is as follows:

[0096] The method of obtaining the question inference step corresponding to each current inference time step in sequence through an autoregressive manner until the current inference time step is the end inference time step includes:

[0097] Through the autoregressive manner, obtain multiple initial inference steps corresponding to the current inference time step, and respectively determine the multiple initial inference steps as the question inference steps of multiple inference paths;

[0098] For each inference path, when there is a next inference time step in the current inference time step, mark the next inference time step as the current inference time step,

[0099] Continue to execute the operation of obtaining multiple initial inference steps corresponding to the current inference time step through the autoregressive manner, and respectively determine the multiple initial inference steps as the question inference steps of multiple inference paths until the current inference time step is the end inference time step.

[0100] Among them, when the question - answering model generates question inference steps, it may explore different inference paths, and each inference path corresponds to a possible inference process.

[0101] For example, in the first current inference time step, the initial inference steps for generating the initial question include a first initial inference step, a second initial inference step, and a third initial inference step. These three initial inference steps are respectively used as the question inference steps corresponding to the first inference path, the second inference path, and the third inference path in the first current inference time step.

[0102] Taking the first reasoning path as an example, when it is determined that there is a next reasoning time step for the first current reasoning time step, that is, in the case of the second current reasoning time step, in the second current reasoning time step, using the initial question and the first initial reasoning step corresponding to the first reasoning path, through an autoregressive manner, the fourth initial reasoning step and the fifth initial reasoning step corresponding to the second current reasoning time step are obtained. At this time, it is equivalent to the first reasoning path branching at the second current reasoning time step, and the first reasoning path is extended to the fourth reasoning path and the fifth reasoning path.

[0103] That is, the question reasoning steps in the fourth reasoning path include the first initial reasoning step corresponding to the first current reasoning time step and the fourth initial reasoning step corresponding to the second current reasoning time step. The fifth reasoning path includes the first initial reasoning step corresponding to the first current reasoning time step and the fifth initial reasoning step corresponding to the second current reasoning time step. And so on. Subsequently, when at least one initial reasoning step is generated at the current reasoning time step, the reasoning path corresponding to the previous current reasoning time step can be extended to form a new reasoning path until multiple reasoning paths can be obtained when the current reasoning time step is the end reasoning time step.

[0104] The data processing method provided by the embodiments of this specification can generate at least one initial reasoning step and determine each initial reasoning step as a question reasoning step. The question-answering model can explore and obtain multiple reasoning paths, increasing the diversity of the reasoning process.

[0105] In one or more embodiments of this specification, by taking each reasoning path as the target reasoning path, the question reasoning steps corresponding to each current reasoning time step in each reasoning path can be determined as the at least one question reasoning step. That is, for an initial question, multiple different at least one question reasoning steps corresponding to multiple reasoning paths can be obtained, thereby increasing the diversity of the reasoning data. The specific implementation is as follows:

[0106] Obtaining the at least one question reasoning step according to the question reasoning steps corresponding to the current reasoning time step obtained in sequence includes:

[0107] Sequentially determining the multiple reasoning paths as the target reasoning paths;

[0108] Obtaining the at least one question reasoning step according to the question reasoning steps corresponding to the current reasoning time step obtained in sequence in the target reasoning path.

[0109] Specifically, the inference path includes the problem inference steps of the entire process from the first step to the last step. That is, for each inference path, the problem inference steps corresponding to each current inference time step included therein can be determined. When each inference path will be used as the target inference path and the at least one problem inference step is the problem inference step included in the target inference path, it is possible to use the problem inference steps included in each inference path as the at least one problem inference step, so as to obtain at least one problem inference step corresponding to different inference paths for an initial problem.

[0110] In the above embodiment, by determining one inference step as the problem inference step among the multiple initial inference steps corresponding to each current inference time step, finally, one inference path corresponding to the initial problem will be obtained, that is, this inference path is composed of one problem inference step corresponding to each current inference time step. Therefore, the way to obtain at least one problem inference step in the above embodiment can be called chain construction, that is, based on one problem inference step corresponding to each current inference time step, at least one problem inference step of at least one inference time step is obtained in a chain.

[0111] In this embodiment, by determining all the multiple initial inference steps corresponding to each current inference time step as the problem inference steps, as the inference process progresses, multiple initial inference steps can be horizontally expanded at each current inference time step. Based on each current inference time step and the multiple problem inference steps corresponding to each current inference time step, multiple inference paths can be formed. Based on each inference path among the multiple inference paths, at least one problem inference step corresponding to each inference path can be obtained. Therefore, the way to obtain at least one problem inference step in this embodiment can be called tree construction.

[0112] The data processing method provided in the embodiments of this specification explores multiple inference paths. And when each inference path is used as the target inference path in turn, according to the problem inference steps corresponding to each current inference time step included in the target inference path, it is possible to obtain at least one different problem inference step corresponding to multiple inference paths for an initial problem, thereby increasing the diversity of inference data.

[0113] Step 204: According to the initial answer corresponding to the initial problem, perform a correctness check on the at least one problem inference step and the inference answer to obtain a check result.

[0114] Among them, the correctness check can be understood as checking and verifying whether the problem inference step and the inference answer are logical and whether they can correctly solve the initial problem. The check result can be understood as the final result of the check process, and the check result includes two results: correct and incorrect.

[0115] Specifically, in the embodiments of this specification, there is a judgment module for verifying the correctness of at least one question reasoning step and the reasoning answer output by the Q&A model, and determining whether the question reasoning step is a correct reasoning step and whether the reasoning answer is the correct answer to the initial question.

[0116] By verifying the correctness of at least one question reasoning step and the reasoning answer with the initial answer to the initial question, it can be determined whether subsequent corrections need to be made to at least one question reasoning step and the reasoning answer.

[0117] In one or more embodiments of this specification, a first verification result, that is, a correct result, can be obtained only when at least one question reasoning step and the reasoning answer are both correct; when there is an error in at least one question reasoning step or the reasoning answer, a second verification result, that is, an incorrect result, is obtained. The specific implementation is as follows:

[0118] Verifying the correctness of at least one question reasoning step and the reasoning answer according to the initial answer corresponding to the initial question to obtain a verification result includes:

[0119] Verifying the correctness of at least one question reasoning step and the reasoning answer according to the initial answer;

[0120] When it is determined that at least one question reasoning step and the reasoning answer are both correct, a first verification result is obtained;

[0121] When it is determined that at least one question reasoning step and / or the reasoning answer is incorrect, a second verification result is obtained.

[0122] Among them, the first verification result can be understood as a correct result, indicating that at least one question reasoning step and the reasoning answer are both correct; the second verification result can be understood as an incorrect result, indicating that there is an error in at least one question reasoning step and / or the reasoning answer.

[0123] It should be noted that when there are multiple question reasoning steps in at least one question reasoning step, a second verification result will be obtained when there is an incorrect question reasoning step among the multiple question reasoning steps.

[0124] Specifically, the judgment module performs correctness verification on the question reasoning steps and reasoning answers generated by the Q&A model based on the initial answer. For example, it checks the consistency between the reasoning answer and the initial answer. When the reasoning answer is consistent with the initial answer, it determines the reasoning answer as the correct answer to the initial question. It also performs logical consistency verification on the question reasoning steps to determine whether each step of derivation conforms to common sense and domain knowledge. When each step of derivation conforms to common sense and domain knowledge and the derived reasoning answer is consistent with the initial answer, it determines that both the question reasoning steps and the reasoning answer are correct.

[0125] If both the question reasoning steps and the reasoning answer are correct, a first verification result is obtained; if there is an error in the question reasoning steps and / or the reasoning answer, a second verification result is obtained.

[0126] The data processing method provided by the embodiments of this specification can ensure the accuracy of the question reasoning steps and reasoning answers through correctness verification. Based on the correct question reasoning steps and reasoning answers, it can provide data support for subsequent model adjustment, and for the incorrect question reasoning steps or reasoning answers, they can also be corrected in time.

[0127] Step 206: Obtain target data according to the verification result, the initial question, and the at least one question reasoning step, where the target data includes the initial question, target reasoning steps, and the target answer corresponding to the initial question.

[0128] Among them, the target reasoning steps can be understood as the correct reasoning steps corresponding to the initial question. When at least one of the question reasoning steps is correct, the target reasoning steps are the at least one question reasoning step. When there are incorrect reasoning steps among the at least one question reasoning step, the at least one question reasoning step can be corrected to obtain the corrected target reasoning steps.

[0129] The target data can be understood as the training data for subsequent model training. The feature data in this training data is the initial question, and the label data includes the target reasoning steps and the target answer corresponding to the initial question.

[0130] Specifically, when the verification result is the first verification result, both the question reasoning steps and the reasoning answer are correct. At this time, the at least one question reasoning step can be directly determined as the target reasoning steps, and the reasoning answer can be determined as the target answer. When the verification result is the second verification result, there is an error in the question reasoning steps and / or the reasoning answer. At this time, it is necessary to correct the question reasoning steps or the reasoning answer to obtain the target reasoning steps and the target answer.

[0131] In practical applications, further integration or correction processing based on the verification results can ensure the accuracy of the target data. Therefore, when training the model based on the correct and high-quality target data, the inference ability of the model can be improved.

[0132] In one or more embodiments of this specification, by determining the initial question as feature data, and determining at least one question inference step and the target answer as label data, target data for training the long inference ability of the model can be obtained, realizing the expansion of ordinary initial questions and initial answers to obtain training data with a long inference chain. The specific implementation is as follows:

[0133] Obtaining the target data according to the verification result, the initial question, and the at least one question inference step includes:

[0134] Determine the initial question as feature data according to the first verification result;

[0135] Determine the at least one question inference step as the target inference step, determine the inference answer as the target answer, and determine the target inference step and the target answer as label data;

[0136] Obtain the target data according to the feature data and the label data.

[0137] Among them, in machine learning, feature data refers to the input data of the model. In the embodiments of this specification, the initial question is used as feature data, representing the input of the model; the target answer can be understood as the finally determined correct answer after verification, and label data is used to represent the output target of the model. In the embodiments of this specification, at least one question inference step and the inference answer are used as label data, representing the output target of the model.

[0138] Specifically, when the verification result is the "first verification result", the initial question is used as feature data, the inference answer is used as the target answer, and at least one question inference step is determined as the target inference step. Thus, the target inference step and the target answer are used as label data, and finally the target data is generated. This target data can be understood as the shortcut sample in the above embodiments, and is directly obtained through the initial question, the correct at least one question inference step output by the question-and-answer model, and the inference answer.

[0139] The data processing method provided by the embodiments of this specification can obtain tag data by reasoning steps and reasoning answers corresponding to at least one problem according to the first verification result, and can expand the original question-and-answer pair composed of the initial question and the initial answer into target data including at least one problem reasoning step such as reasoning data, thereby realizing the expansion of the original data and increasing the training data for training the model's reasoning ability.

[0140] In one or more embodiments of this specification, when the verification result is the second verification result, that is, when at least one problem reasoning step or reasoning answer is incorrect, a correction model can be used to correct the incorrect problem reasoning step or reasoning answer to obtain correct target data. The specific implementation is as follows:

[0141] Obtaining the target data according to the verification result, the initial question, and the at least one problem reasoning step includes:

[0142] According to the second verification result, determine the reference reasoning steps from the at least one problem reasoning step, where the reference reasoning steps are the reasoning steps from the first problem reasoning step to the first incorrect reasoning step among the at least one problem reasoning step;

[0143] Input the reference reasoning steps, the initial question, and the initial answer into the correction model to obtain at least one corrected reasoning step and a corrected answer;

[0144] When the verification result of the at least one corrected reasoning step and the corrected answer is the first verification result, determine the initial question as the feature data, determine the reference reasoning steps and the at least one corrected reasoning step as the target reasoning steps, determine the corrected answer as the target answer, and determine the target reasoning steps, the at least one corrected reasoning step, and the target answer as the tag data;

[0145] Obtain the target data according to the feature data and the tag data.

[0146] Among them, the reference reasoning steps can be understood as the reasoning steps from the first problem reasoning step to the first incorrect reasoning step among the at least one problem reasoning step; for example, for a math problem, the at least one problem reasoning step includes the first problem reasoning step, the second problem reasoning step, and the third problem reasoning step. When the second problem reasoning step in the middle is incorrect, resulting in an incorrect reasoning answer, the reference reasoning steps include the first problem reasoning step and the second problem reasoning step.

[0147] The correction model can be understood as a large model with a larger size and stronger reasoning ability compared to the above-mentioned question-and-answer model. Thus, the correction model can correct the incorrect question reasoning steps output by the question-and-answer model. The corrected reasoning steps can be understood as the revised reasoning steps generated by the correction model, which are corrections and / or supplements to the reference reasoning steps. The corrected answer is the finally determined correct answer after correction.

[0148] Specifically, by inputting the reference reasoning steps, the initial question, and the correct initial answer corresponding to the initial question into the correction model, the correction model can reflect on and correct the previously obtained reference reasoning steps to obtain at least one corrected reasoning step and a corrected answer. In fact, when the verification result of at least one corrected reasoning step and the corrected answer is the first verification result, that is, when the verification result is correct, although tortuous, the finally correct corrected reasoning steps and the target answer can be obtained.

[0149] The data processing method provided by the embodiments of this specification, by using the reference reasoning steps, the corrected reasoning steps, and the corrected answer as labeled data, can enhance the training of the model when training the model subsequently, so that on the basis of the model having long reasoning ability, the model's ability to reflect on and correct errors is enhanced.

[0150] In one or more embodiments of this specification, when obtaining the target data, the target data can be used to train the model to be adjusted to obtain a target model with stronger reasoning ability. By expanding the initial question and the initial answer into the target data of a long reasoning link, the training data of the reasoning model is expanded. The specific implementation is as follows:

[0151] After obtaining the target data according to the verification result, the initial question, and the at least one question reasoning step, it further includes:

[0152] Determine the model to be adjusted;

[0153] Use the target data to train the model to be adjusted to obtain a target model.

[0154] Among them, the model to be adjusted can be understood as a model that needs to be adjusted or trained, which can be an existing large language model or a model specifically used for reasoning tasks, and is not limited here. The target model can be understood as a model obtained after training. This target model has stronger reasoning ability and can generate more accurate reasoning steps and answers corresponding to questions.

[0155] Specifically, the system can select an appropriate model as the model to be adjusted according to the task requirements, and use the target data to train the model to be adjusted, so that the model to be adjusted can better understand and solve the problems that require reasoning; specifically, the training process usually includes inputting feature data (initial questions) and label data (target reasoning steps and target answers), and enabling the model to be adjusted to learn how to derive the correct reasoning steps and answers from the questions.

[0156] The data processing method provided by the embodiments of this specification trains the model to be adjusted by using the target data, so that the obtained target model can better understand and solve the problems that require complex reasoning, thereby improving its reasoning ability. When the target data contains diverse target reasoning steps and target answers, the trained target model can better generalize to new and unseen reasoning problems, and can enhance a certain model reflection ability.

[0157] In one or more embodiments of this specification, the loss function of the prediction result and the label data is calculated, and then the model parameters of the model to be adjusted are adjusted using an optimization algorithm, so as to obtain the target model. The specific implementation is as follows:

[0158] Training the model to be adjusted using the target data to obtain a target model includes:

[0159] Inputting the feature data into the model to be adjusted to obtain the prediction result output by the model to be adjusted;

[0160] Adjusting the parameters of the model to be adjusted according to the prediction result and the label data to obtain the target model.

[0161] Among them, the prediction result can be understood as the output result generated by the model to be adjusted according to the feature data, including medical reasoning steps and predicted answers.

[0162] Specifically, by comparing the prediction result and the label data, calculating the loss function, and then using an optimization algorithm (such as gradient descent) to adjust the model parameters of the model to be adjusted, so that the target model obtained after adjusting the model parameters and the corresponding output result can be closer to the label data.

[0163] The data processing method provided by the embodiments of this specification, during the training process, inputs the feature data into the model to be adjusted to generate a prediction result, and then adjusts the parameters of the model according to the prediction result and the label data, and finally obtains a target model with stronger reasoning ability.

[0164] The data processing method provided in the embodiments of this specification can, for the existing Q&A data with simple answers, automatically construct target data with long inference chains through a Q&A model. By combining model fine-tuning and using the target data, the inference ability of the model to be adjusted can be effectively improved. When constructing inference path data, there are two construction methods, namely chain construction and tree construction, so that more diverse inference path data (the inference path data includes at least one question inference step) can be obtained. And for the inference path data with errors, reflection and correction can be carried out, so that the target data obtained after correction can be used to perform model enhancement training on the model to be adjusted, which is beneficial to the enhancement of the model's reflection and inference abilities.

[0165] In one embodiment of this specification, a model training method is further provided, including:

[0166] Determine the model to be adjusted;

[0167] Use the target data to train the model to be adjusted to obtain a target model, where the target data is obtained through the above data processing method.

[0168] Specifically, when the target data is obtained through the above data processing method, the target data includes feature data and label data. The feature data is input into the model to be adjusted, and the prediction result output by the model to be adjusted is obtained. According to the prediction result and the label data, the parameters of the model to be adjusted are adjusted to obtain a target model.

[0169] When the label data of the target data includes the target inference step, the trained target model can have a stronger inference ability.

[0170] See Figure 3 , Figure 3 Schematic diagram of the chain data construction process of a data processing method provided in one embodiment of this specification.

[0171] In practical applications, in the process of chain data construction, there are two links. The first link is to input the initial question into the Q&A model (this Q&A model can be a large language model), and in the Q&A model, use the initial question and example data (few-shots) to generate the inference process at one time (that is, at least one question inference step in the above embodiment, including the question inference step corresponding to each inference time step. In the embodiments of this specification, the entire process of analyzing the initial question, the first question inference step to the generation of the final inference answer is called a path, that is, the inference path) and the final inference answer. The other link inputs the initial question and the initial answer, and allows the Q&A model to generate the inference process and the inference answer at one time. This link is equivalent to giving the reference answer to generate the inference process and the inference answer.

[0172] Since there is a possibility that the question-answering model may output incorrect or fabricated question reasoning steps during the generation process, a judgment module is introduced. This judgment module can be a larger-sized large language model or expert / manual annotation, so as to collect incorrect and correct reasoning paths using the initial answers.

[0173] For the correct reasoning path, the initial question, reasoning process, and reasoning answer corresponding to this reasoning path are used as shortcut samples (i.e., the first target data) and directly saved. For the incorrect reasoning path, the reasoning process is split, and the determined reference reasoning steps (from the first question reasoning step to the question reasoning step where the error occurs), the initial question, and the initial answer are input into the error correction model (this error correction model can also be a large model). Let the error correction model perform error correction and then generate subsequent steps (i.e., the error correction reasoning steps in the above embodiment), and obtain the correct target answer. Finally, the correct reasoning path is retained, and based on the initial question, target reasoning steps, and target answer, backtracking samples (i.e., the second target data) are obtained.

[0174] See Figure 4 , Figure 4 which shows a schematic diagram of the tree data construction process of a data processing method provided by an embodiment of this specification.

[0175] Figure 4 In this case, when the question-answering model reasons about the initial question, at least one initial reasoning step can be obtained at each reasoning time step, and it continues to expand until the final reasoning answer is generated or the maximum set tree depth (i.e., the maximum number of reasoning time steps in the above embodiment) is reached.

[0176] Similarly, the judgment module is used to perform correctness verification on each generated reasoning answer. When the verification result is correct, the initial question, reasoning process, and reasoning answer on this reasoning path are saved as shortcut samples; when the verification result is incorrect, the initial answer and a larger-sized and more powerful large language model (i.e., the error correction model in the above embodiment) are introduced to further reflect and correct errors, and finally the correct target answer is generated. And based on the same method as above, backtracking samples are obtained.

[0177] The shortcut samples and backtracking samples obtained based on the above chain generation or tree construction can be used to adjust and train the basic model to be adjusted (i.e., the model to be adjusted in the above embodiments), and the method of adjustment and training can be full-parameter fine-tuning or LoRA (Low-Rank Adaptation) or other peft (Parameter-Efficient Fine-Tuning, a parameter-efficient fine-tuning) training methods, which are not limited herein.

[0178] The data processing method provided in the embodiments of this specification can, for the existing Q&A data with simple answers, automatically construct target data with long inference chains through a Q&A model. Combining model fine-tuning can effectively improve the inference ability of existing models. Specifically, through two ways of constructing inference path data, namely chain and tree, shortcut long inference data and even longer backtracking inference data can be constructed. At the same time, combined with the horizontal expansion of the tree, it is possible to generate more diverse inference path data; and combined with a larger-scale large model to implement a correction function, that is, to reflect on and correct the problem inference steps, and obtain longer inference data that is tortuous but ultimately still correct. Then, when using these longer inference data for model enhancement training, it is beneficial to enhance the model's reflection and inference ability.

[0179] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing device. Figure 5 Shows a schematic structural diagram of a data processing device provided in an embodiment of this specification. As Figure 5 shown, the device includes:

[0180] An inference module 502, configured to determine at least one problem inference step corresponding to an initial problem, and an inference answer corresponding to the initial problem obtained based on the at least one problem inference step;

[0181] A verification module 504, configured to perform a correctness verification on the at least one problem inference step and the inference answer according to the initial answer corresponding to the initial problem, and obtain a verification result;

[0182] An obtaining module 506, configured to obtain target data according to the verification result, the initial problem, and the at least one problem inference step, where the target data includes the initial problem, target inference steps, and a target answer corresponding to the initial problem.

[0183] Optionally, the inference module 502 is configured to:

[0184] Determine an initial problem and example data;

[0185] Input the initial problem and the example data into the Q&A model, or input the initial problem, the initial answer corresponding to the initial problem, and the example data into the Q&A model, and use the Q&A model to obtain the at least one question reasoning step corresponding to the initial problem and the reasoning answer corresponding to the initial problem, where the reasoning answer is obtained based on the at least one question reasoning step.

[0186] Optionally, the verification module 504 is configured to:

[0187] Verify the correctness of the at least one question reasoning step and the reasoning answer according to the initial answer;

[0188] Obtain a first verification result when it is determined that the at least one question reasoning step and the reasoning answer are both correct;

[0189] Obtain a second verification result when it is determined that the at least one question reasoning step and / or the reasoning answer is incorrect.

[0190] Optionally, the obtaining module 506 is configured to:

[0191] Determine the initial problem as the feature data according to the first verification result;

[0192] Determine the at least one question reasoning step as the target reasoning step, determine the reasoning answer as the target answer, and determine the target reasoning step and the target answer as the label data;

[0193] Obtain the target data according to the feature data and the label data.

[0194] Optionally, the obtaining module 506 is configured to:

[0195] Determine a reference reasoning step from the at least one question reasoning step according to the second verification result, where the reference reasoning step is the reasoning step from the first question reasoning step to the first occurrence of an incorrect reasoning step among the at least one question reasoning steps;

[0196] Input the reference reasoning step, the initial problem, and the initial answer into the correction model to obtain at least one corrected reasoning step and a corrected answer;

[0197] In the case where the check result of the at least one corrective inference step and the corrective answer is the first check result, determine the initial question as the feature data, determine the reference inference step and the at least one corrective inference step as the target inference step, determine the corrective answer as the target answer, and determine the target inference step, the at least one corrective inference step, and the target answer as the label data;

[0198] Obtain the target data according to the feature data and the label data.

[0199] Optionally, the inference module 502 is configured to:

[0200] In the question-and-answer model, obtain the question inference steps corresponding to each current inference time step in an autoregressive manner until the current inference time step is the end inference time step;

[0201] Obtain the at least one question inference step according to the question inference steps corresponding to the current inference time step obtained in sequence.

[0202] Optionally, the inference module 502 is configured to:

[0203] Obtain at least one initial inference step corresponding to the current inference time step in the autoregressive manner, and determine the question inference step from the at least one initial inference step;

[0204] In the case where it is determined that there is a next inference time step for the current inference time step, mark the next inference time step as the current inference time step,

[0205] Continue to execute the operation of obtaining at least one initial inference step corresponding to the current inference time step in the autoregressive manner, and determining the question inference step from the at least one initial inference step until the current inference time step is the end inference time step.

[0206] Optionally, the inference module 502 is configured to:

[0207] Obtain a plurality of initial inference steps corresponding to the current inference time step in the autoregressive manner, and respectively determine the plurality of initial inference steps as the question inference steps of a plurality of inference paths;

[0208] For each inference path, in the case where it is determined that there is a next inference time step for the current inference time step, mark the next inference time step as the current inference time step,

[0209] Continue to execute in the above autoregressive manner to obtain multiple initial inference steps corresponding to the current inference time step, and respectively determine the multiple initial inference steps as the problem inference steps of multiple inference paths until the current inference time step is the end inference time step.

[0210] Optionally, the inference module 502 is configured to:

[0211] Sequentially determine the multiple inference paths as the target inference paths;

[0212] According to the problem inference steps corresponding to the current inference time step sequentially obtained in the target inference path, obtain the at least one problem inference step.

[0213] The apparatus further includes:

[0214] A training module, configured to determine a model to be adjusted; use the target data to train the model to be adjusted to obtain a target model.

[0215] Optionally, the training module is configured to:

[0216] Input the feature data into the model to be adjusted to obtain a prediction result output by the model to be adjusted;

[0217] According to the prediction result and the label data, adjust the parameters of the model to be adjusted to obtain the target model.

[0218] The data processing apparatus provided in the embodiments of this specification can obtain inference path data including problem inference steps corresponding to the initial problem by determining at least one problem inference step corresponding to the initial problem and the inference answer corresponding to the initial problem obtained based on the at least one problem inference step, and when verifying the correctness of the at least one problem inference step and the inference answer according to the initial answer corresponding to the initial problem, correct target data can be obtained through the obtained verification result, the initial problem, and the at least one problem inference step, where the target data includes the initial problem, target inference steps, and the target answer corresponding to the initial problem; compared with the simple question-and-answer pair including the initial problem and the initial answer, the target data adds target inference steps, thereby realizing the extension of the long inference link for the simple question-and-answer pair, and providing effective training data for improving the model inference ability subsequently.

[0219] The above is a schematic solution of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.

[0220] Figure 6 FIG. 4 shows a structural block diagram of a computing device 600 according to an embodiment of this specification. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.

[0221] The computing device 600 further includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).

[0222] In an embodiment of this specification, the above components of the computing device 600 and Figure 6 other components not shown in FIG. 4 may also be connected to each other, for example, through a bus. It should be understood that Figure 6 the structural block diagram of the computing device shown in FIG. 4 is only for illustrative purposes and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0223] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.

[0224] Wherein, the processor 620 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above data processing method are implemented.

[0225] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the data processing method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the data processing method embodiment.

[0226] An embodiment of this specification also provides a computer-readable storage medium, which stores computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above data processing method are implemented.

[0227] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the data processing method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the data processing method embodiment.

[0228] An embodiment of this specification also provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above data processing method are implemented.

[0229] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above data processing method.

[0230] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0231] The computer instructions include computer program code, which can be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0232] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0233] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0234] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Determining, according to an initial question, example data, and a question-and-answer model, at least one question reasoning step corresponding to the initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step, wherein the example data contains a reasoning process corresponding to a question for enabling the question-and-answer model to generate question reasoning steps; Performing a correctness verification on the at least one question reasoning step and the reasoning answer according to an initial answer corresponding to the initial question to obtain a verification result; Obtaining target data according to the verification result, the initial question, and the at least one question reasoning step, wherein the target data includes the initial question, target reasoning steps, and a target answer corresponding to the initial question, The target reasoning steps are correct reasoning steps corresponding to the initial question. In the case where there are incorrect reasoning steps in the at least one question reasoning step, correcting the at least one question reasoning step until the corrected question reasoning steps are correct to obtain the target reasoning steps.

2. The data processing method according to claim 1, wherein the determining, according to an initial question, example data, and a question-and-answer model, at least one question reasoning step corresponding to the initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step, comprises: Determining the initial question and the example data; Inputting the initial question and the example data into the question-and-answer model, or inputting the initial question, an initial answer corresponding to the initial question, and the example data into the question-and-answer model, and using the question-and-answer model to obtain the at least one question reasoning step corresponding to the initial question and the reasoning answer corresponding to the initial question, wherein the reasoning answer is obtained based on the at least one question reasoning step.

3. The data processing method according to claim 1, wherein the performing a correctness verification on the at least one question reasoning step and the reasoning answer according to an initial answer corresponding to the initial question to obtain a verification result, comprises: Performing a correctness verification on the at least one question reasoning step and the reasoning answer according to the initial answer; Obtaining a first verification result when it is determined that both the at least one question reasoning step and the reasoning answer are correct; Obtaining a second verification result when it is determined that at least one of the at least one question reasoning step and the reasoning answer is incorrect.

4. The data processing method according to claim 3, wherein the obtaining target data according to the verification result, the initial question, and the at least one question reasoning step, comprises: Determining the initial question as feature data according to the first verification result; Determining the at least one question reasoning step as target reasoning steps, determining the reasoning answer as the target answer, and determining the target reasoning steps and the target answer as label data; Obtaining the target data according to the feature data and the label data.

5. The data processing method according to claim 3, wherein obtaining the target data according to the verification result, the initial question, and the at least one question reasoning step comprises: Determining a reference reasoning step from the at least one question reasoning step according to the second verification result, where the reference reasoning step is the reasoning steps between the first question reasoning step and the first error reasoning step among the at least one question reasoning step; Inputting the reference reasoning step, the initial question, and the initial answer into a correction model to obtain at least one corrected reasoning step and a corrected answer; In the case where the verification result of the at least one corrected reasoning step and the corrected answer is the first verification result, determining the initial question as the feature data, determining the reference reasoning step and the at least one corrected reasoning step as the target reasoning steps, determining the corrected answer as the target answer, and determining the target reasoning steps and the target answer as the label data; Obtaining the target data according to the feature data and the label data.

6. The data processing method according to claim 2, wherein obtaining the at least one question reasoning step corresponding to the initial question by using the question and answer model comprises: In the question and answer model, obtaining the question reasoning step corresponding to the current reasoning time step in sequence by an autoregressive manner until the current reasoning time step is the end reasoning time step; Obtaining the at least one question reasoning step according to the question reasoning steps corresponding to the current reasoning time step obtained in sequence.

7. The data processing method according to claim 6, wherein obtaining the question reasoning step corresponding to the current reasoning time step in sequence by an autoregressive manner until the current reasoning time step is the end reasoning time step comprises: Obtaining at least one initial reasoning step corresponding to the current reasoning time step by the autoregressive manner, and determining the question reasoning step from the at least one initial reasoning step; In the case where it is determined that there is a next reasoning time step for the current reasoning time step, marking the next reasoning time step as the current reasoning time step, Continuing to execute obtaining at least one initial reasoning step corresponding to the current reasoning time step by the autoregressive manner, and determining the question reasoning step from the at least one initial reasoning step until the current reasoning time step is the end reasoning time step.

8. The data processing method according to claim 6, wherein obtaining the question reasoning step corresponding to the current reasoning time step in sequence by an autoregressive manner until the current reasoning time step is the end reasoning time step comprises: Obtaining multiple initial reasoning steps corresponding to the current reasoning time step by the autoregressive manner, and respectively determining the multiple initial reasoning steps as the question reasoning steps of multiple reasoning paths; For each reasoning path, in the case where there is a next reasoning time step for the current reasoning time step, marking the next reasoning time step as the current reasoning time step, Continue to execute in the above autoregressive manner to obtain multiple initial inference steps corresponding to the current inference time step, and respectively determine the multiple initial inference steps as the problem inference steps of multiple inference paths until the current inference time step is the end inference time step.

9. The data processing method according to claim 8, wherein obtaining the at least one problem inference step according to the problem inference steps corresponding to the current inference time step obtained in sequence comprises: Sequentially determine the multiple inference paths as target inference paths; Obtain the at least one problem inference step according to the problem inference steps corresponding to the current inference time step obtained in sequence in the target inference path.

10. The data processing method according to claim 4 or 5, after obtaining the target data according to the verification result, the initial problem, and the at least one problem inference step, further comprises: Determine the model to be adjusted; Train the model to be adjusted with the target data to obtain a target model.

11. The data processing method according to claim 10, wherein training the model to be adjusted with the target data to obtain a target model comprises: Input the feature data into the model to be adjusted to obtain a prediction result output by the model to be adjusted; Adjust the parameters of the model to be adjusted according to the prediction result and the label data to obtain the target model.

12. A model training method, comprising: Determine the model to be adjusted; Train the model to be adjusted with target data to obtain a target model, wherein the target data is obtained by the data processing method according to any one of claims 1-11.

13. A data processing device, comprising: An inference module configured to determine at least one problem inference step corresponding to the initial problem and an inference answer corresponding to the initial problem obtained based on the at least one problem inference step according to the initial problem, example data, and a question-and-answer model, wherein the example data contains an inference process corresponding to the question for enabling the question-and-answer model to generate problem inference steps; A verification module configured to perform a correctness verification on the at least one problem inference step and the inference answer according to the initial answer corresponding to the initial problem to obtain a verification result; An obtaining module configured to obtain target data according to the verification result, the initial problem, and the at least one problem inference step, wherein the target data includes the initial problem, target inference steps, and a target answer corresponding to the initial problem, The target inference steps are correct inference steps corresponding to the initial problem. In the case where there are incorrect inference steps among the at least one problem inference steps, correct the at least one problem inference step until the corrected problem inference steps are correct to obtain the target inference steps.

14. A computing device, comprising: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

16. A computer program product includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Large model illusion treatment method, device and equipment and storage medium

    CN117556920A

  • Method and device for evaluating model reasoning ability and storage medium

    CN119066381A

  • Question and answer task processing model training method and device, equipment and storage medium

    CN119493849A