Data processing method and device

By determining and verifying the inference steps and answers of the initial problem, target data containing the inference steps is generated, which solves the problem that large language models lack training data when dealing with problems that require long-link inference, and improves the model's inference ability.

CN119990333AActive Publication Date: 2025-05-13ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510461488.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing large language models lack effective training data when dealing with mathematical and logical problems that require strong dependence on logical reasoning and long-link reasoning, because these problems usually lack intermediate inference step processes.

Method used

By determining at least one problem reasoning step and inference answer corresponding to the initial question, a correctness check is performed, a verification result is obtained, and target data is generated based on this information, which includes the initial question, inference step and target answer. This target data is used to train the model and extends the long inference link of simple question-and-answer pairs.

Benefits of technology

A long inference link extension for simple question-and-answer data is realized, providing effective training data to improve the model's inference ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990333A_ABST
    Figure CN119990333A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, and the method comprises the steps: determining at least one question reasoning step corresponding to an initial question, and obtaining a reasoning answer corresponding to the initial question based on the at least one question reasoning step; according to an initial answer corresponding to the initial question, performing correctness verification on the at least one question reasoning step and the reasoning answer to obtain a verification result; target data are obtained according to the verification result, the initial question and the at least one question reasoning step, and the target data comprise the initial question, a target reasoning step and a target answer corresponding to the initial question; compared with a simple question-answer pair containing an initial question and an initial answer, the target data has the advantage that a target reasoning step is added, so that extension of a long reasoning link for the simple question-answer pair is realized, and effective training data is provided for improving the model reasoning capability subsequently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence, and in particular, to a data processing method. One or more embodiments of this specification also relate to a data processing device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art. Background Art

[0002] In recent years, technologies related to large language models (LLMs) have continued to develop, and LLM capabilities have continued to increase. However, there is still room for improvement in some problems that require strong reliance on logical reasoning and long-chain reasoning to obtain results, such as common mathematical reasoning and logical reasoning.

[0003] However, when fine-tuning the model to enhance the model's reasoning ability, there is often a lack of effective training data, because many mathematical logic questions are multiple-choice questions, true-or-false questions, or directly give the final answer without the key reasoning steps in the middle, and cannot be directly used as training data for the model's reasoning ability. Summary of the invention

[0004] In view of this, an embodiment of this specification provides a data processing method. One or more embodiments of this specification also relate to a model training method, a data processing device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a data processing method is provided, including: Determine at least one question reasoning step corresponding to an initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step; According to the initial answer corresponding to the initial question, the correctness of the at least one question reasoning step and the reasoning answer is verified to obtain a verification result; Target data is obtained according to the verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the reasoning step, and a target answer corresponding to the initial question.

[0006] According to a second aspect of an embodiment of this specification, there is provided a data processing device, including: A reasoning module, configured to determine at least one question reasoning step corresponding to an initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step; A verification module is configured to verify the correctness of the at least one question reasoning step and the reasoning answer according to the initial answer corresponding to the initial question, and obtain a verification result; An acquisition module is configured to obtain target data based on the verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the reasoning step, and a target answer corresponding to the initial question.

[0007] According to a third aspect of an embodiment of this specification, a model training method is provided, including: Determine the model to be adjusted; The model to be adjusted is trained using target data to obtain a target model, wherein the target data is obtained by the above-mentioned data processing method.

[0008] According to a fourth aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.

[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the computer program / instruction implements the steps of the above-mentioned data processing method when executed by a processor.

[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.

[0011] A data processing method provided by an embodiment of the present specification can obtain reasoning path data including the question reasoning steps corresponding to the initial question by determining at least one question reasoning step corresponding to the initial question and the reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step, and when the correctness of the at least one question reasoning step and the reasoning answer is verified according to the initial answer corresponding to the initial question, correct target data can be obtained through the obtained verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the reasoning steps, and the target answer corresponding to the initial question; compared with a simple question-answer pair including the initial question and the initial answer, the target data adds a target reasoning step, thereby realizing the expansion of a long reasoning chain for a simple question-answer pair, and subsequently providing effective training data for improving the model reasoning capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a scenario diagram of a data processing method provided by an embodiment of this specification; Figure 2 is a flow chart of a data processing method provided by an embodiment of this specification; Figure 3 It is a schematic diagram of a chain data construction process of a data processing method provided by an embodiment of this specification; Figure 4 It is a schematic diagram of a tree data construction process of a data processing method provided by an embodiment of this specification; Figure 5 is a structural schematic diagram of a data processing device provided by an embodiment of this specification; Figure 6 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0013] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0014] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0017] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM, Large Language Model), a multi-modal pre-training model, etc.

[0018] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image description (IC, Image Caption), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0019] First, the terms involved in one or more embodiments of this specification are explained.

[0020] few-shots: some sample data used to provide reference for the question-answering model so that the question-answering model can obtain the final answer based on the reasoning steps.

[0021] Shortcut sample: A shortcut sample refers to data that directly infers the correct answer in the embodiments of this specification.

[0022] Backtracking sample: A backtracking sample in the embodiments of this specification refers to data in which there are some errors in the reasoning process, but the correct answer is finally obtained after self-reflection and error correction by the correction model.

[0023] The traditional method of fine-tuning the model to enhance the model's reasoning ability involves a large amount of manual labeling work, which not only requires expensive manpower and time costs, but also requires further proofreading for the consistency of manual labeling. In addition, because many math and logic questions are multiple-choice questions, true-or-false questions, or directly give the final answer without the key reasoning steps in the middle, many existing logic question-answering data cannot be well utilized.

[0024] For example, the common question and answer data example is: Question: "There are 15 different circles. What is the maximum possible number of points at which the circles intersect?" Options: (A) 390 (B) 100 (C) 110 (D) 180 (E) 210. Answer: E.

[0025] Example of long inference data: Question: "There are 15 different circles. What is the maximum possible number of points at which the circles intersect?" Options: (A) 390 (B) 100 (C) 110 (D) 180 (E) 210.

[0026] Answer: Analysis: Each pair of different circles has at most two intersection points. For 15 different circles, the maximum possible number of intersection points can be calculated using combinatorial mathematics.

[0027] 1. Calculate the number of pairs of circles: 15×14 / 2=105 pairs.

[0028] 2. Each pair of circles has at most 2 intersection points, so the total number of intersection points is 105×2=210.

[0029] Therefore, the maximum possible number of intersection points is 210, which means the answer is (E)210.

[0030] Example of general question-answering data: Question: "What is the mouth of the body of water on which the Bartram Suspension Bridge is located?" Answer: The Delaware River.

[0031] Example of long inference data: Question: "What is the mouth of the body of water on which the Bartram Suspension Bridge is located?" Answer: The mouth of the body of water on which the Bartram Suspension Bridge is located, namely the mouth of Croom Creek, is the Delaware River in Edston, Pennsylvania.

[0032] Through the above examples, we can see that many existing data, such as math problems, coding problems, logic problems, etc., are simple answers or multiple-choice questions. The answers are too short, and the model may guess the answer and it is not conducive to understanding the complete reasoning process. For use scenarios with high requirements for reasoning ability, the reasoning process and the reasoning result are equally important. Therefore, the embodiment of this specification proposes a data processing method that automatically expands ordinary question and answer data into long reasoning link data and then fine-tunes the model to efficiently improve the long reasoning ability of the model.

[0033] In order to solve the above technical problems, a data processing method is provided in this specification. This specification also involves a model training method, a data processing device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0034] See also Figure 1 , Figure 1 A scenario diagram of a data processing method provided according to an embodiment of the present specification is shown.

[0035] Specifically, the data processing method is applied to a data processing system, which includes an end-side device 102 and a server 104. The end-side device 102 is used to send initial questions and initial answers to the server 104. In actual applications, the user can input the initial questions and initial answers in the end-side device 102 by text or voice. If voice is used, the end-side device 102 will also include a corresponding voice processing part, such as voice analysis, voice-to-text, voice synthesis and other modules, which are used to convert the questions input by the user by voice into text. This manual does not limit this.

[0036] A question-answering model is trained in the server 104, and when the server 104 receives an initial question sent by the terminal device 102, the question-answering model is used to determine at least one question reasoning step corresponding to the initial question, and an inference answer corresponding to the initial question obtained based on the at least one question reasoning step; according to the initial answer corresponding to the initial question, the correctness of the at least one question reasoning step and the reasoning answer are verified to obtain a verification result; according to the verification result, the initial question and the at least one question reasoning step, target data is obtained, wherein the target data includes the initial question, the target reasoning step, and the target answer corresponding to the initial question; the target data can be returned to the terminal device 102; or when the request of the terminal device 102 carries a model training request, the target data can be used to train the model to be adjusted to obtain the target model, thereby returning the calling interface or the target model corresponding to the target model to the terminal device 102, so that the terminal device 102 can call the target model through the calling interface, or deploy the target model on the terminal device 102.

[0037] The end-side device 102 may include a browser, an APP (Application), or a web application such as an H5 (Hyper Text Markup Language 5, version 5 of Hypertext Markup Language) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The end-side device may be based on a software development kit (SDK) of the corresponding service provided by the server, such as based on the real-time communication (RTC, Real Time Communication) SDK development and acquisition. The end-side device may be deployed in an electronic device and needs to rely on the device to run or some APPs in the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications may also be configured in the electronic device, such as human-computer dialogue applications, model training applications, data processing applications, web browser applications, shopping applications, search applications, instant messaging tools, mailbox clients, social platform software, etc.

[0038] Server 104 can be understood as a server that provides various services, including physical servers and cloud servers, such as a server that provides communication services to multiple clients, a server for background training that supports the model used on the client, and a server that processes the data sent by the client. It should be noted that server 104 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. Server 104 can also be a server for a distributed system, or a server combined with a blockchain. Server 104 can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0039] It is worth noting that the data processing method provided in the embodiments of this specification can be executed by the server 104. In other embodiments of this specification, the question-and-answer model can be deployed in the terminal device 102, so that the terminal device 102 can also have similar functions as the server 104, thereby executing the data processing method provided in the embodiments of this specification; in other embodiments, the data processing method provided in the embodiments of this specification can also be executed jointly by the terminal device 102 and the server 104.

[0040] The data processing method provided in the embodiments of this specification can obtain reasoning path data including the question reasoning steps corresponding to the initial question by determining at least one question reasoning step corresponding to the initial question and the reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step, and when the correctness of the at least one question reasoning step and the reasoning answer is verified according to the initial answer corresponding to the initial question, correct target data can be obtained through the obtained verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the reasoning steps, and the target answer corresponding to the initial question; compared with a simple question-answer pair including the initial question and the initial answer, the target data adds a target reasoning step, thereby realizing the expansion of a long reasoning chain for a simple question-answer pair, and subsequently providing effective training data for improving the model reasoning capability.

[0041] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0042] Step 202: Determine at least one question reasoning step corresponding to the initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step; Among them, the initial question can be understood as the original question raised by the user, such as the original question can be a math problem, a logic problem or a code problem, etc., which is not limited here. The problem reasoning steps can be understood as a series of logical deductions or calculation steps that need to be taken to solve the initial problem; each problem reasoning step is a key node in the reasoning process. For example, in a math problem, the problem reasoning steps can be "calculate A first", "calculate B then", and "finally get C". The reasoning answer can be understood as the final result obtained through the problem reasoning steps. The reasoning answer is the solution to the initial question. For example, the reasoning answer can be the final answer to a math problem, the conclusion of a logic problem, etc., which is not limited here.

[0043] For example, the initial question is a math problem: "Xiao Li has 4 apples. He ate 2 and bought 1. How many apples does he have now?"; at this time, at least one reasoning step corresponding to the initial question includes "Xiao Ming had 4 apples at the beginning", "He ate 2, leaving 4-2=2", "He bought 1, now he has 2+1=3"; through the above reasoning steps, the reasoning answer corresponding to the initial question is 3 apples.

[0044] In fact, at least one question reasoning step corresponding to the initial question can be determined through manual annotation or question-answering model, and the reasoning answer can be derived based on at least one question reasoning step. In order to save labor costs and improve data processing efficiency, in one or more embodiments of this specification, the question-answering model and example data can be used to perform detailed reasoning on the initial question or the initial question and the initial answer, obtain the question reasoning steps corresponding to the initial question, and then derive the reasoning answer based on the question reasoning steps. The specific implementation is as follows: The determining of at least one question reasoning step corresponding to the initial question, and the reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step, include: Identify initial questions and sample data; The initial question and the example data are input into a question-answering model, or the initial question, the initial answer corresponding to the initial question, and the example data are input into a question-answering model, and the question-answering model is used to obtain at least one question reasoning step corresponding to the initial question and an reasoning answer corresponding to the initial question, wherein the reasoning answer is obtained based on the at least one question reasoning step.

[0045] Among them, the initial answer can be understood as the short answer corresponding to the initial question. The short answer can be the final answer given directly to the initial question (such as the options of a multiple-choice question or "true / false" of a judgment question). The initial answer may be incomplete or lack a detailed reasoning process.

[0046] Sample data can be understood as reference data or examples related to the initial question. Sample data contains detailed reasoning process of the question. For example, sample data can contain questions similar to the initial question and their answer reasoning process. Sample data can help the question-answering model better understand the question and generate reasoning steps for the question. The question-answering model can be understood as a large language model used to generate reasoning steps and reasoning answers for questions. The main task of the question-answering model is to generate detailed reasoning steps and answers based on the input questions and sample data.

[0047] In practical applications, in addition to inputting the initial question and example data into the question-answering model to obtain the question reasoning steps, the initial answer can also be input into the question-answering model. The initial answer can be used as a reference answer for the question-answering model to help the question-answering model better understand the goal of the question and generate corresponding question reasoning steps.

[0048] Specifically, the question-answering model will generate at least one question reasoning step based on the input initial question, example data (and possible initial answers). The question reasoning step is the key logical deduction process for solving the initial question. After generating at least one question reasoning step, the question-answering model will derive the final reasoning answer based on these steps.

[0049] The data processing method provided in the embodiments of this specification generates detailed question reasoning steps through a question-answering model. The question-answering model can not only derive the final reasoning answer, but also display the complete reasoning process, which helps the question-answering model to be more logical, clear and accurate when solving complex problems. By clarifying the question reasoning steps, the question-answering model needs to deduce step by step according to logic, thereby reducing the possibility of guessing, improving the accuracy of the reasoning answer, and providing data support for subsequent adjustments to other models.

[0050] In one or more embodiments of this specification, the question-answering model will gradually generate the question reasoning steps corresponding to the current reasoning time step by autoregression, thereby obtaining at least one question reasoning step corresponding to the initial question based on the question reasoning steps of each current reasoning time step from the initial question to the reasoning answer. The specific implementation is as follows: The step of obtaining the at least one question reasoning step corresponding to the initial question by using the question-answering model includes: In the question-answering model, the question reasoning steps corresponding to the current reasoning time step are obtained in sequence by an autoregressive method until the current reasoning time step is the end reasoning time step; The at least one problem reasoning step is obtained according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence.

[0051] Among them, autoregression can be understood as a way of generating text. When the model generates text, it only generates one word or one fragment in the current time step, and then continues to generate the next word or fragment based on the already generated content in the next time step; in the embodiment of this specification, an inference step is obtained in the current time step, and then the next inference step is generated based on the already generated inference step in the next time step.

[0052] The reasoning time step can be understood as each step or stage when the question-answering model generates reasoning steps. Each reasoning time step generates a new reasoning step based on the already generated reasoning steps; the current reasoning time step is the currently being processed and reasoning time step.

[0053] Specifically, in the question-answering model, the question-answering model generates question reasoning steps step by step (i.e., at the current reasoning time step determined sequentially) through autoregression. Each question reasoning step generated depends on the output of the previous reasoning time step. The question reasoning steps generated at each step are combined until a complete reasoning process (including at least one question reasoning step) is generated.

[0054] The data processing method provided in the embodiments of this specification can gradually generate question reasoning steps by the question-answering model through autoregression, ensuring that each step is based on the logical deduction generated previously. This method makes the reasoning process clearer and more coherent, and because each step depends on the output of the previous reasoning time step, the question-answering model can be continuously adjusted and corrected during the generation process to reduce the possibility of error accumulation; by gradually generating question reasoning steps, it is easier to optimize and expand each question reasoning step, such as backtracking and correcting when an error occurs in a certain reasoning step.

[0055] In one or more embodiments of the present specification, in at least one initial reasoning step corresponding to the current reasoning time step, a problem reasoning step is determined, so that when the next reasoning time step is marked as the current reasoning time step and reasoning is performed, at least one corresponding initial reasoning step is obtained based on the initial question and the obtained problem reasoning steps by autoregression, and the problem reasoning steps corresponding to the current reasoning time step are obtained in sequence by looping through the above steps. The specific implementation is as follows: The problem reasoning steps corresponding to the current reasoning time step are obtained in sequence by the autoregressive method until the current reasoning time step is the end reasoning time step, including: Obtaining at least one initial reasoning step corresponding to the current reasoning time step by means of the autoregression method, and determining the problem reasoning step from the at least one initial reasoning step; When it is determined that there is a next reasoning time step for the current reasoning time step, marking the next reasoning time step as the current reasoning time step, Continue to execute the autoregressive method to obtain at least one initial reasoning step corresponding to the current reasoning time step, and determine the problem reasoning step from the at least one initial reasoning step until the current reasoning time step is the end reasoning time step.

[0056] Specifically, in the current reasoning time step, when the question-answering model generates reasoning steps, it usually generates multiple possible candidate reasoning steps (i.e., initial reasoning steps), and when generating multiple candidate reasoning steps, each candidate reasoning step has a corresponding probability value, which indicates the possibility that the question-answering model believes that this candidate reasoning step is correct. Therefore, the question-answering model can select the candidate reasoning step with a higher probability value as the question reasoning step of the current reasoning time step based on the probability value corresponding to the candidate reasoning step.

[0057] For example, using the above example, in the current reasoning time step, the initial reasoning steps sampled by the question-answering model include the first initial reasoning step being "Xiao Li initially had 4 apples", the second initial reasoning step being "Xiao Li has 4 apples", and the third initial reasoning step being "Xiao Li initially had 4 apples". The question-answering model will calculate the probability of each initial reasoning step. For example, the probability of the first initial reasoning step is 0.5, the probability of the second initial reasoning step is 0.2, and the probability of the third initial reasoning step is 0.3. Finally, the first initial reasoning step with the highest probability is selected as the question reasoning step for the current reasoning time step.

[0058] If the current reasoning time step has not generated a complete reasoning process, that is, there is still a next reasoning time step, the question-answering model needs to continue to generate subsequent reasoning steps. Specifically, the next reasoning time step is marked as the current reasoning time step, and at least one initial reasoning step corresponding to the current reasoning time step is continued to be generated. In fact, in the new current reasoning time step, the question-answering model will continue to generate at least one initial reasoning step by autoregression, and determine the question reasoning step of the current reasoning time step from at least one initial reasoning step. This process will be repeated until the current reasoning time step is the end reasoning time step and the reasoning process ends.

[0059] In specific implementation, the ending reasoning time step may be the reasoning time step for generating the reasoning answer, or the reasoning time step that reaches the set maximum number of reasoning time steps; in actual applications, when the question-answering model is gradually performing question reasoning, the reasoning step corresponding to the ending reasoning time step usually includes an end symbol, so when the question reasoning step generated by the current reasoning time step includes an end symbol, the current reasoning time step is the ending reasoning time step; in addition, by setting the maximum number of reasoning time steps, the question-answering model can be prevented from falling into an infinite loop when generating reasoning steps, ensuring that the reasoning process can be terminated within a reasonable time and avoiding unnecessary waste of resources, so when the current reasoning time step reaches the maximum number of reasoning time steps, the current reasoning time step is determined as the ending reasoning time step.

[0060] The data processing method provided in the embodiments of this specification selects a problem reasoning step in at least one initial reasoning step corresponding to each current reasoning time step, thereby being able to obtain at least one problem reasoning step corresponding to the initial problem from the initial problem to the end reasoning time step, which is composed of the problem reasoning steps corresponding to each current reasoning time step.

[0061] In one or more embodiments of this specification, when the initial question and the initial answer are expanded to obtain feature data that enhances the model reasoning ability, in order to obtain as much feature data as possible, it is necessary to generate various question reasoning steps based on an initial question. The specific implementation is as follows: The problem reasoning steps corresponding to each current reasoning time step are obtained in sequence by the autoregressive method until the current reasoning time step is the end reasoning time step, including: Obtaining multiple initial reasoning steps corresponding to the current reasoning time step by means of the autoregressive method, and determining the multiple initial reasoning steps as problem reasoning steps of multiple reasoning paths respectively; For each reasoning path, if there is a next reasoning time step in the current reasoning time step, mark the next reasoning time step as the current reasoning time step, Continue to execute the autoregressive method to obtain multiple initial reasoning steps corresponding to the current reasoning time step, and determine the multiple initial reasoning steps as problem reasoning steps of multiple reasoning paths, until the current reasoning time step is the end reasoning time step.

[0062] Among them, the question-answering model may explore different reasoning paths when generating question reasoning steps, and each reasoning path corresponds to a possible reasoning process.

[0063] For example, in the first current reasoning time step, the initial reasoning steps for generating the initial problem include a first initial reasoning step, a second initial reasoning step, and a third initial reasoning step. These three initial reasoning steps are respectively used as the problem reasoning steps corresponding to the first reasoning path, the second reasoning path, and the third reasoning path in the first current reasoning time step.

[0064] Taking the first reasoning path as an example, when it is determined that there is a next reasoning time step for the first current reasoning time step, that is, the second current reasoning time step, in the second current reasoning time step, the initial problem and the first initial reasoning step corresponding to the first reasoning path are used to obtain the fourth initial reasoning step and the fifth initial reasoning step corresponding to the second current reasoning time step through autoregression. At this time, it is equivalent to the first reasoning path branching at the second current reasoning time step, expanding the first reasoning path into the fourth reasoning path and the fifth reasoning path.

[0065] That is, the problem reasoning steps in the fourth reasoning path include the first initial reasoning step corresponding to the first current reasoning time step and the fourth initial reasoning step corresponding to the second current reasoning time step, and the fifth reasoning path includes the first initial reasoning step corresponding to the first current reasoning time step and the fifth initial reasoning step corresponding to the second current reasoning time step. By analogy, when at least one initial reasoning step is generated in the current reasoning time step, the reasoning path corresponding to the previous current reasoning time step can be expanded to form a new reasoning path, until the current reasoning time step is the end reasoning time step, multiple reasoning paths can be obtained.

[0066] The data processing method provided in the embodiments of this specification generates at least one initial reasoning step and determines each initial reasoning step as a question reasoning step. The question-answering model can explore and obtain multiple reasoning paths, thereby increasing the diversity of the reasoning process.

[0067] In one or more embodiments of the present specification, by taking each reasoning path as the target reasoning path, the problem reasoning step corresponding to each reasoning path at each current reasoning time step can be determined as the at least one problem reasoning step, that is, at least one different problem reasoning step corresponding to multiple reasoning paths can be obtained for an initial problem, thereby increasing the diversity of reasoning data. The specific implementation is as follows: The obtaining of the at least one problem reasoning step according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence includes: Determining the multiple reasoning paths as target reasoning paths in sequence; The at least one problem reasoning step is obtained according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence in the target reasoning path.

[0068] Specifically, the reasoning path includes the problem reasoning steps of the entire process from the first step to the last step, that is, for each reasoning path, the problem reasoning steps corresponding to each current reasoning time step can be determined. When each reasoning path is used as the target reasoning path and the at least one problem reasoning step is a problem reasoning step included in the target reasoning path, the problem reasoning steps included in each reasoning path can be used as the at least one problem reasoning step, so that for an initial problem, at least one problem reasoning step corresponding to different reasoning paths can be obtained.

[0069] In the above embodiment, by determining one reasoning step as a problem reasoning step among multiple initial reasoning steps corresponding to each current reasoning time step, a reasoning path corresponding to the initial problem will eventually be obtained, that is, the reasoning path is composed of a problem reasoning step corresponding to each current reasoning time step; the way in which the above embodiment obtains at least one problem reasoning step can be called chain construction, that is, based on a problem reasoning step corresponding to each current reasoning time step, at least one problem reasoning step of at least one reasoning time step is obtained in a chain.

[0070] In this implementation, multiple initial reasoning steps corresponding to each current reasoning time step are determined as problem reasoning steps. As the reasoning process proceeds, multiple initial reasoning steps can be horizontally expanded at each current reasoning time step. Based on each current reasoning time step and the multiple problem reasoning steps corresponding to each current reasoning time step, multiple reasoning paths can be formed. Based on each reasoning path in the multiple reasoning paths, at least one problem reasoning step corresponding to each reasoning path can be obtained. Therefore, the way in which this embodiment obtains at least one problem reasoning step can be called tree construction.

[0071] The data processing method provided in the embodiments of this specification obtains multiple reasoning paths through exploration, and when each reasoning path is taken as the target reasoning path in turn, according to the problem reasoning steps corresponding to each current reasoning time step contained in the target reasoning path, it is possible to obtain at least one different problem reasoning step corresponding to multiple reasoning paths for an initial problem, thereby increasing the diversity of the reasoning data.

[0072] Step 204: According to the initial answer corresponding to the initial question, the correctness of the at least one question reasoning step and the reasoning answer is verified to obtain a verification result.

[0073] Among them, correctness verification can be understood as checking and verifying whether the reasoning steps and answers to the problem are logical and whether they can correctly solve the initial problem; the verification result can be understood as the final result of the verification process, and the verification result includes both correct and incorrect results.

[0074] Specifically, in an embodiment of the present specification, there is a judgment module for verifying the correctness of at least one question reasoning step and reasoning answer output by the question-answering model to determine whether the question reasoning step is a correct reasoning step and whether the reasoning answer is the correct answer to the initial question.

[0075] By verifying the correctness of at least one question reasoning step and the reasoning answer through the initial answer to the initial question, it can be determined whether the at least one question reasoning step and the reasoning answer need to be corrected later.

[0076] In one or more embodiments of this specification, the first verification result, i.e., the correct result, can be obtained only when at least one question reasoning step and the reasoning answer are correct; the second verification result, i.e., the wrong result, can be obtained when at least one question reasoning step or the reasoning answer is wrong. The specific implementation is as follows: The step of performing correctness verification on the at least one question reasoning step and the reasoning answer according to the initial answer corresponding to the initial question to obtain a verification result includes: Performing a correctness check on the at least one question reasoning step and the reasoning answer according to the initial answer; When it is determined that the at least one question reasoning step and the reasoning answer are both correct, obtaining a first verification result; In the case where it is determined that the at least one question reasoning step and / or the reasoning answer is wrong, a second verification result is obtained.

[0077] Among them, the first verification result can be understood as a correct result, indicating that at least one question reasoning step and the reasoning answer are correct; the second verification result can be understood as an incorrect result, indicating that there is an error in at least one question reasoning step and / or the reasoning answer.

[0078] It should be noted that, in the case where at least one question reasoning step is a plurality of question reasoning steps, when there is an erroneous question reasoning step among the plurality of question reasoning steps, a second verification result will be obtained.

[0079] Specifically, the judgment module verifies the correctness of the question reasoning steps and reasoning answers generated by the question-answering model based on the initial answer, such as checking the consistency of the reasoning answer with the initial answer. If the reasoning answer is consistent with the initial answer, the reasoning answer is determined to be the correct answer to the initial question; the logical consistency of the question reasoning steps is verified to determine whether each step of the deduction is consistent with common sense and domain knowledge. If each step of the deduction is consistent with common sense and domain knowledge and the derived reasoning answer is consistent with the initial answer, the question reasoning steps and reasoning answers are determined to be correct.

[0080] If the question reasoning steps and the reasoning answer are correct, a first verification result is obtained; if the question reasoning steps and / or the reasoning answer are incorrect, a second verification result is obtained.

[0081] The data processing method provided in the embodiments of this specification can ensure the accuracy of problem reasoning steps and reasoning answers through correctness verification, and can provide data support for subsequent model adjustments based on correct problem reasoning steps and reasoning answers, and can also correct erroneous problem reasoning steps or reasoning answers in a timely manner.

[0082] Step 206: Obtain target data based on the verification result, the initial question, and the at least one question reasoning step, wherein the target data includes the initial question, the target reasoning step, and the target answer corresponding to the initial question.

[0083] Among them, the target reasoning step can be understood as the correct reasoning step corresponding to the initial problem. When at least one problem reasoning step is correct, the target reasoning step is the at least one problem reasoning step. When there is an erroneous reasoning step in at least one problem reasoning step, at least one problem reasoning step can be corrected to obtain the corrected target reasoning step.

[0084] The target data can be understood as training data used for subsequent training models. The feature data in the training data is the initial question, and the label data includes the target reasoning steps and the target answer corresponding to the initial question.

[0085] Specifically, when the verification result is the first verification result, the problem reasoning steps and the reasoning answers are correct. At this time, at least one problem reasoning step can be directly determined as the target reasoning step, and the reasoning answer can be determined as the target answer; when the verification result is the second verification result, there are errors in the problem reasoning steps and / or the reasoning answers. At this time, the problem reasoning steps or reasoning answers need to be corrected to obtain the target reasoning steps and target answers.

[0086] In practical applications, further integration or correction of the verification results can ensure the accuracy of the target data, thereby improving the reasoning ability of the model when the model is trained based on correct, high-quality target data.

[0087] In one or more embodiments of this specification, by determining the initial question as feature data, and determining at least one question reasoning step and the target answer as label data, target data for training the long reasoning ability of the model can be obtained, so that training data for the long reasoning link can be obtained by expanding the common initial question and initial answer. The specific implementation is as follows: The step of obtaining target data according to the verification result, the initial question and the at least one question reasoning step includes: According to the first verification result, determining the initial question as feature data; Determine the at least one question reasoning step as a target reasoning step, determine the reasoning answer as the target answer, and determine the target reasoning step and the target answer as label data; The target data is obtained according to the feature data and the label data.

[0088] Among them, in machine learning, feature data refers to the input data of the model. In the embodiment of this specification, the initial question is used as feature data to represent the input of the model; the target answer can be understood as the final correct answer determined after verification, and the label data is used to represent the output target of the model. In the embodiment of this specification, at least one question reasoning step and reasoning answer are used as label data to represent the output target of the model.

[0089] Specifically, when the verification result is the "first verification result", the initial question is used as the feature data, the inference answer is used as the target answer, and at least one question reasoning step is determined as the target reasoning step, so that the target reasoning step and the target answer are used as label data to finally generate the target data. The target data can be understood as the shortcut sample in the above embodiment, which is obtained directly through the initial question, at least one correct question reasoning step and reasoning answer output by the question-answering model.

[0090] The data processing method provided in the embodiments of this specification can expand the original question-answer pair consisting of an initial question and an initial answer into target data containing reasoning data such as at least one question reasoning step, by obtaining label data based on at least one question reasoning step and reasoning answer corresponding to a first verification result, thereby achieving an expansion of the original data and increasing the training data for training the model's reasoning ability.

[0091] In one or more embodiments of the present specification, when the verification result is the second verification result, that is, when at least one question reasoning step or reasoning answer is wrong, the correction model can be used to correct the wrong question reasoning step or reasoning answer to obtain correct target data. The specific implementation is as follows: The step of obtaining target data according to the verification result, the initial question and the at least one question reasoning step includes: Determine a reference reasoning step from the at least one problem reasoning step according to the second verification result, wherein the reference reasoning step is a reasoning step in the at least one problem reasoning step, starting from the first problem reasoning step to the reasoning step where the error initially occurs; Inputting the reference reasoning step, the initial question and the initial answer into a correction model to obtain at least one corrected reasoning step and a corrected answer after correction; In a case where the verification result of the at least one correcting reasoning step and the correcting answer is a first verification result, determining the initial question as feature data, determining the reference reasoning step and the at least one correcting reasoning step as the target reasoning step, determining the correcting answer as the target answer, and determining the target reasoning step, the at least one correcting reasoning step, and the target answer as label data; The target data is obtained according to the feature data and the label data.

[0092] Among them, the reference reasoning step can be understood as the reasoning step in at least one problem reasoning step, starting from the first problem reasoning step to the reasoning step where the error first occurs; for example, for a math problem, at least one problem reasoning step includes the first problem reasoning step, the second problem reasoning step and the third problem reasoning step. When an error occurs in the second problem reasoning step, resulting in an incorrect reasoning answer, the reference reasoning step includes the first problem reasoning step and the second problem reasoning step.

[0093] The correction model can be understood as a large model with a larger size and stronger reasoning ability than the above-mentioned question-answering model, so that the correction model can correct the erroneous question reasoning steps output by the question-answering model; the correction reasoning steps can be understood as the corrected reasoning steps generated by the correction model, which are corrections and / or supplements to the reference reasoning steps; the corrected answer is the final correct answer determined after correction.

[0094] Specifically, by inputting the reference reasoning steps, the initial question and the correct initial answer corresponding to the initial question into the correction model, the correction model can reflect on and correct the reference reasoning steps obtained previously, and obtain at least one correct reasoning step and a correct answer; in fact, when the verification result of at least one correct reasoning step and the correct answer is the first verification result, that is, the verification result is correct, a correct reasoning step and a target answer that are tortuous but ultimately correct can be obtained.

[0095] The data processing method provided in the embodiments of this specification uses the reference reasoning steps, the corrected reasoning steps after correction, and the corrected answers as label data, so that the model can be enhanced in the subsequent training of the model, that is, the reflective error correction ability of the model can be enhanced on the basis of the model's long reasoning ability.

[0096] In one or more embodiments of this specification, when target data is obtained, the model to be adjusted can be trained with the target data to obtain a target model with stronger reasoning ability, and the training data of the extended reasoning model is realized by expanding the initial question and the initial answer into target data with a long reasoning link. The specific implementation is as follows: After obtaining the target data according to the verification result, the initial question and the at least one question reasoning step, the method further includes: Determine the model to be adjusted; The target data is used to train the model to be adjusted to obtain a target model.

[0097] The model to be adjusted can be understood as a model that needs to be adjusted or trained. It can be an existing large language model or a model specifically used for reasoning tasks. The target model can be understood as a model obtained after training. The target model has stronger reasoning ability and can generate reasoning steps and answers corresponding to questions more accurately.

[0098] Specifically, the system can select a suitable model as the model to be adjusted according to task requirements, and use the target data to train the model to be adjusted, so that the model to be adjusted can better understand and solve problems that require reasoning; the specific training process usually includes inputting feature data (initial question) and label data (target reasoning steps and target answers), so that the model to be adjusted can learn how to derive the correct reasoning steps and answers from the problem.

[0099] The data processing method provided in the embodiments of this specification utilizes target data to train the model to be adjusted so that the obtained target model can better understand and solve problems that require complex reasoning, thereby improving its reasoning ability. When the target data contains a variety of target reasoning steps and target answers, the trained target model can better generalize to new and unseen reasoning problems and can enhance certain model reflection capabilities.

[0100] In one or more embodiments of this specification, the target model is obtained by calculating the loss function of the prediction result and the label data, and then using an optimization algorithm to adjust the model parameters of the model to be adjusted. The specific implementation is as follows: The step of training the model to be adjusted by using the target data to obtain a target model includes: Inputting the characteristic data into the model to be adjusted to obtain a prediction result output by the model to be adjusted; According to the prediction result and the label data, the parameters of the model to be adjusted are adjusted to obtain the target model.

[0101] Among them, the prediction result can be understood as the output result generated by the model to be adjusted based on the feature data, including medical reasoning steps and predicted answers.

[0102] Specifically, by comparing the prediction results and the label data, calculating the loss function, and then using an optimization algorithm (such as gradient descent) to adjust the model parameters of the model to be adjusted, the target model obtained after the model parameters are adjusted and the corresponding output results can be closer to the label data.

[0103] The data processing method provided in the embodiments of this specification inputs feature data into the model to be adjusted during the training process to generate prediction results, and then adjusts the parameters of the model based on the prediction results and label data, ultimately obtaining a target model with stronger reasoning capabilities.

[0104] The data processing method provided in the embodiments of this specification can automatically construct target data of long reasoning links through the question-answering model for existing question-answering data with simple answers. Combined with model fine-tuning and the use of target data, the reasoning ability of the model to be adjusted can be effectively improved; when constructing reasoning path data, two construction methods, chain construction and tree construction, are included, so that more diverse reasoning path data (the reasoning path data contains at least one question reasoning step) can be obtained, and reflection and correction can be performed on erroneous reasoning path data, so that the target data obtained after correction can be used to perform model enhancement training on the model to be adjusted, which is conducive to enhancing model reflection and reasoning ability.

[0105] In one embodiment of the present specification, a model training method is also provided, including: Determine the model to be adjusted; The model to be adjusted is trained using target data to obtain a target model, wherein the target data is obtained by the above-mentioned data processing method.

[0106] Specifically, when the target data is obtained through the above-mentioned data processing method, the target data includes feature data and label data. The feature data is input into the model to be adjusted to obtain the prediction result of the output of the model to be adjusted. According to the prediction result and the label data, the parameters of the model to be adjusted are adjusted to obtain the target model.

[0107] When the label data of the target data includes the target reasoning step, the target model obtained through training can have stronger reasoning ability.

[0108] See also Figure 3 , Figure 3 A schematic diagram of a chain data construction process of a data processing method provided in one embodiment of the present specification.

[0109] In actual applications, there are two links in the chain data construction process. The first link is to input the initial question into the question-answering model (the question-answering model can be a large language model), and use the initial question and sample data (few-shots) in the question-answering model to generate the reasoning process (i.e., at least one question reasoning step in the above embodiment, including the question reasoning step corresponding to each reasoning time step. In the embodiments of this specification, the entire process of analyzing the initial question, the first question reasoning step to the final reasoning answer generation is called a path, i.e., a reasoning path) and the final reasoning answer at one time. The other link inputs the initial question and the initial answer, and allows the question-answering model to generate the reasoning process and the reasoning answer at one time. This link is equivalent to giving a reference answer to generate the reasoning process and the reasoning answer.

[0110] Since the question-answering model may output incorrect question reasoning steps or fabricate question reasoning steps during the generation process, a judgment module is introduced. The judgment module can be a larger language model or expert / manual annotation, so as to use the initial answer to collect incorrect reasoning paths and correct reasoning paths.

[0111] For the correct reasoning path, the initial question, reasoning process and reasoning answer corresponding to the reasoning path are directly saved as shortcut samples (i.e., the first target data). For the wrong reasoning path, the reasoning process is split, and the reference reasoning steps (from the first question reasoning step to the wrong question reasoning step), the initial question and the initial answer determined by the split are input into the error correction model (the error correction model can also be a large model), and the error correction model is allowed to perform error correction and then generate subsequent steps (i.e., the error correction reasoning steps in the above embodiment), and obtain the correct target answer. Finally, the correct reasoning path is retained, and based on the initial question, the target reasoning step and the target answer, a backtracking sample (i.e., the second target data) is obtained.

[0112] See also Figure 4 , Figure 4 A schematic diagram of a tree data construction process of a data processing method provided in an embodiment of the present specification is shown.

[0113] Figure 4 In the example, when the question-answering model performs reasoning on the initial question, at least one initial reasoning step can be obtained at each reasoning time step, and the reasoning step is continuously expanded until the final reasoning answer is generated or the maximum set tree depth (i.e., the maximum number of reasoning time steps in the above embodiment) is reached.

[0114] Similarly, the judgment module is used to verify the correctness of each generated reasoning answer. If the verification result is correct, the initial question, reasoning process and reasoning answer on the reasoning path are saved as shortcut samples. If the verification result is incorrect, the initial answer and a larger and more powerful large language model (i.e., the error correction model in the above embodiment) are introduced to further reflect on the error correction, and finally the correct target answer is generated, and the backtracking sample is obtained based on the same method as above.

[0115] The shortcut samples and backtracking samples obtained based on the above chain generation or tree construction can be used to adjust the training of the basic model to be adjusted (i.e., the model to be adjusted in the above embodiment), and the method of adjusting the training can be full parameter fine-tuning or LoRA (Low-Rank Adaptation) or other peft (Parameter-Efficient Fine-Tuning, a parameter efficient fine-tuning) training method, which is not limited here.

[0116] The data processing method provided in the embodiments of this specification can automatically construct target data of long reasoning links through question-answering models for existing question-answering data with simple answers, and can effectively improve the reasoning ability of existing models in combination with model fine-tuning. Specifically, through the chain and tree-shaped reasoning path data construction methods, shortcut long reasoning data and backtracking longer reasoning data can be constructed, and combined with the horizontal expansion of the tree, more diverse reasoning path data can be generated; and a larger-sized model is combined to realize the correction function, that is, to reflect on and correct the reasoning steps of the problem, and obtain longer reasoning data that is tortuous but ultimately correct. Then, when these longer reasoning data are used for model enhancement training, it can be beneficial to enhance the model reflection and reasoning ability.

[0117] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 5 FIG. 1 is a schematic diagram showing the structure of a data processing device provided by an embodiment of the present specification. Figure 5 As shown, the device comprises: The reasoning module 502 is configured to determine at least one question reasoning step corresponding to the initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step; A verification module 504 is configured to verify the correctness of the at least one question reasoning step and the reasoning answer according to the initial answer corresponding to the initial question, and obtain a verification result; The acquisition module 506 is configured to obtain target data based on the verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the target reasoning step, and the target answer corresponding to the initial question.

[0118] Optionally, the reasoning module 502 is configured to: Identify initial questions and sample data; The initial question and the example data are input into a question-answering model, or the initial question, the initial answer corresponding to the initial question, and the example data are input into a question-answering model, and the question-answering model is used to obtain at least one question reasoning step corresponding to the initial question and an reasoning answer corresponding to the initial question, wherein the reasoning answer is obtained based on the at least one question reasoning step.

[0119] Optionally, the verification module 504 is configured to: Performing a correctness check on the at least one question reasoning step and the reasoning answer according to the initial answer; When it is determined that the at least one question reasoning step and the reasoning answer are both correct, obtaining a first verification result; In the case where it is determined that the at least one question reasoning step and / or the reasoning answer is wrong, a second verification result is obtained.

[0120] Optionally, the obtaining module 506 is configured to: According to the first verification result, determining the initial question as feature data; Determine the at least one question reasoning step as a target reasoning step, determine the reasoning answer as the target answer, and determine the target reasoning step and the target answer as label data; The target data is obtained according to the feature data and the label data.

[0121] Optionally, the obtaining module 506 is configured to: Determine a reference reasoning step from the at least one problem reasoning step according to the second verification result, wherein the reference reasoning step is a reasoning step in the at least one problem reasoning step, starting from the first problem reasoning step to the reasoning step where the error initially occurs; Inputting the reference reasoning step, the initial question and the initial answer into a correction model to obtain at least one corrected reasoning step and a corrected answer after correction; In a case where the verification result of the at least one correcting reasoning step and the correcting answer is a first verification result, determining the initial question as feature data, determining the reference reasoning step and the at least one correcting reasoning step as the target reasoning step, determining the correcting answer as the target answer, and determining the target reasoning step, the at least one correcting reasoning step, and the target answer as label data; The target data is obtained according to the feature data and the label data.

[0122] Optionally, the reasoning module 502 is configured to: In the question-answering model, the question reasoning steps corresponding to each current reasoning time step are obtained in sequence by an autoregressive method until the current reasoning time step is the end reasoning time step; The at least one problem reasoning step is obtained according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence.

[0123] Optionally, the reasoning module 502 is configured to: Obtaining at least one initial reasoning step corresponding to the current reasoning time step by means of the autoregression method, and determining the problem reasoning step from the at least one initial reasoning step; When it is determined that there is a next reasoning time step for the current reasoning time step, marking the next reasoning time step as the current reasoning time step, Continue to execute the autoregressive method to obtain at least one initial reasoning step corresponding to the current reasoning time step, and determine the problem reasoning step from the at least one initial reasoning step until the current reasoning time step is the end reasoning time step.

[0124] Optionally, the reasoning module 502 is configured to: Obtaining multiple initial reasoning steps corresponding to the current reasoning time step by means of the autoregressive method, and determining the multiple initial reasoning steps as problem reasoning steps of multiple reasoning paths respectively; For each reasoning path, if there is a next reasoning time step in the current reasoning time step, mark the next reasoning time step as the current reasoning time step, Continue to execute the autoregressive method to obtain multiple initial reasoning steps corresponding to the current reasoning time step, and determine the multiple initial reasoning steps as problem reasoning steps of multiple reasoning paths, until the current reasoning time step is the end reasoning time step.

[0125] Optionally, the reasoning module 502 is configured to: Determining the multiple reasoning paths as target reasoning paths in sequence; The at least one problem reasoning step is obtained according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence in the target reasoning path.

[0126] The device further comprises: The training module is configured to determine the model to be adjusted; train the model to be adjusted using the target data to obtain a target model.

[0127] Optionally, the training module is configured as follows: Inputting the characteristic data into the model to be adjusted to obtain a prediction result output by the model to be adjusted; According to the prediction result and the label data, the parameters of the model to be adjusted are adjusted to obtain the target model.

[0128] The data processing device provided in the embodiments of the present specification can obtain reasoning path data including the question reasoning steps corresponding to the initial question by determining at least one question reasoning step corresponding to the initial question and the reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step, and when the correctness of the at least one question reasoning step and the reasoning answer is verified according to the initial answer corresponding to the initial question, correct target data can be obtained through the obtained verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the target reasoning step, and the target answer corresponding to the initial question; compared with a simple question-answer pair including the initial question and the initial answer, the target data adds a target reasoning step, thereby realizing the expansion of a long reasoning chain for a simple question-answer pair, and subsequently providing effective training data for improving the model reasoning capability.

[0129] The above is a schematic scheme of a data processing device of this embodiment. It should be noted that the technical scheme of the data processing device and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data processing device that are not described in detail can be referred to the description of the technical scheme of the above data processing method.

[0130] Figure 6 The block diagram of a computing device 600 according to an embodiment of the present specification is shown. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and the database 650 is used to store data.

[0131] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0132] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 6 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0133] The computing device 600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 may also be a mobile or stationary server.

[0134] The processor 620 is used to execute the following computer program / instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.

[0135] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the data processing method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the data processing method embodiment.

[0136] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.

[0137] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the data processing method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the data processing method embodiment.

[0138] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above data processing method when executed by a processor.

[0139] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above data processing method belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of the above data processing method.

[0140] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0141] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0142] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0143] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0144] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Determine at least one question reasoning step corresponding to an initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step; According to the initial answer corresponding to the initial question, the correctness of the at least one question reasoning step and the reasoning answer is verified to obtain a verification result; Target data is obtained according to the verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the target reasoning step, and the target answer corresponding to the initial question.

2. The data processing method according to claim 1, wherein the determining of at least one question reasoning step corresponding to the initial question and the reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step include: Identify initial questions and sample data; The initial question and the example data are input into a question-answering model, or the initial question, the initial answer corresponding to the initial question, and the example data are input into a question-answering model, and the question-answering model is used to obtain at least one question reasoning step corresponding to the initial question and an reasoning answer corresponding to the initial question, wherein the reasoning answer is obtained based on the at least one question reasoning step.

3. The data processing method according to claim 1, wherein the correctness of the at least one question reasoning step and the reasoning answer is verified according to the initial answer corresponding to the initial question to obtain the verification result, comprising: Performing a correctness check on the at least one question reasoning step and the reasoning answer according to the initial answer; When it is determined that the at least one question reasoning step and the reasoning answer are both correct, obtaining a first verification result; In the case where it is determined that the at least one question reasoning step and / or the reasoning answer is wrong, a second verification result is obtained.

4. The data processing method according to claim 3, wherein obtaining target data according to the verification result, the initial question and the at least one question reasoning step comprises: According to the first verification result, determining the initial question as feature data; Determine the at least one question reasoning step as a target reasoning step, determine the reasoning answer as the target answer, and determine the target reasoning step and the target answer as label data; The target data is obtained according to the feature data and the label data.

5. The data processing method according to claim 3, wherein obtaining target data according to the verification result, the initial question and the at least one question reasoning step comprises: Determine a reference reasoning step from the at least one problem reasoning step according to the second verification result, wherein the reference reasoning step is a reasoning step from the first problem reasoning step to the reasoning step where the error initially occurs in the at least one problem reasoning step; Inputting the reference reasoning step, the initial question and the initial answer into a correction model to obtain at least one corrected reasoning step and a corrected answer after correction; In a case where the verification result of the at least one corrective reasoning step and the corrective answer is a first verification result, determining the initial question as feature data, determining the reference reasoning step and the at least one corrective reasoning step as the target reasoning step, determining the corrective answer as the target answer, and determining the target reasoning step and the target answer as label data; The target data is obtained according to the feature data and the label data.

6. The data processing method according to claim 2, wherein the step of obtaining the at least one question reasoning step corresponding to the initial question using the question-answering model comprises: In the question-answering model, the question reasoning steps corresponding to the current reasoning time step are obtained in sequence by an autoregressive method until the current reasoning time step is the end reasoning time step; The at least one problem reasoning step is obtained according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence.

7. The data processing method according to claim 6, wherein the step of sequentially obtaining the problem reasoning steps corresponding to the current reasoning time step by means of autoregression until the current reasoning time step is the end reasoning time step comprises: Obtaining at least one initial reasoning step corresponding to the current reasoning time step by means of the autoregression method, and determining the problem reasoning step from the at least one initial reasoning step; When it is determined that there is a next reasoning time step for the current reasoning time step, marking the next reasoning time step as the current reasoning time step, Continue to execute the autoregressive method to obtain at least one initial reasoning step corresponding to the current reasoning time step, and determine the problem reasoning step from the at least one initial reasoning step until the current reasoning time step is the end reasoning time step.

8. The data processing method according to claim 6, wherein the step of sequentially obtaining the problem reasoning steps corresponding to the current reasoning time step by means of autoregression until the current reasoning time step is the end reasoning time step comprises: Obtaining multiple initial reasoning steps corresponding to the current reasoning time step by means of the autoregressive method, and determining the multiple initial reasoning steps as problem reasoning steps of multiple reasoning paths respectively; For each reasoning path, if there is a next reasoning time step in the current reasoning time step, mark the next reasoning time step as the current reasoning time step, Continue to execute the autoregressive method to obtain multiple initial reasoning steps corresponding to the current reasoning time step, and determine the multiple initial reasoning steps as problem reasoning steps of multiple reasoning paths, until the current reasoning time step is the end reasoning time step.

9. The data processing method according to claim 8, wherein the step of obtaining the at least one problem reasoning step according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence comprises: Determining the multiple reasoning paths as target reasoning paths in sequence; The at least one problem reasoning step is obtained according to the problem reasoning steps corresponding to the current reasoning time step obtained in sequence in the target reasoning path.

10. The data processing method according to claim 4 or 5, after obtaining the target data according to the verification result, the initial question and the at least one question reasoning step, further comprising: Determine the model to be adjusted; The target data is used to train the model to be adjusted to obtain a target model.

11. The data processing method according to claim 10, wherein the step of training the model to be adjusted using the target data to obtain the target model comprises: Inputting the characteristic data into the model to be adjusted to obtain a prediction result output by the model to be adjusted; According to the prediction result and the label data, the parameters of the model to be adjusted are adjusted to obtain the target model.

12. A model training method, comprising: Determine the model to be adjusted; The model to be adjusted is trained using target data to obtain a target model, wherein the target data is obtained by the data processing method described in any one of claims 1-11.

13. A data processing device, comprising: A reasoning module, configured to determine at least one question reasoning step corresponding to an initial question, and a reasoning answer corresponding to the initial question obtained based on the at least one question reasoning step; A verification module is configured to verify the correctness of the at least one question reasoning step and the reasoning answer according to the initial answer corresponding to the initial question, and obtain a verification result; An acquisition module is configured to obtain target data based on the verification result, the initial question and the at least one question reasoning step, wherein the target data includes the initial question, the reasoning step, and a target answer corresponding to the initial question.

14. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

16. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Question and answer model training method, question and answer method and device

    CN114925178A

  • Large model correction result determination method and device, computer equipment and storage medium

    CN117033960A

  • Large model illusion treatment method, device and equipment and storage medium

    CN117556920A

  • Time reasoning method based on large model, electronic equipment and storage medium

    CN118211655A

  • Question and answer model training method, text processing method and reward model training method

    CN118350463A

Cited By

  • Instruction fine tuning data set generation method, electronic equipment and storage medium

    CN121579077A