Data processing method, device and program product
By using cross-task advanced reflective information guidance model reasoning in autonomous language agents, the problem of accuracy and appropriateness of content generation of large language models is solved, and effective self-correction and self-reflection in the absence of real labels is achieved.
Patent Information
- Application Number
- CN202510233116.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-27
AI Technical Summary
In autonomous language agents, there are problems with accuracy and appropriateness in the content generated by large language models, especially in the absence of real labels, which makes it difficult to effectively self-correct and self-reflection.
By using advanced reflection information across tasks, we guide the generation of specific reflection information of the previous round of model reasoning, and determine the prompt information of the current round of model reasoning based on this reflection information, and then adjust the model reasoning steps to achieve self-correction and self-reflection.
It improves the accuracy and appropriateness of the content generated by the model, avoids dependence on real tags, and enhances the feasibility of the model in practical application scenarios.
Smart Images

Figure CN120218232A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to, but is not limited to, the field of computer technology, and in particular, to a data processing method, apparatus, and program product. Background Art
[0002] Currently, large language models (LLMs) are designed as autonomous language agents, enabling them to perform goal-oriented multi-step tasks. However, in autonomous language agents, there are still accuracy and appropriateness issues in the content generated by large language models. Summary of the Invention
[0003] In view of this, the present disclosure provides at least a data processing method, apparatus, and program product.
[0004] The technical solution of the present disclosure is implemented as follows:
[0005] On the one hand, the present disclosure provides a data processing method, the method including:
[0006] Determining prompt information for the current round of model inference based on specific reflection information corresponding to the previous round of model inference; the specific reflection information corresponding to the previous round of model inference is generated based on the process information and high-level reflection information of the previous round of model inference; the high-level reflection information represents cross-task reflection information generated by the target model;
[0007] An inference module, based on the prompt information for the current round of model inference, using the target model to perform the current round of model inference, obtaining the inference result of the current round of model inference;
[0008] A second determination module, determining a target inference result based on the inference results of at least two rounds of model inference.
[0009] On the other hand, the present disclosure provides a data processing apparatus, the apparatus including:
[0010] A first determination module, determining prompt information for the current round of model inference based on specific reflection information corresponding to the previous round of model inference; the specific reflection information corresponding to the previous round of model inference is generated based on the process information and high-level reflection information of the previous round of model inference; the high-level reflection information represents cross-task reflection information generated by the target model;
[0011] An inference module, based on the prompt information for the current round of model inference, using the target model to perform the current round of model inference, obtaining the inference result of the current round of model inference;
[0012] A second determination module, determining a target inference result based on the inference results of at least two rounds of model inference.
[0013] In another aspect, the present disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented.
[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0016] Figure 1 It is a schematic flowchart of the implementation of a data processing method provided by the present disclosure;
[0017] Figure 2 It is a schematic flowchart of the implementation of an embodiment of the data processing method provided by the present disclosure;
[0018] Figure 3 It is a schematic diagram of updating high-level reflection information in the data processing method provided by the present disclosure;
[0019] Figure 4 It is a schematic diagram of the composition structure of a data processing device provided by the present disclosure;
[0020] Figure 5 It is a schematic diagram of the hardware entity of an electronic device provided by the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations to the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present disclosure.
[0022] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0023] The terms "first / second / third" involved herein are merely used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this disclosure belongs. The terms used herein are for the purpose of describing this disclosure only and are not intended to limit this disclosure.
[0025] To improve the accuracy and appropriateness of the content generated by the LLM, related technologies have proposed using self-correction and self-reflection to make up for the limitations of the LLM. For example, in the Reflexion framework, when the model executes an autonomous language agent task, instead of updating the weight parameters of the model, it reinforces the language agent through language feedback. Specifically, the Reflexion framework utilizes the feedback signals of natural language reflection tasks (such as the correctness of the answer, etc.) and generates a specific reflection text buffered in the episodic memory for making better decisions in subsequent model inferences.
[0026] However, there are the following problems in using the self-correction and self-reflection capabilities of the model to improve the accuracy and appropriateness of the content generated by the model:
[0027] First, real labels are needed to guide the self-correction process. For example, after obtaining the initial response, in addition to generating a specific reflection text, the LLM will also add external feedback information for this initial response in the context: "The answer is correct" or "The answer is wrong". However, in the actual test field, these external feedback information cannot be used (that is, the standard answer to the question cannot be known in advance);
[0028] Second, real labels are needed to determine when to stop the self-correction loop, that is, real labels are needed to verify the answer at each step. If the answer is wrong, self-correction will be performed, and if the answer is correct, self-correction will not be performed. It can be seen that the way of using real labels here is counterintuitive, because only correcting wrong answers means that the answer is known;
[0029] Third, in the case of directly canceling the use of real labels and letting the model decide when to stop self-correction, after testing, the accuracy of the model (for example, the GPT-3.5 model and the GPT-4 model in the Generative Pre-trained Transformer (GPT) series of models) has decreased on datasets such as HotPotQA and CommonSenseQA.
[0030] Based on this, embodiments of the present disclosure provide a data processing method. First, based on the specific reflection information corresponding to the previous round of model inference, the prompt information for the current round of model inference is determined. Among them, the specific reflection information corresponding to the previous round of model inference is generated based on the process information and high-level reflection information of the previous round of model inference, and the high-level reflection information represents cross-task reflection information generated by the target model. Then, based on the prompt information for the current round of model inference, the target model is used to perform the current round of model inference to obtain the inference result of the current round of model inference. Finally, based on the inference results of at least two rounds of model inference, the target inference result is determined. In this way, on the one hand, the cross-task high-level reflection information is used to guide the generation of the specific reflection information corresponding to the previous round of model inference, and based on this specific reflection information, the prompt information for this round of model inference is determined, realizing the use of cross-task and summary experience information to guide the current model inference, so that the target model can more easily determine the correct model inference steps. On the other hand, the cross-task high-level reflection information is used to guide the model to generate specific reflection information. Compared with the solution in the related art that uses real labels to guide the model for self-correction, the present disclosure has higher feasibility in actual application scenarios.
[0031] The method provided by the present disclosure can be executed by an electronic device. The electronic device can be various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (such as a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device), etc., or can be implemented as a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0032] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present disclosure.
[0033] Figure 1 It is a schematic flowchart of the implementation of a data processing method provided by the present disclosure. As Figure 1 shown, the method includes the following steps S101 to step S103:
[0034] Step S101: Determine the prompt information for the current round of model inference based on the specific reflection information corresponding to the previous round of model inference. The specific reflection information corresponding to the previous round of model inference is generated based on the process information and high-level reflection information of the previous round of model inference. The high-level reflection information represents cross-task reflection information generated by the target model.
[0035] Here, after the target model receives the target task, for this target task, the target model iteratively performs multiple rounds of model inference to determine the target inference result of the target task.
[0036] The target model can be any type of model. For example, the target model can be an LLM, a vision language model (VLMs) based on an LLM, etc. In some embodiments, the target model can be a model improved to support the Reflexion framework.
[0037] The target task can be any type of model processing task. For example, the target task can be a text generation text task, a text generation image task, an image recognition task, an image generation text task, an image generation image task, etc. In some embodiments, the target task can be obtained in any suitable way. For example, the target task can be a task obtained through user input; or for another example, the target task can be a task determined by automatically identifying the interaction data between the user and other applications running on the electronic device; etc.
[0038] Each round of model inference is a model processing process completed by the target model according to the specified model inference rules for the target task, and each round of model inference has a corresponding inference result. In some embodiments, the model inference rules can be model inference rules determined according to the specified model inference examples, or can also be model inference rules determined according to the inference logic specified in natural language. The inference result corresponding to each round of model inference can indicate that the inference is completed, or can also indicate that the inference is not completed. For example, if the result of the target task is not obtained after completing the specified maximum number of inference steps, it indicates that the inference is not completed; if the result of the target task is obtained after completing the specified maximum number of inference steps, it indicates that the inference is completed.
[0039] The iterative execution of multiple rounds of model inference by the target model means that the target model sequentially executes multiple rounds of model inference, and in two adjacent rounds of model inference, in the latter round of model inference, at least the inference process of this round of model inference is adjusted based on the process information of the previous round of model inference. In the embodiments of the present disclosure, the current round of model inference determines the prompt information of the current round of model inference based on the specific reflection information corresponding to the previous round of model inference; wherein, the specific reflection information corresponding to the previous round of model inference is generated based on the process information of the previous round of model inference and the high-level reflection information; the high-level reflection information represents the cross-task reflection information generated by the target model.
[0040] Here, the process information of model inference, that is, the inference trajectory information of each round of model inference, refers to the steps or paths followed by the model when performing model inference. In implementation, the process information of model inference may include, but is not limited to, the processing information for input data, the feature extraction information for input data, the application information of the inference algorithm, the generation algorithm information for the inference result, etc.
[0041] Reflection information refers to the information obtained after analyzing and evaluating the model inference behavior and / or output during the inference process.
[0042] Reflection information can be information in any form. In some embodiments, the reflection information can be information presented in natural language, for example, text information presented in natural language. In some embodiments, the reflection information can be information presented in feature representation, for example, feature representation information presented in any form that can be understood by the model, such as a numerical vector, matrix, or higher-dimensional tensor.
[0043] Reflection information can be information for analyzing and evaluating any model inference stage. In some embodiments, the reflection information can be information related to any type such as the correctness of inference steps, the rationality of output results, the transparency of the decision-making process, performance performance, and the model inference process and intermediate results.
[0044] High-level reflection information refers to the cross-task reflection information corresponding to the model and can be used to guide the model to process different specific tasks.
[0045] In some embodiments, the high-level reflection information can be information presented in natural language and / or feature representation.
[0046] In some embodiments, the high-level reflection information can be the empirical information manually summarized by the staff based on the inference performance of the model in processing one or more tasks.
[0047] In some embodiments, the high-level reflection information can be the summary empirical information generated by the model based on its own inference process in processing one or more tasks.
[0048] In some embodiments, the high-level reflection information may be reflection information updated according to a preset time period. For example, according to a preset time period, the inference tasks completed within that time period are analyzed and evaluated manually or by a model, so as to update the high-level reflection information.
[0049] In some embodiments, during the execution of a specific task, the model may analyze and evaluate each round of model inference, and update the original high-level reflection information according to the analysis and evaluation results to obtain the updated high-level reflection information.
[0050] It can be seen that the high-level reflection information can provide multiple reflection perspectives and experiences for different historical problems to better guide the generation of specific reflection information.
[0051] The specific reflection information refers to the reflection information generated by the model for the complete inference process of the task and / or one round of model inference when executing a specific task. In the embodiments of the present disclosure, the specific reflection information is the reflection information generated by the target model for one round of model inference when executing the target task. For example, the specific reflection information may be "In this round of inference, I did not get the answer because... In the next round of inference, I will first search for..." etc.
[0052] In some embodiments, the specific reflection information may be information presented using natural language and / or feature representations. In some embodiments, the specific reflection information may be hidden state information generated during the model inference process.
[0053] The specific reflection information corresponding to the previous round of model inference is generated based on the process information and high-level reflection information of the previous round of model inference. That is, under the guidance of the high-level reflection information, the target model analyzes and evaluates the process information of the previous round of model inference to obtain the specific reflection information corresponding to the previous round of model inference to guide the current round of model inference process.
[0054] In some embodiments, the specific reflection information corresponding to the previous round of model inference may be used as part of the prompt information for the current round of model inference, so as to determine the prompt information for the current round of model inference based on the specific reflection information corresponding to the previous round of model inference.
[0055] In some example embodiments, the prompt information for the previous round of model inference may be adjusted under the guidance of the specific reflection information corresponding to the previous round of model inference to obtain the prompt information for the current round of model inference. For example, when the specific reflection information corresponding to the previous round of model inference is "Adjust the model inference steps", the inference steps in the inference example of the previous round of prompt information may be appropriately adjusted to obtain the prompt information for the current round of model inference.
[0056] Step S102: Based on the prompt information of the current round of model inference, use the target model to perform the current round of model inference to obtain the inference result of the current round of model inference.
[0057] Here, after determining the prompt information of the current round of model inference, input the prompt information into the target model; based on this prompt information, the target model performs model inference and obtains the inference result of the current round of model inference.
[0058] In some embodiments, after executing the specified maximum number of inference steps, if the target model generates an inference result for the target task, the target model outputs an identification information indicating that the inference is completed. In some embodiments, this identification information indicating that the inference is completed can be the inference result presented in natural language output by the target model. In some embodiments, this identification information indicating that the inference is completed can be a keyword used to identify the completion of the inference in the inference trajectory of the target model, such as "finished".
[0059] In some embodiments, after executing the specified maximum number of inference steps, if the target model does not obtain an inference result for the target task, the target model outputs an identification information indicating that the inference is not completed. In some embodiments, this identification information indicating that the inference is not completed can be an error message presented in natural language output by the target model. In some embodiments, this identification information indicating that the inference is not completed can be a keyword used to identify the non - completion of the inference in the inference trajectory of the target model, such as "unfinished".
[0060] During implementation, after obtaining the inference result of the current round of model inference, the inference result of the current round of model inference can be stored in an inference result list, so as to determine the target inference result based on at least one inference result in the inference result list.
[0061] Step S103: Determine the target inference result based on the inference results of at least two rounds of model inference.
[0062] Here, after performing at least two rounds of model inference, determine the target inference result of the target task based on the obtained inference results.
[0063] Since the model inference may be in a state where the inference is not completed, therefore, the inference results corresponding to at least two rounds of model inference may be at least two inference results, or there may be only one inference result. In this way, when the inference result of at least two rounds of model inference is only one, use this one inference result as the target inference result; when the inference results of at least two rounds of model inference are at least two, the inference result corresponding to the last round of model inference can be used as the target inference result, and the inference result with the most occurrences among at least two inference results can also be used as the target inference result.
[0064] In the data processing method provided by the present disclosure, first, based on the specific reflection information corresponding to the previous round of model inference, the prompt information for the current round of model inference is determined, where the specific reflection information corresponding to the previous round of model inference is generated based on the process information and the high-level reflection information of the previous round of model inference, and the high-level reflection information represents the cross-task reflection information generated by the target model; then, based on the prompt information for the current round of model inference, the target model is used to perform the current round of model inference to obtain the inference result of the current round of model inference; finally, based on the inference results of at least two rounds of model inference, the target inference result is determined. In this way, on the one hand, the cross-task high-level reflection information is used to guide the generation of the specific reflection information corresponding to the previous round of model inference, and based on this specific reflection information, the prompt information for this round of model inference is determined, realizing the use of cross-task and summary experience information to guide the current model inference, making it easier for the target model to determine the correct model inference steps; on the other hand, the cross-task high-level reflection information is used to guide the model to generate specific reflection information. Compared with the solution in the related art that uses real labels to guide the model for self-correction, the present disclosure has higher feasibility in actual application scenarios.
[0065] In some embodiments, the data processing method provided by the present disclosure further includes steps S104 to S105:
[0066] Step S104, based on the inference result of the current round of model inference and the inference results of at least one previous round of model inference executed, determine the pseudo-feedback information corresponding to the current round of model inference.
[0067] Here, the pseudo-feedback information is determined based on the inference result of the current round of model inference and the inference results of at least one previous round of model inference executed, that is, determined based on the current inference result and at least one previous inference result.
[0068] The pseudo-feedback information can be information in any form. In some embodiments, the pseudo-feedback information can be information presented in natural language, for example, text information presented in natural language. In some embodiments, the pseudo-feedback information can be information presented in feature representation, for example, it can be information presented in any form that the target model can understand, such as a numerical vector, a matrix, or a higher-dimensional tensor. In some embodiments, the pseudo-feedback information can be the hidden state information generated by the target model during multiple rounds of model inference for the target task.
[0069] The pseudo-feedback information can indicate the relationship between the current inference result and at least one prior inference result in any suitable manner. In some embodiments, the pseudo-feedback information can be used to indicate whether the current inference result is the same as at least one prior inference result. For example, the pseudo-feedback information can be used to indicate that the current inference result is the same as all, some, or none of the prior inference results. In some embodiments, the pseudo-feedback information can be used to indicate the semantic similarity between the current inference result and at least one prior inference result. For example, algorithms such as Euclidean distance and cosine similarity can be used to determine the semantic similarity between the current inference result and at least one prior inference result.
[0070] Step S105: Generate specific reflection information corresponding to the current round of model inference based on the process information of the current round of model inference, the pseudo-feedback information corresponding to the current round of model inference, and the high-level reflection information.
[0071] Here, under the guidance of the high-level reflection information, referring to the pseudo-feedback information of the current round of model inference, the process information of the current round of model inference is reflected to obtain the specific reflection information corresponding to the current round of model inference.
[0072] In some embodiments, a target model or a human is used to generate specific reflection information corresponding to the current round of model inference based on the high-level reflection information, the pseudo-feedback information corresponding to the current round of model inference, and the process information.
[0073] In some embodiments, specific reflection information corresponding to the current round of model inference can be generated based on the process information of the current round of model inference, the process information of at least one prior round of model inference executed earlier, the high-level reflection information, and the pseudo-feedback information.
[0074] In the embodiments of the present disclosure, on the one hand, the pseudo-feedback information is used to assist the target model in judging whether the process information of the current round of model inference is correct or reasonable from the perspective of the inference result, so as to determine more reasonable specific reflection information; on the other hand, since the pseudo-feedback information is determined based on the inference result of the current round of model inference and at least one prior inference result, compared with the scheme of using real labels to guide self-correction in the related art, the present disclosure has higher feasibility in actual application scenarios.
[0075] In some embodiments, determining the pseudo-feedback information corresponding to the current round of model inference based on the inference result of the current round of model inference and the inference results of at least one prior round of model inference executed earlier, that is, the above step S104, can be implemented as the following step S1041:
[0076] Step S1041, determine the pseudo-feedback information based on whether the inference result of the current round of model inference matches the inference results of at least one round of model inference executed previously.
[0077] Here, whether the inference result of the current round of model inference matches the inference results of at least one round of model inference executed previously means whether the current inference result matches at least one previous inference result.
[0078] In some embodiments, whether the current inference result matches at least one previous inference result means whether the current inference result is the same as or consistent with at least one previous inference result. For example, when the inference result is a numerical value, it is determined whether they match by judging whether the numerical value of the current inference result is the same as the numerical values of at least one previous inference result.
[0079] In some embodiments, whether the current inference result matches at least one previous inference result means whether the current inference result is close to at least one previous inference result. For example, when the inference result is a numerical value, it is determined whether they match by judging whether the numerical value of the current inference result is the same as the numerical values of at least one previous inference result or whether the absolute value of the difference is within a specified threshold range; or for example, when the inference result is text, it is determined whether they match by judging whether the semantic information of the current inference result is close to the semantic information of at least one previous inference result.
[0080] In this way, when the current inference result matches at least one previous inference result, the pseudo-feedback information can be, for example, "The answers are the same. It is possible that the answer is correct, but it is also possible that you guessed wrong due to overconfidence."; when the current inference result does not match at least one previous inference result, the pseudo-feedback information can be, for example, "The answers are inconsistent. There must be an error in one of the inferences. Please reflect on the error and give a solution." In this way, through the pseudo-feedback information, the target model can reflect on the process information of the current round of model inference and / or the process information of at least one round of model inference executed previously to determine a more reasonable and accurate model inference process.
[0081] In some embodiments, the data processing method provided by the present disclosure further includes the following step S106:
[0082] Step S106, for each round of model inference, after generating the specific reflection information corresponding to the model inference, update the high-level reflection information by using the specific reflection information corresponding to the model inference.
[0083] Here, for each round of model inference, after generating the specific reflection information corresponding to this round of model inference based on the process information, high-level reflection information, and / or pseudo-feedback information of the model inference, the high-level reflection information is updated based on this specific reflection information. In this way, the updated high-level reflection information can be used to guide the generation of the specific reflection information corresponding to the next round of model inference.
[0084] In implementation, the reflection information that can be used for other model inference tasks in the specific reflection information corresponding to this round of model inference can be determined as general reflection information, and the high-level reflection information is updated using this general reflection information.
[0085] In some implementation manners, after determining the general reflection information in the specific reflection information corresponding to this round of model inference, any type of abstract processing algorithm can be used to determine the abstract information of this general reflection information, and the abstract information is added to the high-level reflection information, so as to update the high-level reflection information. In implementation, the specified abstract processing algorithm can include but is not limited to extraction-based abstract algorithms (for example, Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, TextRank algorithm, etc.), and can also be a generation-based abstract algorithm (for example, Sequence-to-Sequence (Seq2Seq) algorithm, GPT series models, etc.).
[0086] In some implementation manners, the general reflection information of the specific reflection information corresponding to this round of model inference can be determined manually, the abstract information of this general reflection information can be determined, and the abstract information is added to the high-level reflection information, so as to update the high-level reflection information.
[0087] In some implementation manners, the step of updating the high-level reflection information using the specific reflection information corresponding to the model inference in step S106 can be implemented as the following step S1061:
[0088] Step S1061: Generate abstract information based on the specific reflection information corresponding to the model inference and the high-level reflection information, and use the abstract information as the updated high-level reflection information.
[0089] Here, abstract processing is performed on the specific reflection information corresponding to the model inference and the high-level reflection information to obtain abstract information, and the abstract information is used as the updated high-level reflection information. In some implementation manners, a specified abstract processing algorithm can be used to generate abstract information based on the specific reflection information corresponding to this round of model inference and the high-level reflection information. In some implementation manners, the abstract information of the specific reflection information corresponding to this round of model inference and the high-level reflection information can be determined manually.
[0090] For example, when the high-level reflection information is "Ensure the consistency between questions and answers", and the specific reflection information is "There are flaws in my reasoning process in this attempt... This error highlights the importance of verifying information and the fact that assumptions should not be made based on only partial connections. To improve, I need to conduct more targeted searches and ensure the accuracy of information before drawing conclusions", performing summary processing on this high-level reflection information and specific reflection information can obtain summary information, that is, the updated high-level reflection information is "Ensure the consistency between questions and answers, and do not make assumptions based on only partial connections".
[0091] For another example, when the high-level reflection information is "Ensure the consistency between questions and answers, and do not make assumptions based on only partial connections", and the specific reflection information is "My reasoning process in this attempt is accurate... This attempt proves the importance of thorough research and logical reasoning for obtaining the correct answer", performing summary processing on this high-level reflection information and specific reflection information can obtain summary information, that is, the updated high-level reflection information is "Ensure the consistency between questions and answers, and do not make assumptions based on only partial connections. Thorough research and logical reasoning are necessary".
[0092] In the embodiments provided by the present disclosure, after each round of model reasoning, by using the specific reflection information corresponding to this round of model reasoning to update the high-level reflection information, the high-level reflection information is continuously updated as the number of model reasoning times performed by the target model increases, becoming dynamically evolving, cross-task, and summary reflection information, thereby providing multi-angle guidance for generating specific reflections, and enabling the target model to more easily obtain accurate specific reflections and reasoning processes.
[0093] In some embodiments, determining the prompt information for the current round of model reasoning based on the specific reflection information corresponding to the previous round of model reasoning, that is, the above step S101, can be implemented as the following step S1011:
[0094] Step S1011: Update the prompt information for the previous round of model reasoning by using the specific reflection information and process information corresponding to the previous round of model reasoning, and use the updated prompt information as the prompt information for the current round of model reasoning.
[0095] In some embodiments, when determining the prompt information for the current round of model reasoning, update the specific reflection information and reasoning process information corresponding to the previous round of model reasoning into the prompt information for the previous round of model reasoning, and use the updated prompt information for the previous round as the prompt information for the current round of model reasoning.
[0096] In some embodiments, when determining the prompt information for the current round of model inference, the process information of the previous round of model inference is updated into the prompt information of the previous round of model inference. At the same time, the specific reflection information of the previous round of model inference is used to adjust the prompt information of the previous round (for example, adjusting the inference steps, intermediate result generation methods, inference result generation methods, etc. defined in the prompt information of the previous round), so as to obtain the adjusted prompt information of the previous round, and the adjusted prompt information of the previous round is used as the prompt information for the current round of model inference.
[0097] In this way, when performing the current round of model inference, the target model can adjust the inference process of this round of model inference based on the specific reflection information of the previous round of model inference in the prompt information, predict the inference results of the target task from different perspectives, or predict the inference results of the target task based on information such as the adjusted inference steps and inference result generation methods; in addition, after completing this round of model inference, the target model can, under the guidance of high-level reflection, generate the specific reflection information corresponding to this round of model inference based on the process information of this round of model inference and the process information of the previous round of model inference in the prompt information.
[0098] In some embodiments, the prompt information for the current round of model inference may include the process information of a specified number of previously executed model inferences. In this way, the target model can reflect on the process information of multiple rounds of model inference under the designation of high-level reflection, so that the specific reflection information is more accurate.
[0099] In some embodiments, when the target task corresponding to the target inference result is the first inference task executed by the target model, the data processing method provided by the present disclosure further includes one of the following steps S107 and step S108:
[0100] Step S107, when the inference result of the first round of model inference of the target model for the target task represents that the inference is completed, using the first initial information as the high-level reflection information; the first initial information is used to instruct the target model to maintain the consistency between the inference result and the target task.
[0101] Here, the high-level reflection information is continuously evolving and updated reflection information. When the target task is the first inference task executed by the target model, it is necessary to initialize the high-level reflection information.
[0102] As described above, the completion of inference means that in a round of model inference, after the target model performs model inference for a specified number of steps, an inference result of the target task is generated, that is, the inference result of this round of model inference. In some embodiments, when the target model outputs the inference result of the first round of model inference, or there is a keyword representing the completion of inference in the inference trajectory, it is considered that the first round of model inference of the target model for the target task is in a state of completed inference.
[0103] In this way, when the first-round inference is in a completed state, the high-level reflection information is initialized as the first initial information for indicating that the target model maintains the consistency between the inference result and the target task, thereby prompting the target model to maintain the accuracy and appropriateness of the inference result.
[0104] Step S108, when the inference result of the first-round model inference of the target model for the target task indicates that the inference is not completed, use the second initial information as the high-level reflection information; the second initial information is used to instruct the target model to avoid the same inference process.
[0105] As described above, the inference not being completed means that in one round of model inference, the target model does not generate an inference result for the target task after executing the maximum number of inference steps. In some embodiments, if the target model does not output the inference result of the first-round model inference, or there is a keyword indicating that the inference is not completed in the inference trajectory, it is considered that the first-round model inference of the target model for the target task is in an uncompleted state.
[0106] In this way, when the first-round model inference is in an uncompleted state, the high-level reflection information is initialized as the second initial information for instructing the target model to avoid the same inference process, so as to instruct the target model to adjust the inference process in the case of incomplete inference and attempt to execute the next-round model inference using different inference strategies, thereby increasing the possibility that the target model obtains the correct inference result.
[0107] In some embodiments, determining the target inference result based on the inference results of at least two rounds of model inference, that is, the above step S103, can be implemented as the following step S1031:
[0108] Step S1031, determine the target inference result from the inference results of the at least two rounds of model inference based on the relevance of the inference results of the at least two rounds of model inference.
[0109] Here, when only one inference result is generated in at least two rounds of model inference (that is, only one round of model inference is in a completed state), determine this one inference result as the target inference result; when at least two inference results are generated in at least two rounds of model inference, determine the target inference result based on the relevance of the inference results of the at least two rounds of model inference, that is, based on the relevance of the at least two inference results.
[0110] In some embodiments, the relevance of at least two inference results may be whether the at least two inference results are the same. That is, if the at least two inference results are the same, the relevance is considered high; if the at least two inference results are different, the relevance is considered low. In this way, the target inference result is determined based on the relevance of at least two inference results, that is, the same inference results among the at least two inference results are determined as the target inference result. For example, in the case where the at least two inference results are result A, result B, and result C, if result A is the same as result B, and both result A and result B are different from result C, then result A and result B are used as the target inference results.
[0111] In some embodiments, the relevance of at least two inference results may be the semantic similarity of the at least two inference results. That is, the semantic similarity of the at least two inference results is used as the relevance of the at least two inference results. In this way, the target inference result is determined based on the relevance of at least two inference results, that is, the target inference result is determined based on at least two inference results with high similarity. For example, in the case where the at least two inference results are result D, result E, and result F, if the semantic similarity between result D and result E is relatively high, and the semantic similarity between both result D and result E and result F is relatively low, then the target inference result is determined based on the semantic information of result D and result E. For example, abstract processing is performed on result D and result E, or the average value of result D and result E is calculated to obtain the target inference result.
[0112] In some embodiments, in the case where the relevance of any two of the at least two inference results is relatively low, the inference result of the last round of model inference is used as the target inference result.
[0113] Next, in combination with Figure 2 , an embodiment of the data processing method provided according to the present disclosure will be described. As Figure 2 shown, this embodiment includes the following steps S201 to step S214; among them, steps S202 to S205 shown in the dashed box 220 are the first-round model inference for the user question, steps S206 to S209 shown in the dashed box 230 are the second-round model inference for the user question, and steps S210 to S213 shown in the dashed box 240 are the third-round model inference for the user question:
[0114] Step S201, obtain the user question; then, execute step S202;
[0115] Step S202, generate a first prompt word based on the user question; then, execute step S203;
[0116] Here, the first prompt word contains the user's question, as well as at least one reasoning example and an explanation of the keywords in the reasoning example. Among them, the reasoning example is a reasoning example based on the ReAct framework of Synergizing Reasoning and Acting in Language Models and executed on the HotPotQA dataset, as shown below:
[0117] "Thought 1: I need to search for information related to the Livsihal War Memorial...
[0118] Action 1: Search [Livsihal War Memorial]
[0119] Observation 1: Unable to find [Livsihal War Memorial]. Similar information: ['Danny Glover filmography'...
[0120] Thought 2: I need to adjust my search... Try searching for "Livsihal War Memorial address"...
[0121] Action 2: Search [Livsihal War Memorial address]
[0122] Observation 2: Unable to find [Livsihal War Memorial address]. Similar information: ['James Stewart'],...
[0123] Thought 3: I need to try a different approach... Try searching for "Livsihal War Memorial history"...
[0124] Action 3: Search [Livsihal War Memorial history]
[0125] Observation 3: Unable to find [Livsihal War Memorial history]. Similar information: ['James Stewart'],...
[0126] Thought 4: I am facing challenges in my search... Obtain more detailed information,... This will help me overcome the lack of specific search results and provide a more accurate answer.
[0127] Action 4: Search [War memorials commemorating over 60 million casualties]
[0128] Observation 4: World War II is the deadliest military conflict in human history...
[0129] Thought 5: Based on the information collected, World War II had over 60 million casualties...
[0130] Action 5: Complete [World War II]".
[0131] Among them, the description of the keywords in the inference example in the prompt can include, for example, the descriptions of "thinking", "action", "observation", and "completion" in the above-mentioned inference example.
[0132] Step S203, based on the first prompt, perform model inference to obtain a first inference result and first process information; then, perform step S204;
[0133] Step S204, store the first inference result in the answer list; generate first specific reflection information based on the first process information and advanced reflection information; then, perform step S205 and step S206;
[0134] Step S205, update the advanced reflection information based on the first specific reflection information; then, perform step S208;
[0135] Here, perform summary processing on the first specific reflection information and the advanced reflection information to obtain summary information, and use this summary information as the updated advanced reflection information.
[0136] Step S206, update the first prompt with the first process information and the first specific reflection information to obtain a second prompt; then, perform step S207;
[0137] Step S207, based on the second prompt, perform model inference to obtain a second inference result and second process information; then, perform step S208;
[0138] Step S208, store the second inference result in the answer list; generate first pseudo-feedback information based on the consistency between the second inference result and the first inference result; generate second specific reflection information based on the first process information, the second process information, the first pseudo-feedback information, and the updated advanced reflection information; then, perform step S209 and step S210;
[0139] Here, when the second inference result is the same as the first inference result, the first pseudo-feedback information is "The answers are consistent. The possible answer may be correct, but it is also possible that you guessed the answer wrong due to overconfidence"; when the second inference result is different from the first inference result, the first pseudo-feedback information is "The answers are inconsistent. There must be an error in one of the inferences. Please reflect on the error and give a solution"; then, under the guidance of the advanced reflection information updated in step S205, generate second specific reflection information based on the first process information, the second process information, and the first pseudo-feedback information.
[0140] Step S209, update the advanced reflection information based on the second specific reflection information; then, perform step S212;
[0141] Here, summary processing is performed on the second specific reflection information and the updated high-level reflection information based on step S205 to obtain summary information, and this summary information is used as the high-level reflection information after being updated again.
[0142] Step S210: Update the second process information and the second specific reflection information to the second prompt word to obtain a third prompt word; then, execute step S211;
[0143] Step S211: Based on the third prompt word, perform model reasoning to obtain a third reasoning result and a third process information; then, execute step S212;
[0144] Step S212: Store the third reasoning result in the answer list; generate second pseudo-feedback information based on the consistency between the third reasoning result and the first and second reasoning results; generate third specific reflection information based on the first process information, the second process information, the third process information, the second pseudo-feedback information, and the updated high-level reflection information; then, execute step S213 and step S214;
[0145] Here, when the third reasoning result is the same as the first and / or second reasoning results, the first pseudo-feedback information is "The answers are consistent. The possible answer may be correct, but it is also possible that you guessed the answer wrong due to overconfidence"; when the third reasoning result is different from both the first and second reasoning results, the first pseudo-feedback information is "The answers are inconsistent. There must be an error in one of the inferences. Please reflect on the error and provide a solution"; then, under the guidance of the high-level reflection information updated based on step S209, generate third specific reflection information based on the first process information, the second process information, the third process information, and the second pseudo-feedback information.
[0146] Step S213: Update the high-level reflection information based on the third specific reflection information;
[0147] Here, summary processing is performed on the third specific reflection information and the high-level reflection information updated based on step S209 to obtain summary information, and this summary information is used as the high-level reflection information after being updated again. The high-level reflection information after being updated again is stored in a specified memory so that when the target model receives the next model processing task, the updated high-level reflection information is used to guide the generation of specific reflection information.
[0148] Step S214: Generate the target answer to the user's question based on the consistency of the first, second, and third reasoning results in the answer list.
[0149] Here, when the first inference result, the second inference result, and the third inference result are all the same, or when two of the inference results are the same, the same inference result is used as the target answer; when the first inference result, the second inference result, and the third inference result are all different, the third inference result is used as the target answer.
[0150] Next, in conjunction with Figure 3 , an embodiment of dynamically updating the high-level reflection information according to the data processing method provided by the present disclosure will be described. As Figure 3 shown, this embodiment includes the following steps S301 to step S306:
[0151] Step S301, obtain user question A; then, execute step S302;
[0152] Step S302, the target model iteratively executes multiple rounds of model inferences for question A, and in each round of model inference, under the guidance of the high-level reflection information, generate the specific reflection information for this round of model inference, and use this specific reflection information to update the high-level reflection information; then, execute step S303;
[0153] Here, in the multiple rounds of iterative model inferences of the target model for question A, the high-level reflection information is iteratively updated.
[0154] Step S303, determine and store the high-level reflection information updated based on the inference process of the target model for question A; then, execute step S304;
[0155] Here, after the target model finishes the last round of model inference for question A, the updated high-level reflection information corresponding to the last round of model inference is used as the high-level reflection information updated based on the inference process of the target model for question A.
[0156] Step S304, obtain user question B; then, execute step S305;
[0157] Step S305, the target model iteratively executes multiple rounds of model inferences for question B, and in each round of model inference, under the guidance of the updated high-level reflection information, generate the specific reflection information for this round of model inference, and use this specific reflection information to update the high-level reflection information; then, execute step S306;
[0158] Step S306, determine and store the high-level reflection information updated based on the inference process of the target model for question B.
[0159] As can be seen from the above, in the data processing method provided by the present disclosure, by dynamically updating the high-level reflection information in different problem environments, the high-level reflection information becomes dynamically evolving, cross-problem, and summary reflection information, which can provide multi-angle reflections and experiences when guiding the generation of specific reflection information, making the specific reflection information more reasonable and accurate.
[0160] To verify the effectiveness of the data processing method provided by the present disclosure, corresponding test cases are designed in the present disclosure.
[0161] First, prepare a test data set; among them, the test data set is the test questions and standard answers selected from HotPotQA;
[0162] Then, take GPT-3.5 and GPT-4 as target models respectively, complete single inference according to the inference examples in the ReAct framework, and design pseudo-feedback information and high-level reflection information for GPT-3.5 and GPT-4 respectively according to the implementation methods of the pseudo-feedback information and high-level reflection information described above, so as to obtain the improved GPT-3.5 and GPT-4;
[0163] Finally, according to the model inference process described above, use the improved GPT-3.5 and GPT-4 to infer the test questions in the test data set respectively to obtain the inference results corresponding to each test question; among them, for each test question, GPT-3.5 and GPT-4 both iteratively execute 3 rounds of model inference. In this way, each test question corresponds to 3 inference results.
[0164] After the above tests, compare the 3 inference results generated by the improved GPT-3.5 and GPT-4 for each test question with the standard answers respectively to obtain the accuracy rates of the inference results of each round of model inference, as shown in Table 1 below:
[0165] Model Name Accuracy of the First Round of Inference Results Accuracy of the Second Round of Inference Results Accuracy of the Third Round of Inference Results GPT-3.5 0.35 0.38 0.41 GPT-4 0.43 0.49 0.49
[0166] Table 1: List of accuracy rates of inference results
[0167] As can be seen from Table 1, under the standard prompt words, the accuracy rates of the inference results of the first round of model inference obtained by inferring using the ReAct inference examples are 0.35 and 0.43 respectively; on the basis of the first round of model inference, giving play to the role of the continuously evolving high-level reflection information, the accuracy rates of the inference results of the second round are 0.38 and 0.49 respectively; on the basis of the first two rounds of model inference, jointly giving play to the roles of the pseudo-feedback information and the high-level reflection information, the accuracy rates of the inference results of the third round are 0.41 and 0.49 respectively. It can be seen that under the joint action of the high-level reflection information and the pseudo-feedback information, the accuracy rate of the inference results of the third round is improved compared with that of the first round of inference results, which indicates the effectiveness of the data processing method provided by the present disclosure.
[0168] Based on the foregoing embodiments, the present disclosure provides a data processing device, which includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the process of implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0169] Figure 4 It is a schematic diagram of the composition structure of a data processing device provided by the present disclosure, as Figure 4 shown, the data processing device 400 includes: a first determination module 410, an inference module 420, and a second determination module 430, where:
[0170] The first determination module 410 determines the prompt information for the current round of model inference based on the specific reflection information corresponding to the previous round of model inference; the specific reflection information corresponding to the previous round of model inference is generated based on the process information and the high-level reflection information of the previous round of model inference; the high-level reflection information represents the cross-task reflection information generated by the target model.
[0171] The inference module 420 uses the target model to perform the current round of model inference based on the prompt information for the current round of model inference, and obtains the inference result of the current round of model inference.
[0172] The second determination module 430 determines the target inference result based on the inference results of at least two rounds of model inference.
[0173] In some embodiments, the device 400 further includes:
[0174] A pseudo-feedback information module, configured to determine the pseudo-feedback information corresponding to the current round of model inference based on the inference result of the current round of model inference and the inference results of at least one previous round of model inference performed earlier;
[0175] A specific reflection information module, configured to generate the specific reflection information corresponding to the current round of model inference based on the process information of the current round of model inference, the pseudo-feedback information corresponding to the current round of model inference, and the high-level reflection information.
[0176] In some embodiments, the pseudo-feedback information module is configured to determine the pseudo-feedback information based on whether the inference result of the current round of model inference matches the inference results of at least one previous round of model inference that has been executed.
[0177] In some embodiments, the apparatus 400 further includes:
[0178] An update module, configured to, for each round of model inference, after generating the specific reflection information corresponding to the model inference, use the specific reflection information corresponding to the model inference to update the high-level reflection information.
[0179] In some embodiments, the update module is configured to generate summary information based on the specific reflection information corresponding to the model inference and the high-level reflection information, and use the summary information as the updated high-level reflection information.
[0180] In some embodiments, the first determination module 410 is configured to update the hint information of the previous round of model inference by using the specific reflection information and process information corresponding to the previous round of model inference, and use the updated hint information as the hint information of the current round of model inference.
[0181] In some embodiments, when the target task corresponding to the target inference result is the first inference task executed by the target model, the apparatus 400 further includes a high-level reflection information module; the high-level reflection information module is configured to perform one of the following:
[0182] When the inference result of the first round of model inference of the target model for the target task represents that the inference is completed, use the first initial information as the high-level reflection information; the first initial information is used to instruct the target model to maintain the consistency between the inference result and the target task;
[0183] When the inference result of the first round of model inference of the target model for the target task represents that the inference is not completed, use the second initial information as the high-level reflection information; the second initial information is used to instruct the target model to avoid the same inference process.
[0184] In some embodiments, the second determination module 430 is configured to determine the target processing result from the inference results of the at least two rounds of model inference based on the relevance of the inference results of the at least two rounds of model inference.
[0185] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.
[0186] It should be noted that in the embodiments of the present disclosure, if the above data processing method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0187] The embodiments of the present disclosure provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0188] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.
[0189] The embodiments of the present disclosure provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.
[0190] An embodiment of the present disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0191] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.
[0192] It should be noted that Figure 5 is a schematic diagram of a hardware entity of an electronic device in the present disclosure. As Figure 5 shown, the hardware entity of the electronic device 500 includes: a processor 501, a communication interface 502, and a memory 503, where:
[0193] The processor 501 generally controls the overall operation of the electronic device 500.
[0194] The communication interface 502 can enable the electronic device to communicate with other terminals or servers through a network.
[0195] The memory 503 is configured to store instructions and applications executable by the processor 501, and can also cache data to be processed or already processed by the processor 501 and each module in the electronic device 500 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be performed between the processor 501, the communication interface 502, and the memory 503 through a bus 504.
[0196] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the magnitude of the serial numbers of the above steps / processes does not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0197] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0198] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0199] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0200] In addition, each functional unit in the embodiments of the present disclosure can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.
[0201] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks and other various media that can store program codes.
[0202] Alternatively, if the above integrated unit is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks and other various media that can store program codes.
[0203] As described above, the above are only the implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure.
Claims
1. A data processing method, comprising: Based on the specific reflection information corresponding to the previous round of model reasoning, determine the prompt information for the current round of model reasoning; The specific reflection information corresponding to the previous round of model reasoning is generated based on the process information and high-level reflection information of the previous round of model reasoning; the high-level reflection information represents the cross-task reflection information generated by the target model; Based on the prompt information of the current round model reasoning, using the target model to perform the current round model reasoning to obtain the reasoning result of the current round model reasoning; Based on the inference results of at least two rounds of model inference, a target inference result is determined.
2. The method according to claim 1, further comprising: Determine pseudo feedback information corresponding to the current round of model reasoning based on the reasoning result of the current round of model reasoning and the reasoning result of at least one round of model reasoning previously performed; Based on the process information of the current round of model reasoning, the pseudo feedback information corresponding to the current round of model reasoning and the high-level reflection information, the specific reflection information corresponding to the current round of model reasoning is generated.
3. The method according to claim 2, wherein determining the pseudo feedback information corresponding to the current round of model reasoning based on the reasoning result of the current round of model reasoning and the reasoning result of at least one round of model reasoning previously performed comprises: The pseudo feedback information is determined based on whether the reasoning result of the current round of model reasoning matches the reasoning result of at least one round of model reasoning previously performed.
4. The method according to any one of claims 1 to 3, further comprising: For each round of model reasoning, after generating specific reflection information corresponding to the model reasoning, the high-level reflection information is updated using the specific reflection information corresponding to the model reasoning.
5. The method according to claim 4, wherein the step of updating the high-level reflection information by using the specific reflection information corresponding to the model reasoning comprises: Based on the specific reflection information corresponding to the model reasoning and the high-level reflection information, summary information is generated, and the summary information is used as the updated high-level reflection information.
6. The method according to any one of claims 1 to 3, wherein determining the prompt information for the current round of model reasoning based on the specific reflection information corresponding to the previous round of model reasoning comprises: The specific reflection information and process information corresponding to the previous round of model reasoning are used to update the prompt information of the previous round of model reasoning, and the updated prompt information is used as the prompt information of the current round of model reasoning.
7. According to the method of any one of claims 1 to 3, when the target task corresponding to the target reasoning result is the first reasoning task performed by the target model, the method further comprises one of the following: In the case where the reasoning result of the first round of model reasoning of the target model for the target task represents that the reasoning is completed, the first initial information is used as the high-level reflection information; the first initial information is used to instruct the target model to maintain consistency between the reasoning result and the target task; In a case where a reasoning result of a first round of model reasoning of the target model for the target task indicates that reasoning is incomplete, using the second initial information as the high-level reflection information; The second initial information is used to instruct the target model to avoid the same reasoning process.
8. The method according to any one of claims 1 to 3, wherein: The determining of the target reasoning result based on the reasoning results of at least two rounds of model reasoning includes: Based on the correlation of the inference results of the at least two rounds of model reasoning, the target inference result is determined from the inference results of the at least two rounds of model reasoning.
9. A data processing device, comprising: The first determination module determines the prompt information of the current round of model reasoning based on the specific reflection information corresponding to the previous round of model reasoning; The specific reflection information corresponding to the previous round of model reasoning is generated based on the process information and high-level reflection information of the previous round of model reasoning; the high-level reflection information represents the cross-task reflection information generated by the target model; A reasoning module, based on the prompt information of the current round model reasoning, uses the target model to perform the current round model reasoning to obtain a reasoning result of the current round model reasoning; The second determination module determines a target reasoning result based on the reasoning results of at least two rounds of model reasoning.
10. A computer program product, comprising a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, the method according to any one of claims 1 to 8 is implemented.