Large model reasoning method and reasoning equipment based on large model
Through internal and external coordinated dynamic reflection strategy and memory bank update, the problem of error recognition and correction in self-reflection of big models is solved, and the inference accuracy and reliability of big models are improved.
Patent Information
- Application Number
- CN202510232702.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-18
AI Technical Summary
During the self-reflection process, it is difficult for big models to accurately identify and correct the wrong answers generated by themselves, and may even misjudge the correct answer as wrong answers, resulting in a decrease in the accuracy and reliability of the output results.
The internal and external collaborative dynamic reflection strategy is adopted, and after obtaining the basic response through large-scale model inference, the target response is dynamically selected based on external evaluation information and advanced information, and the target response is continuously optimized during the iteration process until the iteration termination conditions are met. The memory bank stores and updates advanced ideas to improve reasoning capabilities.
It effectively improves the accuracy and reliability of inference of large models, reduces the possibility of modifying the correct answer to the wrong answer, and realizes continuous improvement and optimization of large model inference.
Smart Images

Figure CN120338094A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a large model inference method and an inference device based on a large model. Background Art
[0002] With its powerful language understanding and generation capabilities, large model technology has been widely applied in many fields. To further improve the inference ability of large models, a self-reflection mechanism is currently commonly used, that is, through the way of "response → evaluation → correction", the large model optimizes the content it generates. However, in the self-reflection process of the large model, it is difficult to accurately identify the wrong answers it generates, resulting in the inability to effectively correct the errors; more seriously, even the originally correct answers may be misjudged as wrong answers during the reflection process, which greatly reduces the accuracy and reliability of the output results of the large model and seriously affects its effect and value in practical applications. Summary of the Invention
[0003] The present disclosure provides a large model inference method and an inference device based on a large model to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, there is provided a large model inference method, the method comprising:
[0005] Inferring a first basic response to problem information through a large model;
[0006] Based on first evaluation information obtained by evaluating the first basic response, inferring a first reflection response corresponding to the problem information;
[0007] Based on advanced information corresponding to the problem information, determining a first target response among the first basic response and the first reflection response; wherein the advanced information represents cross-task inference ideas;
[0008] Determining feedback information corresponding to the problem information based on the first target response.
[0009] In the above solution, the determining the feedback information corresponding to the problem information based on the first target response includes:
[0010] Using the first target response as a second basic response;
[0011] Based on second evaluation information obtained by evaluating the second basic response, inferring a second reflection response corresponding to the problem information;
[0012] Based on the advanced information, determining a second target response among the second basic response and the second reflection response;
[0013] Iteratively execute the above steps until the iteration termination condition is met, and determine the feedback information corresponding to the problem information based on the Nth target response, where N is a positive integer greater than or equal to 2;
[0014] Wherein, the iteration termination condition includes that the number of iterations reaches the number threshold or the second target response meets the quality requirements.
[0015] In the above solution, before determining the first target response in the first basic response and the first reflection response based on the advanced information corresponding to the problem information, the method further includes:
[0016] Obtain the advanced information generated by the target model;
[0017] Wherein, the target model determines approximate advanced information related to the keyword and / or the attribute information from the memory library based on the keyword and / or the attribute information in the problem information;
[0018] The target model generates the advanced information based on the approximate problem information corresponding to the problem information and the approximate advanced information; or, the target model generates the advanced information based on the approximate problem information corresponding to the problem information, the advanced information corresponding to the approximate problem information, and the approximate advanced information.
[0019] In the above solution, the memory library stores the approximate problem information and the approximate advanced information corresponding to the problem information; or, the memory library stores the approximate problem information, the approximate advanced information corresponding to the problem information, and the advanced information corresponding to the approximate problem information;
[0020] The advanced information generated by the target model is used to update the advanced information corresponding to at least one problem stored in the memory library.
[0021] In the above solution, before determining the first target response in the first basic response and the first reflection response based on the advanced information corresponding to the problem information, the method further includes:
[0022] Obtain the advanced information generated by the target model;
[0023] Wherein, the target model determines at least one approximate advanced information corresponding to each keyword from the memory library based on at least one keyword in the problem information;
[0024] The target model determines the advanced information corresponding to the problem information based on multiple pieces of the approximate advanced information having an association relationship.
[0025] In the above solution, at least one approximate advanced information corresponding to the problem information is stored in the memory bank;
[0026] After the target model generates the advanced information, the problem information is stored in the memory bank.
[0027] In the above solution, before the first evaluation information obtained based on the evaluation of the first basic response, the method further includes:
[0028] Based on at least one of whether the first basic response includes a preset character, the rationality of the first basic response, the logic of the first basic response, and whether the first basic response conforms to the facts, the first evaluation information corresponding to the first basic response is determined.
[0029] In the above solution, the first reflection response corresponding to the problem information is inferred from the first evaluation information obtained based on the evaluation of the first basic response, including:
[0030] Based on the problem information and the first evaluation information, the first basic response is updated so that the updated first basic response matches the first evaluation information;
[0031] The updated first basic response is determined to be the first reflection response corresponding to the problem information.
[0032] In the above solution, the updating of the first basic response specifically includes at least one of the following:
[0033] If the first basic response does not include a preset character, based on the semantics of the first basic response, the preset character is added to the first basic response;
[0034] If the first basic response does not meet the rationality, the content of the first basic response is adjusted so that the adjusted first basic response meets the rationality;
[0035] If the first basic response does not meet the logic, the statement order or content of the first basic response is adjusted so that the adjusted first basic response meets the logic;
[0036] If the first basic response does not conform to the facts, the statement order or content of the first basic response is adjusted so that the adjusted first basic response conforms to the facts.
[0037] According to a second aspect of the present disclosure, there is provided an inference device based on a large model, the device includes:
[0038] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.
[0039] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understandable. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, wherein:
[0041] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0042] Figure 1 The first alternative flowchart diagram of the large model inference method provided by the embodiment of the present disclosure is shown;
[0043] Figure 2 The second alternative flowchart diagram of the large model inference method provided by the embodiment of the present disclosure is shown;
[0044] Figure 3 The static self-reflection diagram in the related art is shown;
[0045] Figure 4 The internal and external collaborative dynamic reflection diagram provided by the embodiment of the present disclosure is shown;
[0046] Figure 5 The third alternative flowchart diagram of the large model inference method provided by the embodiment of the present disclosure is shown;
[0047] Figure 6 The fourth alternative flowchart diagram of the large model inference method provided by the embodiment of the present disclosure is shown;
[0048] Figure 7 The fifth alternative flowchart diagram of the large model inference method provided by the embodiment of the present disclosure is shown;
[0049] Figure 8 The large model inference flowchart diagram provided by the embodiment of the present disclosure is shown;
[0050] Figure 9 The iterative output results of different methods on the same data set are shown;
[0051] Figure 10 The figure shows an optional structural schematic diagram of the large model inference device provided by an embodiment of the present disclosure;
[0052] Figure 11 The figure shows a composition structural schematic diagram of an inference device based on a large model according to an embodiment of the present disclosure. Detailed implementation manners
[0053] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present disclosure.
[0054] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0055] In the following description, the terms "first / second" involved are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present disclosure described here can be implemented in an order other than that illustrated or described here.
[0056] Unless otherwise defined, all technical and scientific terms used in the present disclosure have the same meaning as commonly understood by those skilled in the technical field to which the present disclosure belongs. The terms used in the present disclosure are only for the purpose of describing the embodiments of the present disclosure, and are not intended to limit the present disclosure.
[0057] It should be understood that in various embodiments of the present disclosure, the magnitude of the serial numbers of each implementation process does not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0058] Currently, large models use self-reflection (response → evaluation → correction) to improve their inference ability. However, recent research has shown that when large models perform self-reflection, if they only rely on their own inference ability and do not rely on external feedback, there will be problems that large models are difficult to accurately identify and correct the wrong answers they generate, and may even reflect the correct answers as wrong answers.
[0059] Specifically, the related technology uses Self-Correcting with Tool-Interactive Critiquing (CRITIC) to optimize the output results of large models. This is a method that allows the model to perform self-correction, aiming to solve the problem of self-reflection of large models. With the help of the tool-interactive criticism mechanism, the model no longer relies solely on its own internal reasoning for self-reflection, but obtains feedback through interaction with external tools, and makes more accurate judgments and corrections on the content generated by itself, thereby improving the correctness, logic and rationality of the model's output results, effectively avoiding the model misjudging the correct answer as an incorrect answer or failing to accurately identify errors, improving the performance and reliability of the model in various tasks, and enhancing its value in practical applications.
[0060] However, the above method uses oracle labels about the correctness of the answer to guide the self-correction process. For example, the answer output by the large model is compared with the real answer in the data set. If the output answer is consistent with the real answer, the reflection stops, otherwise the reflection continues. However, in practical applications, we cannot obtain oracle labels (we cannot know the answer to the question in advance), so using oracle labels to guide self-reflection cannot adapt to real application scenarios; in addition, due to the overconfidence and high randomness of the large model, the agent may modify the correct answer to the wrong answer during the self-reflection process, or it is difficult to correct the wrong answer to the correct answer. After reflection, the accuracy is reduced; finally, the above scheme uses a static self-reflection strategy (reflecting on the i-1 output result to obtain the i-th result). Experiments show that this method is easily disturbed by wrong answers during iteration: if the large model generates a wrong answer through reflection, the result generated by CRITIC will deviate more and more from the correct answer as the iteration proceeds. Therefore, choosing which output result to reflect on is a key.
[0061] Therefore, based on the problems existing in the related technology, the embodiments of the present disclosure provide a large model reasoning method, which enables the Agent to identify and correct the wrong answers generated by itself during the iterative reflection process, while reducing the possibility of modifying the correct answer to the wrong answer, and efficiently combining internal reflection and external feedback to achieve reasoning optimization in reflection.
[0062] Figure 1 A first optional flow chart of the large model reasoning method provided by an embodiment of the present disclosure is shown and will be explained according to each step.
[0063] Step S101, inferring the question information through a large model to obtain a first basic response.
[0064] In some embodiments, the problem information is input into a large model to obtain a first basic response inferred and output by the large model for the problem information.
[0065] Step S102: Based on the first evaluation information obtained by evaluating the first basic response, infer a first reflection response corresponding to the problem information.
[0066] In some embodiments, a carrier for implementing the large model inference method (hereinafter referred to as the carrier) evaluates the first basic response to obtain first evaluation information. The first evaluation information is used to represent at least one of the correctness, logic, and rationality of the first basic response. The correctness includes whether the first basic response conforms to objective facts, follows professional knowledge and rules, and whether the provided information is accurate without errors, misleading, or confusion. The logic includes the coherence, consistency, self-consistency of the reasoning process, and the absence of logical errors. The rationality includes whether it conforms to common sense and general knowledge, the suitability of the goal and scenario, and the trade-off and comprehensive consideration of various factors.
[0067] In some embodiments, the carrier infers based on the first evaluation information to obtain a first reflection response corresponding to the problem information.
[0068] In some embodiments, the first reflection response and the first basic response may be the same or different; if the first basic response meets the requirements of correctness, logic, and rationality, the first reflection response and the first basic response are the same; if the first basic response does not meet the requirements of correctness, logic, and rationality, the first reflection response and the first basic response are different.
[0069] Among them, the carrier may be a computer program, electronic circuit, database, mobile application, electronic device, cloud computing platform, distributed system, artificial intelligence framework, mathematical model, automation tool, microcontroller, etc., which are software or hardware capable of implementing algorithms and method processes.
[0070] Step S103: Based on the advanced information corresponding to the problem information, determine a first target response among the first basic response and the first reflection response.
[0071] In some embodiments, the carrier obtains the advanced information corresponding to the problem information; the advanced information can be obtained according to other problem information other than the problem information, and is used to represent cross-task reasoning ideas. Specifically, the advanced information can be obtained based on the problem information and response information corresponding to other tasks, and can be the reasoning process or reasoning idea for obtaining the response information according to the other problem information.
[0072] In some embodiments, based on the advanced information, the carrier determines a first target response from the first basic response and the first reflective response. Specifically, among the first basic response and the first reflective response, the response that better matches the advanced information is determined as the first target response.
[0073] The matching may include semantic matching or feature matching in the feature space; the feature matching may be implemented based on the distance between the features corresponding to the response in the feature space and the features corresponding to the advanced information.
[0074] Step S104, determining feedback information corresponding to the problem information based on the first target response.
[0075] In some embodiments, the carrier determines the first target response as the feedback information corresponding to the problem information and outputs the feedback information.
[0076] In this way, through the large model inference method provided by the embodiments of the present disclosure, by evaluating the basic response to obtain evaluation information, and obtaining a reflective response based on the evaluation information to achieve internal self-reflection; obtaining advanced ideas based on the problem information, and determining a target response based on the advanced ideas to achieve external dynamic selection; guiding the large model to perform inference by combining internal self-reflection and external dynamic selection, avoiding guiding the large model to output incorrect results only relying on the self-reflection mechanism, and improving the accuracy and reliability of the large model inference.
[0077] Figure 2 Fig. shows a second alternative flowchart of the large model inference method provided by the embodiments of the present disclosure. Figure 3 Fig. shows a schematic diagram of static self-reflection in the related art. Figure 4 Fig. shows a schematic diagram of internal and external collaborative dynamic reflection provided by the embodiments of the present disclosure, which will be described according to each step.
[0078] As Figure 3 shown, for the static self-reflection strategy (Critic) in the related art, first an initial response is generated based on the problem information and prompt 1, then internal self-reflection is performed based on the initial response and prompt 2 to generate a reflective response; reflection is performed based on the reflective response and prompt 2 to generate a new reflective response until the iteration terminates, and the output result of the large model for the last time is determined as the feedback information corresponding to the problem information.
[0079] However, due to the overconfidence and high randomness of the large model, it is difficult for the Agent to accurately identify and correct wrong answers through self-reflection, and even correct answers may be modified to wrong answers. Based on the static self-reflection framework, the iteration is vulnerable to interference from wrong answers.
[0080] Based on this, as Figure 4As shown in the figure, an internal and external collaborative dynamic reflection strategy (Meta-Critic), that is, a large model reasoning method, is provided in an embodiment of the present disclosure. After the large model generates an initial response based on the problem information and prompt 1, it reflects based on prompt 2 (i.e., the first evaluation information) to generate a reflective response; the reflective response, the initial response, and the advanced idea (i.e., prompt 3) are input into the large model, and based on the advanced idea outside the large model, dynamic selection between the reflective response and the initial response is performed to determine a better response (i.e., the target response). Iteration of internal self-reflection and external dynamic selection is performed again, that is, the better response is reflected to obtain a new reflective response 2. The two obtained reflective responses and the advanced idea are input into the large model, and based on the advanced idea outside the large model, dynamic selection between the reflective response and the initial response is performed to determine a better response until the iteration termination condition is met, and the final reflective response output by the large model is determined as the feedback information corresponding to the problem information. The specific steps are as follows:
[0081] Step S201, the large model is used to reason about the problem information to obtain a first basic response.
[0082] The specific processes of step S201 and step S101 are the same, and will not be repeated here.
[0083] Step S202, based on the first evaluation information obtained by evaluating the first basic response, a first reflective response corresponding to the problem information is reasoned out.
[0084] The specific processes of step S202 and step S102 are the same, and will not be repeated here.
[0085] Step S203, based on the advanced information corresponding to the problem information, a first target response in the first basic response and the first reflective response is determined.
[0086] The specific processes of step S203 and step S103 are the same, and will not be repeated here.
[0087] Step S204, based on the second evaluation information obtained by evaluating the second basic response, a second reflective response corresponding to the problem information is reasoned out.
[0088] In some embodiments, the carrier uses the first target response as the second basic response, evaluates the second basic response, and obtains second evaluation information. The specific evaluation method is the same as that of step S102, and will not be repeated here.
[0089] In some embodiments, the carrier updates the second basic response based on the second evaluation information, and makes the content in the second basic response that does not meet rationality, logic, and correctness meet rationality, logic, and correctness to obtain second reflective information.
[0090] Step S205: Based on the advanced information, determine the second target response among the second basic response and the second reflective response.
[0091] In some embodiments, based on the advanced information involved in Step S203, the carrier determines the second target response from the second basic response and the second reflective response. Specifically, among the second basic response and the second reflective response, the response that better matches the advanced information is determined as the second target response.
[0092] In some embodiments, the carrier uses the second target response as the third basic response, evaluates the third basic response to obtain the third evaluation information, and repeats Step S204 to Step S205 until the iteration termination condition is met. Based on the Nth target response, the feedback information corresponding to the problem information is determined, where N is a positive integer greater than or equal to 2, representing the total number of times of executing Step S203 and Step S205. That is, if only Step S203 and Step S205 are executed, then N is 2. After executing Step S205, if the iteration termination condition is not met and Step S204 and Step S205 are executed again, then N is 3.
[0093] In this way, through the large model reasoning method provided by the embodiments of the present disclosure, by evaluating the basic response to obtain the evaluation information, and obtaining the reflective response based on the evaluation information, internal self-reflection is realized; by obtaining the advanced thought based on the problem information and determining the target response based on the advanced thought, external dynamic selection is realized; by combining internal self-reflection and external dynamic selection to guide the large model to perform reasoning, and repeating the internal self-reflection and external dynamic selection, the reasoning ability of the large model is improved.
[0094] Figure 5 Fig. 3 shows a third optional flowchart of the large model reasoning method provided by the embodiments of the present disclosure, and will be described according to each step.
[0095] Step S301: Use the large model to reason about the problem information to obtain the first basic response.
[0096] The specific step flow of Step S301 is the same as that of Step S101, and will not be repeated here.
[0097] Step S302: Obtain the first evaluation information by evaluating the first basic response.
[0098] In some embodiments, based on at least one of whether the first basic response includes a preset character, the rationality of the first basic response, the logic of the first basic response, and whether the first basic response conforms to the facts, the first evaluation information corresponding to the first basic response is determined.
[0099] In specific implementation, the preset character may include at least one keyword in the problem information, and the carrier determines whether at least one keyword in the problem information is included in the first basic response; based on the evaluation result, relevance evaluation information is determined.
[0100] In specific implementation, the carrier evaluates the rationality of the first basic response, that is, evaluates whether the first basic response conforms to common sense and general knowledge, the suitability of the goal and the scenario, and the trade-off and comprehensive consideration of various factors; based on the evaluation result, rationality evaluation information is determined.
[0101] In specific implementation, the carrier evaluates the logic of the first basic response, that is, evaluates the coherence, consistency, self-consistency of the reasoning process of the first basic response, and whether there are logical errors; based on the evaluation result, logic evaluation information is determined.
[0102] In specific implementation, the carrier evaluates the correctness of the first basic response, that is, evaluates whether the first basic response conforms to objective facts, whether it follows professional knowledge and rules, and whether the provided information is accurate, and whether there are errors, misleading or confusing situations; based on the evaluation result, correctness evaluation information is determined.
[0103] In some embodiments, the carrier determines first evaluation information based on at least one of the relevance evaluation information, rationality evaluation information, logic evaluation information, and correctness evaluation information.
[0104] Step S303, update the first basic response based on the first evaluation information to obtain the first reflective response corresponding to the problem information.
[0105] In some embodiments, the carrier updates the first basic response based on the problem information and the first evaluation information, so that the updated first basic response matches the first evaluation information; determining the updated first basic response as the first reflective response corresponding to the problem information may specifically include at least one of the following: if the preset character is not included in the first basic response, then based on the semantics of the first basic response, add the preset character to the first basic response; if the first basic response does not meet the rationality, then adjust the content of the first basic response so that the adjusted first basic response meets the rationality; if the first basic response does not meet the logic, then adjust the statement order or content of the first basic response so that the adjusted first basic response meets the logic; if the first basic response does not conform to the facts, then adjust the statement order or content of the first basic response so that the adjusted first basic response conforms to the facts.
[0106] Step S304, obtain advanced information corresponding to the problem information based on the approximate problem information corresponding to the problem information.
[0107] In some embodiments, the advanced information is obtained by the target model. The target model is independent of the large model.
[0108] In some embodiments, the target model obtains advanced information based on a memory bank.
[0109] During specific implementation, if the advanced information corresponding to the question information is included in the memory bank, the target model directly obtains the advanced information.
[0110] Alternatively, during specific implementation, if the advanced information corresponding to the question information is not included in the memory bank, approximate advanced information related to the keywords and / or the attribute information in the question information is obtained from the memory bank. Wherein, the attribute information may include the classification to which the question information belongs. The approximate advanced information may be related information of the keyword, such as an explanation of the keyword, or related information similar to the keyword, such as related information of other keywords in the same classification as the keyword; or, the approximate advanced information may be an explanation of other question information in the classification to which the question information belongs.
[0111] For example, if the question information is "How many times can a bee sting a person?", the keywords may include bee, sting, person; further, the approximate advanced information may include an explanation of the bee, such as biological characteristics, living habits, ecological functions, and relationship with humans; it may also include an explanation of "sting", such as the basic meaning, attributes, etc.
[0112] Or, if the question information is "How many times can a bee sting a person", the classification to which the question belongs may include an explanation of "How many times can a wasp sting a person?".
[0113] In some embodiments, after the target model obtains approximate advanced information related to the keywords and / or the attribute information, the advanced information is generated based on approximate question information corresponding to the question information and the approximate advanced information.
[0114] For example, the question information is "How to make the WiFi signal at home stronger?", the approximate question information may include "What to do if the WiFi signal is weak?", "How to improve the network coverage at home?", "Where is the best place to put the router for the best signal?", "What are the methods to enhance the WiFi signal?", and the approximate advanced information may include: WiFi signal: wireless network, router, signal strength, network coverage; stronger: enhance, improve, optimize, upgrade; home: indoor, room layout, wall obstruction, distance. Based on the approximate question information and the approximate advanced information, the advanced idea corresponding to the question information can be obtained.
[0115] In some other embodiments, after the target model obtains approximate advanced information related to the keyword and / or the attribute information, the target model generates the advanced information based on the approximate problem information corresponding to the problem information, the advanced information corresponding to the approximate problem information, and the approximate advanced information.
[0116] For example, if the problem information is "How many times can a bee sting a person?", the approximate problem information is "How many times can a wasp sting a person", and the approximate advanced information of the problem information may include the biological characteristics of bees; the advanced information of the approximate problem information includes "The stinger of a wasp is relatively smooth and will not leave the stinger in the skin of the stung person after stinging, so the wasp can smoothly pull out the stinger after stinging and can then launch another attack to sting again." Then the advanced information may include "Comparing the physiological structures of bees and wasps, it is learned that the stinger of a bee consists of a dorsal stinger and two ventral stingers, the ends of which are connected to venom glands and internal organs, and there are barbs on the stinger, and the stinger will remain in the skin after stinging."
[0117] In some alternative embodiments, after the target model generates the advanced information, it transmits the advanced information to the memory bank; so that the memory bank stores the problem information and the advanced information corresponding to the problem information; and, based on the problem information and the advanced information, performs reasoning, and according to the reasoning result or reasoning logic, updates the advanced information corresponding to at least one problem information stored in the memory bank.
[0118] That is, the memory bank learns the reasoning ability of the target model, and based on the reasoning result of the target model, updates and evolves the advanced information stored in the memory bank to make the advanced information more correct, reasonable, and logical.
[0119] Step S305, based on the advanced information corresponding to the problem information, determine the first target response in the first basic response and the first reflective response.
[0120] In some embodiments, the carrier compares the first basic response and the first reflective response; in response to the first basic response and the first reflective response being the same, confirm any one as the first target response; in response to the first basic response and the first reflective response being different, determine the one that is more matched with the advanced information as the first target response.
[0121] Step S306, based on the first target response, determine the feedback information corresponding to the problem information.
[0122] In some embodiments, step S306 may be the same as step S104, and will not be repeated here.
[0123] In some other embodiments, the carrier may also iterate after obtaining the first target response, that is, repeatedly execute step S302, step S303, and step S305 until the iteration termination condition is met, and determine the last target response output by the large model as the feedback information corresponding to the problem information.
[0124] Thus, through the large model inference method provided by the embodiments of the present disclosure, based on the dynamic reflection of internal and external collaboration, an evolvable memory bank is provided. The memory bank stores the advanced ideas for solving different problem information in the form of <problem, advanced idea> pairs, which may include the knowledge required to solve the problem, the analysis of the problem, and the basic ideas for providing solutions to the problem. These ideas are refined from the problem-solving processes of different tasks. As the interaction with the target model and the task execution progress, the memory bank will update and upgrade the stored advanced ideas, realizing the evolution and improvement of the memory bank; the advanced model will obtain the advanced ideas of approximate problem information from the memory bank for analogical learning, and adaptively generate the advanced ideas for the current problem information, and at the same time add the generated advanced ideas to the memory bank. This process strengthens the abstract reasoning ability of the target model. The target model enables the large language model to rise from the perspective of an answerer to the height of a teacher to externally select and globally examine the basic response and the reflection response, improving the reasoning ability of the large model.
[0125] Figure 6 Fig. 4 shows a fourth optional flowchart of the large model inference method provided by the embodiments of the present disclosure, and will be described according to each step.
[0126] Step S401, the large model infers the problem information to obtain the first basic response.
[0127] The specific step flow of step S401 is the same as that of step S101, and will not be repeated here.
[0128] Step S402, obtain the first evaluation information by evaluating the first basic response.
[0129] The specific step flow of step S402 is the same as that of step S302, and will not be repeated here.
[0130] Step S403, update the first basic response based on the first evaluation information to obtain the first reflection response corresponding to the problem information.
[0131] The specific step flow of step S403 is the same as that of step S303, and will not be repeated here.
[0132] Step S404, obtain the advanced information corresponding to the problem information based on the approximate advanced information corresponding to the problem information.
[0133] In some embodiments, the advanced information is obtained by the target model. The target model is independent of the large model.
[0134] In some embodiments, if the advanced information corresponding to the question information is not included in the memory bank, the target model determines at least one approximate advanced information corresponding to each keyword from the memory bank based on at least one keyword in the question information; the target model determines the advanced information corresponding to the question information based on multiple pieces of the approximate advanced information having an association relationship.
[0135] Specifically, during implementation, the target model performs word segmentation on the question information to determine at least one keyword in the question information; obtains at least one approximate advanced information corresponding to each keyword from the memory bank; determines the approximate advanced information among the at least one approximate advanced information corresponding to each keyword that has an association relationship with the approximate advanced information corresponding to other keywords, and determines the advanced information corresponding to the question information based on multiple pieces of the approximate advanced information having an association relationship.
[0136] For example, if the question information includes "How many times can a bee sting a person?", the keywords may include bee, sting, and person; further, the approximate advanced information may include an explanatory description of the bee, such as biological characteristics, living habits, ecological roles, and relationship with humans; it may also include an explanatory description of "sting", such as basic meaning, attributes, etc.; biological characteristics and living habits of humans, etc.
[0137] Among the above approximate advanced information, there is an association relationship between the biological characteristics of the bee, the attributes of the sting (such as verb, attack), and the biological characteristics of the person. The target module can determine the advanced information corresponding to the question information based on the biological characteristics of the bee, the attributes of the sting, and the biological characteristics of the person.
[0138] In some alternative embodiments, after the target model generates the advanced information, it transmits the advanced information to the memory bank; so that the memory bank stores the question information and the advanced information corresponding to the question information; and, based on the question information and the advanced information, performs reasoning, and updates the advanced information corresponding to at least one question information stored in the memory bank according to the reasoning result or reasoning logic.
[0139] That is, the memory bank learns the reasoning ability of the target model, and based on the reasoning result of the target model, updates and evolves the advanced information stored in the memory bank to make the advanced information more correct, reasonable, and logical.
[0140] Step S405, based on the advanced information corresponding to the question information, determine the first target response in the first basic response and the first reflection response.
[0141] The specific step process of step S405 is the same as that of step S305, and will not be repeated here.
[0142] Step S406, determine the feedback information corresponding to the problem information based on the first target response.
[0143] The specific step process of step S406 is the same as that of step S306, and will not be repeated here.
[0144] In this way, through the large model inference method provided by the embodiments of the present disclosure, based on the dynamic reflection of internal and external collaboration, an evolvable memory bank is provided. The memory bank stores the advanced ideas for solving different problem information in the form of <problem, advanced idea> pairs, which may include the knowledge required to solve the problem, analyze the problem, and provide the basic ideas for solving the problem. These ideas are extracted from the problem-solving processes of different tasks. With the interaction with the target model and task execution, the memory bank will update and upgrade the stored advanced ideas, realizing the evolution and improvement of the memory bank; the advanced model will obtain the advanced ideas of approximate problem information from the memory bank for analogical learning, and adaptively generate the advanced ideas for the current problem information, and at the same time add the generated advanced ideas to the memory bank. This process strengthens the abstract reasoning ability of the target model. The target model enables the large language model to rise from the perspective of an answerer to the height of a teacher to externally select and globally review the basic response and reflection response, improving the reasoning ability of the large model.
[0145] Figure 7 Shows the fifth optional process schematic diagram of the large model inference method provided by the embodiments of the present disclosure; Figure 8 Shows the schematic diagram of the large model inference process provided by the embodiments of the present disclosure, and will be described according to each step.
[0146] Step S701, generate an initial response.
[0147] In some embodiments, as Figure 8 shown, the carrier combines k input-output examples (i.e., problem information and feedback information pairs) and the input problem information x into a prompt and inputs it into the self-reflection part (Self-reflectAgent) of the large model, so that the self-reflection part learns from the k input-outputs to obtain the first basic response response0 (Basic response) of the problem information x.
[0148] Step S702, the target model generates advanced ideas.
[0149] In some embodiments, as Figure 8As shown, there is a memory bank outside the large model, which stores the advanced thoughts generated when solving different tasks (or problem information). The memory bank stores the advanced thoughts for solving different problem information in the form of <problem, advanced thought> pairs, which may include the knowledge required to solve the problem, the analysis of the problem, and the basic thoughts for providing solutions to the problem. These thoughts are refined from the process of solving problem information for different tasks.
[0150] In some embodiments, as Figure 8 shown, the target model (Teacher Agent) retrieves in the memory bank based on the problem information x, obtains the problem information and advanced thoughts related to the problem information x for analogical learning, and generates the advanced thought (meta-thought) of the problem information x; updates the advanced thought of the problem information x to the memory bank to update and upgrade the advanced thought information in the memory, realizing the evolution and improvement of the memory bank.
[0151] Step S703, conduct self-reflection on the first basic response.
[0152] In some embodiments, as Figure 8 shown, self-reflection includes two parts: self-assessment and self-correction.
[0153] In some embodiments, the self-reflection part of the large model evaluates the quality of the first basic response response0 according to predefined metrics, generating the first evaluation information. The predefined metrics may include at least one of: whether it includes a preset character, the rationality of the first basic response, the logic of the first basic response, and whether the first basic response conforms to the facts. Specifically, the predefined metrics can be set according to the requirements of the problem information. For example, in experiments, it is found that for some questions in StrategyQA, the large model has uncertain answers. Therefore, when self-assessing, the self-reflection part is prompted to check whether the first basic response in the book directly gives the answer of 'true' or 'false', and rationality is also set up to verify whether the reasoning is logical, and authenticity is used to verify whether the output conforms to the facts.
[0154] In some embodiments, self-correction includes that in the iterative process, the self-reflection part generates the first reflection response response according to the problem information, the first evaluation information, and the first basic response reflect .
[0155] Step S704, external abstract reasoning.
[0156] In some embodiments, as Figure 8As shown, the large model cannot directly determine which of the first reflective response and the first basic response is the more appropriate target response. Therefore, the target model needs to perform abstract reasoning based on the advanced idea meta-. The target model will compare whether the first reflective response is consistent with the first basic response, judge whether the two responses meet the advanced idea, and select a better response from them as the first target response. Especially when the first reflective response is inconsistent with the first basic response, the target model can accurately select a more appropriate response.
[0157] Step S705, iterative update.
[0158] In some embodiments, steps S703 and S704 are repeated iteratively until a specific stop condition is met, such as reaching the maximum number of iterations N. In the experiment, the termination condition we set is to reach the maximum number of iterations 4 (i.e., the fourth execution of steps S703 and S704). In the last iteration, the target response obtained in S704 is used as the final output feedback information.
[0159] In this way, through the large model reasoning method provided by the embodiments of the present disclosure, it is jointly completed by the self-reflection part and the target model. The self-reflection part is responsible for internal self-reflection, and the target model is responsible for the external dynamic selection of responses: selecting a better response from the basic response and the reflective response for the next reflection, realizing the dynamic update of the better response in the iteration, avoiding the influence of the wrong answer caused by the reflection hallucination on the next iteration, and achieving the continuous improvement and dynamic optimization of the large model reasoning.
[0160] Next, an example will be used for illustration.
[0161] The question information is "Can a bee sting a person multiple times?" Correspondingly, the first basic response of the large model is: "Let's think about this problem step by step. When a bee stings a person, its stinger will become stuck in the person's skin and will be torn off from the bee's body, which will cause the bee to die soon after. So, a bee can only sting once because they cannot survive after stinging. Therefore, the answer is: False." That is, "No, a bee cannot sting a person multiple times."
[0162] The advanced ideas generated by the target model include: "Study the physiological structure of the bee's stinger and the reaction of the stinger when the bee stings a mammalian target. Think about what happens to the bee and its stinger after the stinging event."
[0163] The self-reflection part of the large model conducts self-assessment to obtain the first assessment information. The process includes:
[0164] What problems exist in the above answer?
[0165] Clarity of the answer: The given answer concludes with "wrong", which is consistent with the information provided. However, this conclusion does not directly address the question of whether a single bee can sting multiple times, making the answer unclear.
[0166] Rationality: The answer explains the mechanism of a bee's stinger getting stuck (in human skin) and the bee dying after stinging. Although this information is correct, it does not fully answer the question of whether a single bee can sting multiple times.
[0167] Authenticity: The information provided about the bee's stinger being torn off and the bee dying after stinging is accurate. However, this answer does not consider that not all bees behave the same way. Bees are special in this regard, but other species of bees, such as bumblebees, can sting multiple times without dying.
[0168] Based on the question information, the first evaluation information, and the first basic response, the large model obtains the first reflective response, which includes: "The following is a better answer: A bee can only sting once because its stinger gets stuck in the person's skin and is torn off, which causes the bee to die. However, it should be noted that not all bees are like this. For example, bumblebees can sting multiple times without dying. So the answer is: Correct. (×)". At this time, after the large model reflects, the response output by the large model is updated from "wrong" to "correct". However, the updated reflective response does not match the facts. Therefore, the target model makes a dynamic selection between the first basic response and the first reflective response based on the advanced idea, including: According to the provided question and prompt, recommending Thought Chain 1 (COT 1) as a better choice because it gives a clear, well-organized, and directly relevant answer to the question without introducing unnecessary information about other bee species. It closely follows the provided prompt and gives a concise response based on the specific behavior of bees. The better thought chain after comparison: Thought Chain 1 (COT 1), that is, the first basic response is the target response, which is the feedback response finally output by the large model.
[0169] In some embodiments, experiments are conducted on the StrategyQA commonsense reasoning dataset. The GPT model is accessed through the OpenAI API, and GPT-3.5-turbo is used as the base large model for the self-reflection part and the target model of the large model. Three types of final outputs are defined in the experiment:
[0170] Critic means that in each iteration, the reflective response generated by the Critic method (prior art method) is used as the output result;
[0171] pred means that in each iteration, the Meta-critic method (the large model reasoning method described in this disclosure) is used, and the reflective response generated by the self-reflection part of the large model is used as the output result;
[0172] best_pred represents using the Meta-critic method in each iteration and taking the better response generated by the target model as the output result.
[0173] Figure 9 The iterative output results of different methods on the same dataset are shown.
[0174] In the said results, the final output of the Meta-Critic method of the present disclosure includes the accuracy rate of the target response finally determined by the Meta-Critic target model; the intermediate output of the Meta-Critic method of the present disclosure includes the output of the self-reflection part after Meta-Critic iteration; the existing method Critic is the output obtained by using the existing method Critic.
[0175] As Figure 9 shown, as the number of iterations increases, the accuracy of the existing method Critic shows a continuous downward trend, while the large model inference method Meta-Critic provided by the embodiments of the present disclosure shows a continuous upward trend, indicating that using Meta-Critic can retain correct answers, correct wrong answers in each iteration, and continuously improve the accuracy rate. The accuracy rate in the fourth iteration has increased by 4.6% compared to the initial response, and has increased by 11.8% compared to the accuracy rate of Critic in the fourth iteration.
[0176] In addition, more iteration rounds will cause the inference cost to become higher. From Figure 9 it can be seen that the accuracy rate gain of Meta-Critic from the 0th to the 1st time is the largest. As the number of iteration rounds increases, the accuracy rate gradually stabilizes. Meta-Critic can reach a high accuracy rate within a short iteration cycle. Therefore, Meta-Critic can achieve a balance between short iteration rounds and high accuracy rate.
[0177] In some embodiments, statistical analysis is performed on the answers to 500 questions in StrategyQA, and they are classified according to whether they are from correct to correct, from correct to incorrect, from wrong to correct, and from one wrong to another wrong. The experimental comparison results are shown in the following table.
[0178] Table 1 Changes in answers of the Meta-Critic method every two iterations
[0179]
[0180]
[0181] Table 2 Changes in answers of the Critic method every two iterations
[0182]
[0183] Comparing Table 1 and Table 2, it can be seen that using Meta-Critic can identify and correct the wrong answers generated by itself, while reducing the possibility of modifying the correct answers into wrong answers, effectively reducing the impact of reflection hallucination, and achieving continuous improvement and optimization of the large model inference.
[0184] Figure 10 The optional structural schematic diagram of the large model inference device provided by the embodiments of the present disclosure is shown, and will be described according to each step.
[0185] In some embodiments, the large model inference device 900 includes an inference unit 901, a self-reflection unit 902, a dynamic determination unit 903, and an output unit 904.
[0186] The inference unit 901 is configured to infer a first basic response from the problem information through the large model;
[0187] The self-reflection unit 902 is configured to infer a first reflection response corresponding to the problem information based on the first evaluation information obtained by evaluating the first basic response;
[0188] The dynamic determination unit 903 is configured to determine a first target response among the first basic response and the first reflection response based on the advanced information corresponding to the problem information; wherein, the advanced information represents the inference idea across tasks;
[0189] The output unit 904 is configured to determine the feedback information corresponding to the problem information based on the first target response.
[0190] Specifically, the output unit 904 uses the first target response as the second basic response;
[0191] Based on the second evaluation information obtained by evaluating the second basic response, infer a second reflection response corresponding to the problem information;
[0192] Based on the advanced information, determine a second target response among the second basic response and the second reflection response;
[0193] Iteratively execute the above steps until the iteration termination condition is satisfied, and determine the feedback information corresponding to the problem information based on the Nth target response, where N is a positive integer greater than or equal to 2;
[0194] Wherein, the iteration termination condition includes that the number of iterations reaches the number threshold or the second target response meets the quality requirement.
[0195] Before determining the first target response among the first basic response and the first reflective response based on the advanced information corresponding to the question information, the dynamic determination unit 903 is further configured to obtain the advanced information generated by the target model.
[0196] Wherein, the target model determines approximate advanced information related to the keyword and / or the attribute information from the memory bank based on the keyword in the question information and / or the attribute information of the question information; the target model generates the advanced information based on the approximate question information corresponding to the question information and the approximate advanced information; alternatively, the target model generates the advanced information based on the approximate question information corresponding to the question information, the advanced information corresponding to the approximate question information, and the approximate advanced information.
[0197] In some embodiments, the memory bank stores the approximate question information and the approximate advanced information corresponding to the question information; alternatively, the memory bank stores the approximate question information, the approximate advanced information corresponding to the question information, and the advanced information corresponding to the approximate question information; the advanced information generated by the target model is used to update the advanced information corresponding to at least one question information stored in the memory bank.
[0198] Before determining the first target response among the first basic response and the first reflective response based on the advanced information corresponding to the question information, the dynamic determination unit 903 is further configured to obtain the advanced information generated by the target model.
[0199] Wherein, the target model determines at least one piece of approximate advanced information corresponding to each keyword from the memory bank based on at least one keyword in the question information; the target model determines the advanced information corresponding to the question information based on multiple pieces of the approximate advanced information having an association relationship.
[0200] In some embodiments, the memory bank stores at least one piece of approximate advanced information corresponding to the question information.
[0201] After the target model generates the advanced information, the question information is stored in the memory bank.
[0202] The self-reflection unit 902 is specifically configured to determine the first evaluation information corresponding to the first basic response based on at least one of whether the first basic response includes a preset character, the rationality of the first basic response, the logic of the first basic response, and whether the first basic response conforms to the facts.
[0203] The self - reflection unit 902 is specifically configured to update the first basic response based on the problem information and the first evaluation information, so that the updated first basic response matches the first evaluation information;
[0204] Determine that the updated first basic response is the first reflection response corresponding to the problem information.
[0205] The self - reflection unit 902 is specifically configured to, if the preset character is not included in the first basic response, add the preset character to the first basic response based on the semantics of the first basic response;
[0206] If the first basic response does not meet the rationality, adjust the content of the first basic response so that the adjusted first basic response meets the rationality;
[0207] If the first basic response does not meet the logic, adjust the statement order or content of the first basic response so that the adjusted first basic response meets the logic;
[0208] If the first basic response does not conform to the facts, adjust the statement order or content of the first basic response so that the adjusted first basic response conforms to the facts.
[0209] According to an embodiment of the present disclosure, the present disclosure also provides an inference device based on a large - model and a readable storage medium.
[0210] Figure 11 A schematic block diagram of an exemplary large - model - based inference device 800 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0211] As Figure 11As shown, the large model-based inference device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the large model-based inference device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0212] Multiple components in the large model-based inference device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the large model-based inference device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0213] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the large model inference method. For example, in some embodiments, the large model inference method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the large model-based inference device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the large model inference method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the large model inference method in any other appropriate way (e.g., by means of firmware).
[0214] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0215] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0216] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0217] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0218] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0219] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.
[0220] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. No limitation is imposed herein.
[0221] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one such feature. In the description of this disclosure, "a plurality" means two or more unless otherwise specifically defined.
[0222] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims described above.
Claims
1. A large model inference method, the method comprising: Inferring a first basic response from the problem information through a large model; Based on the first evaluation information obtained by evaluating the first basic response, inferring a first reflection response corresponding to the problem information; Based on the advanced information corresponding to the problem information, determining a first target response among the first basic response and the first reflection response; wherein, the advanced information represents cross-task inference ideas; Determining feedback information corresponding to the problem information based on the first target response.
2. The method according to claim 1, wherein the determining the feedback information corresponding to the problem information based on the first target response includes: Using the first target response as a second basic response; Based on the second evaluation information obtained by evaluating the second basic response, inferring a second reflection response corresponding to the problem information; Based on the advanced information, determining a second target response among the second basic response and the second reflection response; Iteratively executing the above steps until an iteration termination condition is met, and determining the feedback information corresponding to the problem information based on the Nth target response, where N is a positive integer greater than or equal to 2; wherein, the iteration termination condition includes that the number of iterations reaches a threshold or the second target response meets the quality requirement.
3. The method according to claim 1, before determining the first target response among the first basic response and the first reflection response based on the advanced information corresponding to the problem information, the method further includes: Obtaining the advanced information generated by the target model; wherein, the target model determines approximate advanced information related to the keyword and / or the attribute information from the memory bank based on the keyword and / or the attribute information in the problem information; The target model generates the advanced information based on the approximate problem information corresponding to the problem information and the approximate advanced information; or, the target model generates the advanced information based on the approximate problem information corresponding to the problem information, the advanced information corresponding to the approximate problem information, and the approximate advanced information.
4. The method according to claim 3, The memory bank stores the approximate problem information and the approximate advanced information corresponding to the problem information; or, the memory bank stores the approximate problem information, the approximate advanced information corresponding to the problem information, and the advanced information corresponding to the approximate problem information; The advanced information generated by the target model is used to update the advanced information corresponding to at least one problem information stored in the memory bank.
5. The method according to claim 1, before determining the first target response among the first basic response and the first reflection response based on the advanced information corresponding to the problem information, the method further includes: Obtaining the advanced information generated by the target model; wherein, the target model determines at least one approximate advanced information corresponding to each keyword from the memory bank based on at least one keyword in the problem information; The target model determines the advanced information corresponding to the problem information based on multiple pieces of the approximate advanced information with an association relationship.
6. The method according to claim 5, at least one piece of approximate advanced information corresponding to the problem information is stored in the memory bank; after the target model generates the advanced information, the problem information is stored in the memory bank.
7. The method according to claim 1, before the first evaluation information obtained by evaluating the first basic response, the method further includes: determining the first evaluation information corresponding to the first basic response based on at least one of whether the first basic response includes a preset character, the rationality of the first basic response, the logic of the first basic response, and whether the first basic response conforms to facts.
8. The method according to claim 7, the first reflection response corresponding to the problem information inferred from the first evaluation information obtained by evaluating the first basic response includes: updating the first basic response based on the problem information and the first evaluation information so that the updated first basic response matches the first evaluation information; determining the updated first basic response as the first reflection response corresponding to the problem information.
9. The method according to claim 8, the updating of the first basic response specifically includes at least one of the following: if the first basic response does not include a preset character, adding the preset character to the first basic response based on the semantics of the first basic response; if the first basic response does not meet the rationality, adjusting the content of the first basic response so that the adjusted first basic response meets the rationality; if the first basic response does not meet the logic, adjusting the statement order or content of the first basic response so that the adjusted first basic response meets the logic; if the first basic response does not conform to facts, adjusting the statement order or content of the first basic response so that the adjusted first basic response conforms to facts.
10. An inference device based on a large model, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.