Inference method and device based on large model and program product

By dynamically adjusting the temperature coefficient and filtering candidate sequences during the decoding process of large language models, the problem of difficulty in taking into account both the diversity and accuracy of answers in the prior art is solved, and efficient and accurate answer generation is achieved.

CN119990339AActive Publication Date: 2025-05-13HUNDSUN TECH

Patent Information

Application Number
CN202510473353.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The temperature coefficient of existing large language models in the decoding stage is usually constant, which is difficult to take into account the diversity and accuracy of answers. Methods that improve the diversity of answers will greatly increase the computational overhead.

Method used

By dynamically adjusting the temperature coefficient during the generation of candidate sequences, it decreases from high to low, gradually transitioning from pursuing diversity to pursuing accuracy, and screening candidate sequences with high confidence based on the sequence entropy value for reply text generation.

Benefits of technology

While improving the diversity and accuracy of answers, the calculation overhead of the model is reduced, processing efficiency is improved, and the accuracy of the final output reply text is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990339A_ABST
    Figure CN119990339A_ABST
Patent Text Reader

Abstract

According to the reasoning method and device based on the large model and the program product, through a prompt word requested at a time, on the premise that multiple prompt words are not artificially constructed, the lexical elements in each candidate sequence are made to be more uniform in the generation process of the candidate sequence by means of the decline strategy of the temperature coefficient, and therefore the lexical elements in the candidate sequence are more uniform in the generation process of the candidate sequence. And the pursuit of diversity is gradually transited to the pursuit of accuracy in sequence. When the reply text is constructed, the candidate sequence with low self-confident degree is removed, and the candidate sequence with higher self-confident degree is reserved to generate the reply text. Therefore, the accuracy of the finally output reply text is ensured while the processing efficiency is improved and the model overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to, in particular, a large model-based reasoning method, device and program product. Background Art

[0002] In recent years, large language models (LLMs) have demonstrated powerful text generation and understanding capabilities in the field of natural language processing. In order to ensure the diversity of model output, temperature coefficients are often used to control the smoothness of probability distribution during the decoding stage. However, traditional temperature coefficients are usually constants in the text generation process of a model call. Among them, a higher temperature coefficient can increase the randomness of the generated sequence. A lower temperature coefficient can reduce randomness and make the output more certain and predictable.

[0003] However, the core contradiction of the above decoding mechanism lies in the irreconcilability between diversity requirements and accuracy assurance. On the one hand, although high-temperature sampling can support multi-path exploration, it is prone to factual errors or logical breaks; on the other hand, although the low-temperature strategy improves local consistency, it inhibits the divergent thinking required for complex problems.

[0004] Therefore, in order to compensate for the above problems, existing technologies often use different prompt words to call the model multiple times to obtain different sequences, thereby improving the diversity of answers. However, this method will greatly increase the computational overhead and the improvement effect is limited. Summary of the invention

[0005] The purpose of this application is to provide a large model-based reasoning method, device and program product, which are used to take into account the diversity and accuracy of the large model answers while reducing the large model overhead.

[0006] In order to achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows: In a first aspect, an embodiment of the present application provides a reasoning method based on a large model, including: Based on the input prompt word, multiple candidate sequences are obtained; each candidate sequence has a corresponding sequence entropy value, and each sequence entropy value is used to characterize the diversity and accuracy of the corresponding candidate sequence; the sequence entropy value is composed of the word-unit entropy values ​​of all word-units in the corresponding candidate sequence; the word-unit entropy values ​​of all word-units in the same candidate sequence are generated in sequence based on a decreasing strategy of a temperature coefficient; Screening the plurality of candidate sequences according to a screening rule to obtain at least one target candidate sequence; the screening rule is used to indicate to eliminate candidate sequences with low confidence among the plurality of candidate sequences; Generate a reply text based on all the target candidate sequences.

[0007] Optionally, the step of obtaining a plurality of candidate sequences based on the input prompt word includes: Based on the prompt word, at each generation moment, the word element corresponding to the generation moment is determined according to the current temperature coefficient of each candidate sequence; wherein, corresponding to each generation moment, the value of the current temperature coefficient of the candidate sequence decreases successively according to the decreasing strategy; A corresponding candidate sequence is constructed based on the word elements corresponding to all the generation moments.

[0008] Optionally, the screening rule is a screening range; the screening range is used to characterize the relative relationship between the confidence levels of the plurality of candidate sequences; the step of screening the plurality of candidate sequences according to the screening rule to obtain at least one target candidate sequence comprises: Determine, according to the screening range, at least one candidate sequence to be eliminated from the plurality of candidate sequences; the candidate sequence to be eliminated is a candidate sequence with relatively low confidence among all candidate sequences; All the eliminated candidate sequences are eliminated, and the remaining candidate sequences are used as the target candidate sequences.

[0009] Optionally, the step of generating a reply text according to all the target candidate sequences includes: Determining whether the contents of all the target candidate sequences are the same; If all are the same, generating a reply text according to the target candidate sequence; If they are at least partially different, at least two different target candidate sequences are screened, and a reply text is generated based on the screening result.

[0010] Optionally, the step of determining whether the contents of all the target candidate sequences are the same includes: Determine the demand type corresponding to the prompt word; If the requirement type is a result type, judging whether the contents of all the target candidate sequences are the same according to the results of all the target candidate sequences; If the requirement type is a non-result type, a content identification module is called to identify all the target candidate sequences to obtain the contents of all the target candidate sequences; and it is determined whether the contents of all the target candidate sequences are the same.

[0011] In a second aspect, an embodiment of the present application provides a large model-based reasoning device, including: a generation module, a screening module, and an output module; The generation module is used to obtain multiple candidate sequences based on the input prompt word; each candidate sequence has a corresponding sequence entropy value, and each sequence entropy value is used to characterize the diversity and accuracy of the corresponding candidate sequence; the sequence entropy value is composed of the word-unit entropy values ​​of all word-units in the corresponding candidate sequence; the word-unit entropy values ​​of all word-units in the same candidate sequence are generated in sequence based on the decreasing strategy of the temperature coefficient; The screening module is used to screen the multiple candidate sequences according to a screening rule to obtain at least one target candidate sequence; the screening rule is used to indicate to eliminate candidate sequences with low confidence among the multiple candidate sequences; The output module is used to generate a reply text based on all the target candidate sequences.

[0012] In a third aspect, an embodiment of the present application provides an electronic device, including: A memory for storing one or more programs; processor; When the one or more programs are executed by the processor, the method as described in any one of the first aspects above is implemented.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method as described in any one of the above-mentioned first aspects.

[0014] In a fifth aspect, an embodiment of the present application provides a program product, which, when executed by a processor, implements a method as described in any one of the above-mentioned first aspects.

[0015] Compared with the prior art, the inference method, device and program product based on a large model provided by the embodiment of the present application, through a prompt word for a request, without artificially constructing multiple prompt words, using the decreasing strategy of the temperature coefficient, accompanying the generation process of the candidate sequence, so that each word in each candidate sequence gradually transitions from pursuing diversity to pursuing accuracy. Then, when constructing the reply text, the candidate sequences with lower confidence are eliminated, and the candidate sequences with higher confidence are retained for the generation of the reply text. Thereby, while improving the processing efficiency and reducing the model overhead, the accuracy of the reply text finally output is guaranteed.

[0016] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A schematic diagram of a self-optimization architecture of the prior art; Figure 2 A schematic diagram of a self-optimization architecture provided by an embodiment of the present invention Figure 3 A flowchart of a large model-based reasoning method provided by an embodiment of the present invention; Figure 4 A schematic diagram of a flow chart of another large model-based reasoning method provided in an embodiment of the present invention; Figure 5 A schematic diagram of another self-optimization architecture provided by an embodiment of the present invention; Figure 6 It is a schematic diagram of the temperature coefficient distribution based on the drop amplitude; Figure 7 It is a schematic diagram of the distribution of sequence entropy values ​​under dynamic temperature coefficient; Figure 8 A flowchart of another large model-based reasoning method provided by an embodiment of the present invention; Fig. 9 A schematic diagram of another self-optimization architecture provided by an embodiment of the present invention; Fig.10 A schematic diagram of a self-optimization process of the prior art; Fig.11 A schematic diagram of a flow chart of another large model-based reasoning method provided in an embodiment of the present invention; Fig.12 A schematic diagram of another self-optimization architecture provided by an embodiment of the present invention; Fig.13 A schematic diagram showing the comparison of the results of the temperature coefficient reduction strategy of this application with other methods using a static temperature coefficient; Fig.14 A schematic diagram showing the comparison of the number of calls and token consumption of the present application solution and other methods on a large model; Fig.15 A schematic diagram of an inference device provided by an embodiment of the present invention; Fig.16 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0020] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0021] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0022] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0023] In recent years, large language models (LLMs) have made significant progress in the field of natural language processing (NLP), but they still face many challenges in complex reasoning tasks. In order to improve the reasoning ability of large models, researchers have proposed a variety of methods, including Chain-of-Thought (CoT), Reinforcement Learning (RL), and Self-Optimization. These methods start from different perspectives, by optimizing the internal consistency and reasoning process of the model, reducing the generation of hallucination content, and improving the performance of the model in multi-step reasoning tasks.

[0024] The core idea of ​​the thought chain is to guide the model to show the reasoning process when answering questions by showing a small number of reasoning examples to the large language model. This method breaks down complex problems into multiple intermediate steps, allowing the model to handle problems more systematically. For example, when solving math problems, CoT helps the model better understand and generate the reasoning process by showing the steps to solve the problem. CoT not only improves the model's reasoning accuracy, but also enables it to handle more complex tasks such as common sense reasoning and symbolic reasoning.

[0025] Reinforcement learning significantly expands the reasoning capabilities of large language models by automatically generating high-quality reasoning trajectories through a trial-and-error search algorithm. During training, the model optimizes its reasoning strategy by receiving reward signals. This approach not only improves the model's reasoning accuracy, but also enables it to think more deeply at test time.

[0026] The core goal of self-optimization is to improve the performance and reasoning ability of the model through internal mechanisms. Self-optimization emphasizes that the model achieves better results through self-feedback, self-adjustment, and self-improvement. For example, the self-consistency method improves the reasoning ability of the model by generating multiple reasoning paths and selecting the most consistent answer. The self-refine method improves the accuracy of the model by gradually optimizing the reasoning process of the model. After generating the reasoning path, the model will adjust it based on the feedback of the intermediate steps to generate more accurate results. Self-contrast identifies and corrects model errors by comparing the results of different reasoning paths. This method is similar to the "comparative analysis" when humans solve problems, and potential problems are discovered by comparing the results of different methods.

[0027] Alternatively, the existing large model decoding strategy often falls into the problem of balancing diversity and accuracy. Usually, in the process of text generation, the temperature coefficient is used to control the smoothness of the probability distribution and indirectly control the certainty of the generated results. The following uses self-optimization as an example to illustrate the decoding strategy of the existing technology. Specifically, Figure 1 For a schematic diagram of the self-optimization architecture of the prior art, see Figure 1 , which shows that in the prior art, when a user inputs a request (which may be in the form of a prompt) to a large model, the large model can construct each token in each candidate sequence (candidate) based on a specific temperature coefficient. Figure 1 The candidate sequences shown are candidate sequence 1, ... candidate sequence i, ..., candidate sequence n, etc. Then, the multiple candidate sequences are screened through the screening mechanism to obtain the final reply text.

[0028] It can be seen that tokens, as the basic units of the text generation process (such as words or subwords), are the basis of the generation process. The model generates a probability distribution for the next token by analyzing the context. The temperature coefficient is like a "tuning knob": high temperature (such as T>1) will make the probability distribution of tokens more uniform, increase the possibility of low-probability tokens, thereby increasing the entropy (uncertainty) of the distribution, and make the model tend to generate diverse and even creative texts; low temperature (such as T<1) will strengthen the advantages of high-probability tokens, reduce the entropy value, and make the output more conservative and stable. Finally, based on the adjusted probability distribution, the model uses a sampling strategy to splice tokens into a complete text sequence, such as the candidate sequence above.

[0029] For example, high temperature may generate novel metaphors such as "the sunset glows like flames", while low temperature is more likely to output conventional expressions such as "the weather is fine this evening". Therefore, the temperature coefficient becomes the core hub for balancing the "creativity" and "reliability" of text sequences by regulating the size of entropy, and the choice of word units determines the specific form of the generated content in this dynamic process.

[0030] However, in the process of constructing each candidate sequence, the prior art often uses the traditional temperature coefficient T to construct each token in the candidate sequence. In this process, the traditional temperature coefficient T is a constant. Figure 1 ,Should Figure 1 The probability distribution of the response word is shown, with the vertical axis being the probability value of the word and the horizontal axis being the text length of the candidate sequence. The inventors found that this mechanism means that during the entire production process of the candidate sequence, since each word of the candidate sequence is based on the same temperature coefficient, when T is high, the entire candidate sequence tends to be more "random". When T is low, the entire candidate sequence tends to be more "conservative".

[0031] Based on the above-mentioned problem of balancing diversity and accuracy in the prior art, the present application provides a mechanism for dynamically adjusting the temperature coefficient, the core idea of ​​which is: when constructing any candidate sequence, the corresponding temperature coefficient is systematically controlled to decrease from high to low, so that in the initial stage of candidate sequence construction, a wider range of potential outputs of word units are explored based on higher temperature coefficients. Then, as the candidate sequence generation process progresses and the text length increases, the temperature coefficient is systematically reduced to enhance the coherence and accuracy of the corresponding candidate sequence text in reasoning logic. This adaptive temperature control enables the model to explore diverse outputs in the initial stage, and focus on more accurate reasoning as text generation unfolds, thereby optimizing overall performance.

[0032] Specifically, Figure 2 A schematic diagram of a self-optimization architecture provided by an embodiment of the present invention, see Figure 2 , see the mechanism of dynamically adjusting the temperature coefficient described above. When constructing the candidate sequence, the initial word unit is determined based on a higher temperature coefficient, so that the word units in the early stage of each candidate sequence generation stage have stronger randomness. As the construction process of each candidate sequence progresses, based on the decreasing strategy of the temperature coefficient, the temperature coefficient values ​​of subsequent word units will also become lower and lower. Based on the previous text, when the temperature coefficient value decreases, the model tends to choose to generate word units with better coherence and lower randomness. Based on this idea, as the temperature coefficient of each word unit decreases during the generation process of the candidate sequence, the text gradually transitions from randomness to coherence.

[0033] From the perspective of the temperature coefficient as a factor in regulating entropy, the technical solution of the present application can be found that by systematically reducing the temperature coefficient during the construction of the candidate sequence, the entropy of the candidate sequence can be gradually reduced, that is, the uncertainty can be reduced. Figure 2 The "response word probability distribution" is reduced as the construction process progresses, that is, as the length of the text increases, so as to take into account the randomness and certainty of the text corresponding to the candidate sequence. The following is an exemplary explanation of the reasoning method based on the large model in combination with the temperature coefficient reduction strategy proposed in this application. Specifically, Figure 3 A flowchart of a large model-based reasoning method provided by an embodiment of the present invention is shown in FIG. Figure 3 , the method comprising: Step 100: Based on the input prompt word, multiple candidate sequences are obtained.

[0034] Among them, Figure 2As shown, each candidate sequence has a corresponding sequence entropy value, and each sequence entropy value is used to characterize the diversity and accuracy of the corresponding candidate sequence; the sequence entropy value is composed of the word-unit entropy values ​​of all words in the corresponding candidate sequence; the word-unit entropy values ​​of all words in the same candidate sequence are generated sequentially based on the decreasing strategy of the temperature coefficient.

[0035] Specifically, there are many ways to form the sequence entropy value from the word-unit entropy values ​​of all word-units in the candidate sequence. For example, the sequence entropy value may be the sum of the word-unit entropy values ​​of all word-units in the candidate sequence generation process; or, the sequence entropy value may be the average of the word-unit entropy values ​​of all word-units in the candidate sequence generation process.

[0036] It should be noted that in the process of generating each word unit of each candidate sequence, the generation has been carried out based on the decreasing temperature coefficient. Therefore, the entropy value of each word unit decreases as the corresponding temperature coefficient decreases. Therefore, the sequence entropy value of the final candidate sequence is more balanced than the entropy value obtained by using a constant temperature coefficient in the prior art, and can take into account both diversity and accuracy.

[0037] Step 101: Screen multiple candidate sequences according to a screening rule to obtain at least one target candidate sequence.

[0038] The screening rule is used to indicate the elimination of candidate sequences with low confidence among multiple candidate sequences.

[0039] It should be noted that due to the higher temperature coefficient, the word unit it constitutes has a higher word unit entropy value and a lower confidence. Therefore, since the sequence entropy value is composed of the entropy values ​​of all the corresponding word units, it means that in a possible implementation, among multiple candidate sequences, based on different sequence entropy values, candidate sequences with higher entropy values, that is, lower confidence, can be eliminated.

[0040] Step 102: Generate a reply text based on all target candidate sequences.

[0041] The large model-based reasoning method provided by the embodiment of the present invention uses a prompt word for a request at a time, without artificially constructing multiple prompt words, and uses a decreasing strategy of the temperature coefficient to accompany the generation process of the candidate sequence, so that each word in each candidate sequence gradually transitions from pursuing diversity to pursuing accuracy. Then, when constructing the reply text, the candidate sequence with lower confidence is eliminated, and the candidate sequence with higher confidence is retained to generate the reply text. Thereby, while improving the processing efficiency and reducing the model overhead, the accuracy of the reply text finally output is guaranteed.

[0042] Optionally, a possible implementation method is provided below for step 100. Specifically, Figure 3 On the basis of Figure 4 A flowchart of another large model-based reasoning method provided by an embodiment of the present invention is shown in FIG. Figure 4 , step 100, comprising: Step 100 - 1 : At each generation time, determine the word element corresponding to the generation time according to the current temperature coefficient of each candidate sequence.

[0043] Among them, corresponding to each generation moment, the value of the current temperature coefficient of the candidate sequence decreases successively according to the decreasing strategy.

[0044] Step 100 - 2 : construct a corresponding candidate sequence based on the word units corresponding to all the generation times.

[0045] It should be noted that, for different candidate sequences in the word unit generation process, at the same generation time, their corresponding current temperature coefficients can be the same or different. This is not limited here. Optional, Figure 5 Another schematic diagram of a self-optimization architecture provided for an embodiment of the present invention shows the relationship between a candidate sequence, a word unit, a temperature coefficient, a word unit entropy value, a sequence entropy value, and a generation time.

[0046] For details, see Figure 5 , for a certain generation time, for example, generation time 1, its temperature coefficient is T1, and then based on T1, the word unit 1 of each candidate sequence is generated. Then the word unit entropy value 1 corresponding to the word unit 1 is obtained.

[0047] It should be noted that, since each candidate sequence is random in selecting word 1, the word entropy value 1 of word 1 in different candidate sequences may be different, and thus the sequence entropy value of each candidate sequence may also be different.

[0048] Then, as the candidate sequence generation progresses, for example, for the generation time i, its corresponding temperature coefficient T i It should be smaller than T1 mentioned above.

[0049] Below, a temperature coefficient T is provided. t Possible expressions for :

[0050] in represents the temperature coefficient corresponding to the tth generation moment, represents the initial temperature coefficient, e.g. Figure 5 T1 shown; stands for target temperature coefficient, which represents the temperature with high accuracy; Represents the rate of decrease, the higher it is, the slower the function decreases. It can be seen that this expression converts the temperature coefficient into an exponential function related to the length of the output text.

[0051] therefore, Figure 6 For a schematic diagram of the temperature coefficient distribution based on the drop amplitude, see Figure 6 , the vertical axis is the temperature coefficient value, and the horizontal axis is the text length. It shows the relationship between the dynamic temperature coefficient and the sequence entropy value, text length, and decline rate. Red means a higher temperature coefficient value, and blue means a lower temperature coefficient value.

[0052] Among them, in the prior art, a constant temperature coefficient, such as T=0.8, T=0.2, is Figure 6 The entropy value in is a horizontal straight line.

[0053] Due to the decrease of the temperature coefficient of the present application, the sequence entropy value of the candidate sequence is also in a decaying trend. Therefore, the decreasing strategy of the temperature coefficient of the present application, combined with the screening of candidate sequences based on the sequence entropy value, can be understood as a kind of entropy decay incentive sampling (Entropy Decay Inspiring Sampling, abbreviated as: EDIS). Furthermore, for the change curve of the temperature coefficient using the present application solution (EDIS), the decline range of the curve is different based on different decline rates, such as Figure 6 As shown in , as well as , the corresponding entropy value decreases at different rates as the text length increases.

[0054] Optionally, Figure 7 This is a schematic diagram of the distribution of sequence entropy values ​​under dynamic temperature coefficients, see Figure 7 , the vertical axis is the average entropy of the sequence entropy value, and the horizontal axis is the text length. It shows that as the reasoning process progresses, the temperature coefficient in this scheme decreases systematically, and then the sequence entropy value distribution of different candidate sequences (candidate) at different text lengths.

[0055] Optionally, the relationship between the sequence entropy value and the temperature coefficient is expressed as follows:

[0056] in, is the sequence entropy value of the candidate sequence at the generation time t. T is the temperature coefficient corresponding to the generation time t. is the number of words. For the The raw probability of the word output. For the The probability of a word after controlling the temperature coefficient T.

[0057] Optionally, after constructing multiple candidate sequences in this solution, how to screen them. This application also provides a possible implementation method. First, the screening rules involved in the above text can be a screening range. The screening range is used to characterize the relative relationship between the confidence levels of multiple candidate sequences. The following is an exemplary introduction in conjunction with the screening range. Specifically, in Figure 3 On the basis of Figure 8 A flowchart of another large model-based reasoning method provided by an embodiment of the present invention is shown in FIG. Figure 8 , step 101, comprising: Step 101 - 1 : Determine at least one candidate sequence to be eliminated from a plurality of candidate sequences according to a screening range.

[0058] Among them, the eliminated candidate sequences are candidate sequences with relatively low confidence among all candidate sequences.

[0059] Step 101 - 2 : Eliminate all candidate sequences and use the remaining candidate sequences as target candidate sequences.

[0060] Optionally, the screening range may be implemented based on the sequence entropy value of each candidate sequence. For example, when there are four candidate sequences, two candidate sequences with the highest entropy values ​​among the sequence entropy values ​​are eliminated as candidate sequences for elimination.

[0061] The following uses the Llama series (2-7B-chat, 2-13B-chat, 3.2-3B-Instruct, 3.1-8B-Instruct) models to illustrate the accuracy performance of different screening ranges. See Table 1 for details.

[0062]

[0063] The meaning of "number-number" in the screening range is: 3 candidate sequences - keep 3 candidate sequences as target candidate sequences, 4 candidate sequences - keep 2 candidate sequences as target candidate sequences, and so on. The bold values ​​are used to highlight the best performance data in each column.

[0064] It can be found that in each model version of the current test example, the screening range is: 4 candidate sequences screen 3 target candidate sequences, that is, eliminating a candidate sequence with the highest sequence entropy value, and its average performance index is the best.

[0065] Optional, in Figure 5 On the basis of Fig. 9 Another schematic diagram of a self-optimization architecture provided for an embodiment of the present invention shows how to implement screening after constructing multiple candidate sequences.

[0066] Specifically, in this example, if N candidate sequences are screened, the candidate sequence with the highest sequence entropy value is removed. Fig. 9 , if it constructs 1 to N candidate sequences, and then confirms that the entropy value corresponding to candidate sequence 3 is the highest based on the entropy values ​​of each sequence. Then it can be considered that candidate sequence 3 is the candidate sequence with the lowest confidence. Therefore, it is eliminated, and the remaining candidate sequences are used as target candidate sequences.

[0067] Of course, if the screening range includes N candidate sequences, and the first M candidate sequences with the highest entropy values ​​are eliminated, the processing mechanism is similar and will not be described here.

[0068] Optionally, after obtaining all target candidate sequences, that is, completing the reasoning path, errors in the target candidate sequences constructed by the large model can also be identified and corrected, thereby improving the accuracy of the reply text.

[0069] In order to identify and correct the above errors, the existing technology uses the logic of artificially constructing multiple different prompt words for a request, and then using these prompt words to self-optimize the reply text. Specifically, Fig.10 A self-optimization process diagram of the prior art. Fig.10 In this process, we first construct 1 to N prompt words for a request, and then use the model to generate the corresponding replies for each prompt word. Then, we use the big model to compare these prompt words in pairs, and generate N(N-1) / 2 comparison results. Then, based on these comparison results, we use the big model to summarize and form a checklist, and finally use the checklist to correct the reply.

[0070] The above methods of the prior art aim to solve the stubbornness and overconfidence problems of large models through comparative analysis, but they have huge computational overhead, limited improvement effect, and are still plagued by the stubbornness of large models in the final generation stage. In addition, for large models with weaker capabilities, complex processes and prompt words not only have little effect, but may even have a negative impact. This process is very complicated and redundant.

[0071] Therefore, the present application provides a simplified self-contrast process to solve the above problems. Specifically, Figure 3 On the basis of Fig.11 A flowchart of another large model-based reasoning method provided by an embodiment of the present invention is shown in FIG. Fig.11 , step 102, comprising: Step 102-1: Determine whether the contents of all target candidate sequences are the same.

[0072] Specifically, if they are all the same, execute step 102-2; if they are at least partially different, execute step 102-3.

[0073] Step 102-2: Generate a reply text according to the target candidate sequence.

[0074] Step 102-3: Screen at least two different target candidate sequences, and generate a reply text according to the screening result.

[0075] For ease of explanation Fig.11 The role of the process in this application scheme is Fig. 9 On the basis of Fig.12 Another schematic diagram of a self-optimization architecture provided for an embodiment of the present invention shows how to implement a simplified self-optimization process after determining a target candidate sequence.

[0076] First, as mentioned above, this solution does not need to manually construct multiple different prompt words in order to obtain multiple responses to a request. Instead, it uses the dynamic decreasing mechanism of the temperature coefficient to construct multiple target candidate sequences. While improving efficiency, it also reduces the overhead of repeatedly calling the model.

[0077] The model then judges all target candidate sequences to confirm whether their contents are the same, thereby triggering all the same processing flows and at least some different processing flows.

[0078] For some different processing procedures, this application only screens the target candidate sequences with differences to obtain the corrected response text.

[0079] The above optimization mechanism of the present application does not need to compare multiple reply texts one by one and then construct a checklist to achieve correction, which effectively simplifies the self-optimization process.

[0080] In the above process, there are different types of requirements for requests or prompt words. For example, when calculating a data formula during a request, the requirement for the reply text is clear and belongs to the result type. When the request itself is an open-ended question, the reply texts corresponding to multiple target candidate sequences may differ in words, but they may be the same in semantics. Therefore, these two types of requirements can be treated differently to further improve recognition accuracy. Therefore, a possible implementation of step 102-1 above is: Determine the requirement type that the prompt word corresponds to.

[0081] If the requirement type is a result class, then based on the results of all target candidate sequences, it is determined whether the contents of all target candidate sequences are the same.

[0082] If the requirement type is non-result type, the content identification module is called to identify all target candidate sequences to obtain the content of all target candidate sequences, and then determine whether the content of all target candidate sequences is the same.

[0083] Optionally, the content recognition module can be a third-party semantic recognition tool or semantic recognition model. The content recognition module can effectively recognize the semantic information of all target candidate sequences, and then determine whether they are the same according to the content of the semantic information to improve the accuracy of the determination.

[0084] Optionally, for the above examples, the method of the present application is verified based on the GSM8K, SVAMP, and MATH data sets on the Llama series (2-7B-chat, 2-13B-chat, 3.2-3B-Instruct, 3.1-8B-Instruct) and Qwen2.5 series (3B-Instruct, 7B-Instruct, and 14B-Instruct). The performance comparison between the method of the present application and other methods [Base CoT, Self-Contrast, Self-Refine, and Self-Consistency] is shown in Table 2 below: Table 2 Comparison between this application method and other methods

[0085] As shown in Table 1, the method of the present application has greatly improved over the existing methods in various models.

[0086] Optionally, Fig.13 For a comparison of the results of the temperature coefficient reduction strategy for this application and other methods using static temperature coefficients, see Fig.13 , using the GSM8K dataset as test data, the solution of this application (ours) is compared with the basic chain of thought (Base CoT) and self-consistency (Self-Consistency).

[0087] Specifically, Fig.13 The experimental data comparison based on the basic thinking chain, self-consistency and this scheme is shown for Llama-2-7B, Llama-2-13B, Llama-3.2-3B and Llama-3.1-8B in the LLM model.

[0088] Among them, "Ours (T=0.7)" means: the temperature coefficient is constant at 0.7, and the screening mechanism in the above example of this solution is used, and the corresponding data effect.

[0089] “Ours (EDIS)” refers to the data effect obtained by using the dynamic reduction strategy of the temperature coefficient in the above example and the screening mechanism in the above example of this solution.

[0090] “SC(T=0.7)” means that the temperature coefficient is kept constant at 0.7, and the subsequent processing is performed using the original technical solution of self-consistency.

[0091] “SC(EDIS)” refers to: using the above-mentioned example temperature coefficient dynamic reduction strategy, and then using the self-consistency original technical solution to perform subsequent processing.

[0092] “Base CoT(T=0.7)” means: the temperature coefficient is constant at 0.7, and the subsequent processing is performed using the original technical solution of the Basic Thinking Chain (BaseCoT).

[0093] “Base CoT (EDIS)” refers to: using the dynamic reduction strategy of the above example temperature coefficient, and then using the original technical solution of the Basic Chain of Thought (Base CoT) to perform subsequent processing.

[0094] Visible for Fig.13 The model shown can bring further optimization effects for the dynamic reduction strategy of the temperature coefficient and the optimized screening mechanism shown in the above example of this scheme, whether it is fully adopted or only partially adopted.

[0095] Optionally, Fig.14 This is a diagram comparing the number of calls and token consumption of this application solution and other methods on a large model. Fig.14 Based on the LLM model, with GSM8K and MATH as test data, this application (ours) is compared with Self-Contrast, Self-Refine and Self-Consistency from the average number of model calls (Call Times Avg) and the average consumption of tokens (Tokens Avg). It can be found that this application is at a lower level in terms of model call times (Call Times) and token consumption.

[0096] Optionally, for the screening mechanism in the above example, the following Table 3 provides an accuracy comparison based on GSM8K. Among them, Compare&Checklist&Revision (C&C&R) comes from Self-Contrast, which is a unified selection using LLM after LLM pairwise comparison. Choose does not select, only one is generated; Entropy is selected only based on entropy.

[0097]

[0098] It can be seen that the screening mechanism of the present application has a higher overall accuracy rate than other methods in the prior art.

[0099] Based on the reasoning method provided in the above examples, a reasoning device based on a large model is provided below to implement the steps of each of the above methods, thereby achieving the corresponding technical effects.

[0100] Specifically, Fig.15 A schematic diagram of an inference device provided by an embodiment of the present invention, see Fig.15 The device 20 includes: a generating module 200, a screening module 201 and an output module 202.

[0101] The generation module 200 is used to obtain multiple candidate sequences based on the input prompt word.

[0102] The screening module 201 is used to screen multiple candidate sequences according to screening rules to obtain at least one target candidate sequence.

[0103] The output module 202 is used to generate a reply text according to all target candidate sequences.

[0104] Optionally, the generation module 200 is specifically used to: based on the prompt word, at each generation moment, determine the word element corresponding to the generation moment according to the current temperature coefficient of each candidate sequence; and construct the corresponding candidate sequence according to the word elements corresponding to all generation moments.

[0105] Optionally, the screening module 201 is specifically used to: determine at least one elimination candidate sequence among multiple candidate sequences according to the screening range; the eliminated candidate sequence is a candidate sequence with relatively low confidence among all candidate sequences; eliminate all eliminated candidate sequences and use the remaining candidate sequences as target candidate sequences.

[0106] Optionally, the output module 202 is specifically used to: determine whether the contents of all target candidate sequences are the same; if they are all the same, generate a reply text based on the target candidate sequence; if they are at least partially different, screen at least two different target candidate sequences, and generate a reply text based on the screening result.

[0107] Optionally, the output module 202 is further specifically used to: determine the requirement type corresponding to the prompt word; if the requirement type is a result type, then based on the results of all target candidate sequences, determine whether the contents of all target candidate sequences are the same; if the requirement type is a non-result type, then call the content recognition module to identify all target candidate sequences to obtain the contents of all target candidate sequences; and determine whether the contents of all target candidate sequences are the same.

[0108] The embodiment of the invention further provides an electronic device, which can execute all the steps of the above examples of the embodiment of the invention to achieve the corresponding technical effects. Specifically, Fig.16 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, see Fig.16 , the electronic device 30 , comprises: a memory 301 , a processor 300 ; Memory 301, used to store one or more programs; Processor 300; When one or more programs are executed by the processor, when the electronic device 30 is used to execute the steps shown in the above-mentioned method examples, it can achieve each step and the corresponding technical effects.

[0109] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a part of the module, program segment or code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0110] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0111] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a program product, which is stored in a computer-readable storage medium and includes a number of instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0112] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0113] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential features of the present application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present application. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

Claims

1. A large model-based reasoning method, characterized in that: include: Based on the input prompt word, multiple candidate sequences are obtained; Each candidate sequence has a corresponding sequence entropy value, and each of the sequence entropy values ​​is used to characterize the diversity and accuracy of the corresponding candidate sequence; the sequence entropy value is composed of the word-unit entropy values ​​of all the word-units in the corresponding candidate sequence; the word-unit entropy values ​​of all the word-units in the same candidate sequence are generated sequentially based on the decreasing strategy of the temperature coefficient; Screening the plurality of candidate sequences according to a screening rule to obtain at least one target candidate sequence; the screening rule is used to indicate to eliminate candidate sequences with low confidence among the plurality of candidate sequences; Generate a reply text based on all the target candidate sequences.

2. The method according to claim 1, characterized in that The step of obtaining a plurality of candidate sequences based on the input prompt word comprises: Based on the prompt word, at each generation moment, the word element corresponding to the generation moment is determined according to the current temperature coefficient of each candidate sequence; wherein, corresponding to each generation moment, the value of the current temperature coefficient of the candidate sequence decreases successively according to the decreasing strategy; A corresponding candidate sequence is constructed based on the word elements corresponding to all the generation moments.

3. The method according to claim 1, characterized in that The screening rule is a screening range; the screening range is used to characterize the relative relationship between the confidence levels of the plurality of candidate sequences; The step of screening the plurality of candidate sequences according to the screening rule to obtain at least one target candidate sequence comprises: Determine, according to the screening range, at least one candidate sequence to be eliminated from the plurality of candidate sequences; the candidate sequence to be eliminated is a candidate sequence with relatively low confidence among all candidate sequences; All the eliminated candidate sequences are eliminated, and the remaining candidate sequences are used as the target candidate sequences.

4. The method according to claim 1, characterized in that: The step of generating a reply text according to all the target candidate sequences comprises: Determining whether the contents of all the target candidate sequences are the same; If all are the same, generating a reply text according to the target candidate sequence; If they are at least partially different, at least two different target candidate sequences are screened, and a reply text is generated based on the screening result.

5. The method according to claim 4, characterized in that The step of determining whether the contents of all the target candidate sequences are the same comprises: Determine the demand type corresponding to the prompt word; If the requirement type is a result type, judging whether the contents of all the target candidate sequences are the same according to the results of all the target candidate sequences; If the requirement type is a non-result type, a content identification module is called to identify all the target candidate sequences to obtain the contents of all the target candidate sequences; and it is determined whether the contents of all the target candidate sequences are the same.

6. A reasoning device based on a large model, characterized in that: include: Generate module, filter module and output module; The generation module is used to obtain multiple candidate sequences based on the input prompt word; Each candidate sequence has a corresponding sequence entropy value, and each of the sequence entropy values ​​is used to characterize the diversity and accuracy of the corresponding candidate sequence; the sequence entropy value is composed of the word-unit entropy values ​​of all the word-units in the corresponding candidate sequence; the word-unit entropy values ​​of all the word-units in the same candidate sequence are generated sequentially based on the decreasing strategy of the temperature coefficient; The screening module is used to screen the multiple candidate sequences according to a screening rule to obtain at least one target candidate sequence; the screening rule is used to indicate to eliminate candidate sequences with low confidence among the multiple candidate sequences; The output module is used to generate a reply text based on all the target candidate sequences.

7. The device according to claim 6, characterized in that The generation module is specifically used to: based on the prompt word, at each generation moment, determine the word element corresponding to the generation moment according to the current temperature coefficient of each candidate sequence; wherein, corresponding to each generation moment, the value of the current temperature coefficient of the candidate sequence decreases successively according to the decreasing strategy; and construct a corresponding candidate sequence according to the word elements corresponding to all the generation moments.

8. An electronic device, characterized in that: include: A memory for storing one or more programs; processor; When the one or more programs are executed by the processor, the method according to any one of claims 1 to 5 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

10. A program product, characterized in that When the program product is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Neural network classification model knowledge distillation method for passive domain data

    CN115456166A

  • Multi-modal medical data generation method and related device

    CN119830218A

  • Self-Improving LLMs through Consistency-Based Self-Generated Demonstrations

    US20240249080A1

  • Contrastive learning for implicit neural representations

    US20250111281A1

Cited By

  • Model updating method, electronic equipment, storage medium and program product

    CN120406989A

  • Model updating method, electronic device, storage medium and program product

    CN120406989B