Intelligent reasoning method and system based on global search in combination with external information

By combining global search and intelligent inference methods, multi-step inference in the existing technology is solved, and the problem of local optimality and insufficient external information call is easily trapped, achieving higher inference accuracy and robustness.

CN120068946AActive Publication Date: 2025-05-30HANGZHOU COLOSSEUM DIGITAL INTELLIGENT TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510525553.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing intelligent inference methods are prone to falling into local optimization when multi-step inference, insufficient external information calls, resulting in poor inference effects, and difficult to adapt to the ambiguity and uncertainty in natural language generation, and it is difficult to achieve the optimal balance of efficiency and quality.

Method used

An intelligent reasoning method based on global search and combined with external information is adopted to generate candidate actions through a large language model, calculate the external support compensation weight of candidate actions, calculate the immediate benefits based on local generation probability and external support compensation weight, and calculate the comprehensive selection value through the upper limit confidence interval algorithm, and dynamically adjust the exploration intensity to achieve the global optimal reasoning chain.

Benefits of technology

The deep integration of local generation and global search is achieved, the accuracy, coherence and systematic robustness of multi-step inference tasks are improved, and the information insufficient and error accumulation problems caused by local confidence of a single relying on large language models is overcome, ensuring efficiency and quality in the search process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068946A_ABST
    Figure CN120068946A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an intelligent reasoning method and system based on global search in combination with external information, and the method comprises the following steps: S1, generating candidate actions in a current state through a pre-trained large language model, synchronously outputting a local generation probability, retrieving external support data according to a current state, and calculating an external support compensation weight of the candidate action; s2, calculating the real-time income of the candidate action through the local generation probability and the external support compensation weight of the candidate action; s3, calculating a comprehensive selection value of the candidate action through an upper limit confidence interval algorithm; and S4, outputting a global optimal reasoning chain. According to the method, deep fusion of local generation and global search can be realized, the aspects of candidate action evaluation and external information dynamic integration are remarkably improved, and the accuracy, coherence and system robustness of a multi-step reasoning task are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an intelligent reasoning method and system based on global search combined with external information. Background Art

[0002] In recent years, large language models have made breakthroughs in the fields of natural language understanding and generation. For example, the excellent performance of GPT series and BERT models in tasks such as question answering, text generation, and translation has been widely verified. However, when dealing with multi-step complex reasoning, constructing problem chains, or cross-domain knowledge invocation, relying solely on the local generation ability of large language models is prone to falling into local optima, and the long-range dependence problem often leads to discontinuous reasoning chains and cumulative errors. In addition, large language models have limitations in knowledge representation and external information invocation. When the internal knowledge is insufficient or facing cross-domain problems, simply relying on model generation often cannot cover all necessary information.

[0003] On the other hand, Monte Carlo tree search, as a global search algorithm based on random simulation and statistical feedback, has shown extremely high effectiveness in decision-making problems such as Go and chess, such as the successful implementation of AlphaGo. However, traditional Monte Carlo tree search methods mainly rely on fixed reward feedback and static exploration parameters, making it difficult to fully handle the uncertainty and ambiguity in natural language generation. Moreover, when dealing with complex multi-step logical reasoning, a single search mechanism often ignores the rich semantics and context relationships contained in the information generated by large language models.

[0004] Currently, some existing technologies attempt to combine the candidate steps generated by large language models with traditional search algorithms, such as the chain of thought and self-consistency strategies, to improve reasoning accuracy. However, these methods usually use fixed thresholds or static weights to determine external knowledge invocation, and fail to dynamically adjust according to the uncertainty expressed by candidate actions, thus lacking a flexible adaptive mechanism in the process of knowledge supplementation and global optimization, which limits the deep integration between local generation and global search. Summary of the Invention

[0005] The technical problem to be solved by the present invention: Existing intelligent reasoning methods mainly rely on the confidence of the model's own local generation, and are prone to falling into local optima during multi-step reasoning. At the same time, the external information invocation is insufficient, and the reasoning effect is poor. Moreover, due to the difficulty of the existing global optimization mechanism in adapting to the ambiguity and uncertainty in natural language generation, the efficiency and quality are difficult to achieve the best balance during the search process.

[0006] To solve the above technical problems, the first aspect of the present invention adopts the following technical solution: An intelligent reasoning method based on global search combined with external information, comprising the following steps: S1: Use a pre-trained large language model to generate candidate actions in the current state, synchronously output local generation probabilities, retrieve external support data according to the current state, and calculate the external support compensation weights of the candidate actions; S2: Calculate the immediate rewards of the candidate actions through the local generation probabilities and external support compensation weights of the candidate actions; S3: Calculate the comprehensive selection values of the candidate actions through the Upper Confidence Bound (UCB) algorithm, where a dynamic exploration constant is generated based on the local generation probabilities and / or external support compensation weights of the candidate actions to correct the exploration term in real time; S4: Select the candidate action with the highest comprehensive selection value, update the current state, and transfer to step S1 until the global optimal inference chain is output.

[0007] When the present invention works, it can achieve a deep integration of local generation and global search, with significant improvements in candidate action evaluation and dynamic integration of external information, greatly improving the accuracy, coherence, and system robustness of multi-step inference tasks. Through the fusion mechanism that combines external support compensation weights, it overcomes the problems of insufficient information and error accumulation caused by solely relying on the local confidence of the large language model, and can automatically adjust the exploration intensity based on the generation quality of the current candidate step and the support of external information, ensuring the efficiency and quality during the search process.

[0008] Preferably, in step S1, the following steps are further included: dynamically calculate the uncertainty index of the candidate action. When the uncertainty index of the candidate action is greater than a preset threshold, it is determined that the internal generated information of the large language model is insufficient, and external support data is actively called to supplement the semantic information of the current state.

[0009] When the present invention works, by dynamically calculating the uncertainty index of the candidate action, it can provide real-time feedback and adjustment for the invocation of external information, ensuring the supplementation of external information during the generation process. When the confidence level of the large language model generation is low, the semantics are ambiguous, or there is a shortage of cross-domain knowledge, it can timely and accurately supplement the necessary external knowledge and correct the generation confidence, thereby enhancing the global consistency and stability of the entire inference chain.

[0010] Preferably, when retrieving external support data according to the current state and calculating the external support compensation weights of the candidate actions in step S1, the following steps are specifically adopted: A1: Use an encoder to map the current state and the candidate action into embedding vectors, retrieve external support data according to the current state, and obtain support vectors; A2: Calculate the similarity between the embedding vectors and the support vectors, and normalize the similarity; A3: Obtain the statistical characteristics of the normalized similarity in the candidate set, combine its mean and standard deviation, perform non-linear scaling on the normalized similarity, calculate the external support compensation weight, and output the external support compensation weight for adjusting the local generation probability of the output of the large language model.

[0011] When the present invention works, through the non-linear scaling model, the external support compensation weight of the external information is calculated according to the statistical distribution of the candidate set, so as to dynamically adjust the locally generated information. Through the multi-dimensional evaluation method, the problems of error accumulation and insufficient robustness that may occur due to single dependence on local information can be effectively overcome.

[0012] Preferably, in the step A1, when retrieving the external support data according to the current state to obtain the support vector, the following steps are adopted: measure the frequency of the words and sentences appearing in the current state, and correct it by measuring the rarity of the words and sentences in the external support data, extract the keywords in the current state, and at the same time, by constructing the dependency tree of the current state, supplement the keywords that have not been extracted before, integrate and output the keyword set, query the external support data through parallel processing, and merge the query results to output the final support vector.

[0013] When the present invention works, by constructing the dependency tree, the keywords that have not been extracted by the frequency statistics can be supplemented, which improves the comprehensiveness and accuracy of keyword extraction. At the same time, by querying the external support data through parallel processing, the query process of the external support data is accelerated, and the generation efficiency and generation quality of the large language model can be further improved.

[0014] Preferably, in the step A1, the following steps are further included: cooperate with the multi-head attention mechanism to perform fine-grained analysis on the candidate actions, and improve the matching degree between the output candidate actions and the actual semantic requirements.

[0015] Preferably, in the step A2, when calculating the similarity between the embedding vector and the support vector, the following steps are adopted: calculate the cosine similarity between the embedding vector and the support vector on several attention heads, calculate the multi-head attention weight through weighted summation, and output the similarity between the embedding vector and the support vector.

[0016] When the present invention works, the multi-head attention mechanism is adopted to perform fine-grained analysis on the candidate actions, combining global feature matching and importance assignment in the local subspace, enabling the large language model to pay attention to the similarity in different semantic spaces, and enhancing the robustness of similarity calculation through context-aware similarity enhancement.

[0017] Preferably, in step A3, when obtaining the statistical characteristics of the normalized similarity in the candidate set, combining its mean and standard deviation, performing non-linear scaling on the normalized similarity, and calculating the external support compensation weight for adjusting the local generation probability of the output of the large language model, the following steps are adopted: B1: Obtain the normalized similarity of the candidate actions in the candidate set, and calculate the mean and standard deviation of the normalized similarity; B2: Convert the normalized similarity of each candidate action into a standardized score; B3: Combine the local generation probability of the candidate action and a preset hyperparameter to set a dynamic adjustment coefficient, and correct the standardized score of the candidate action through the dynamic adjustment coefficient; B4: Map the standardized score of the candidate action to a smooth non-linear interval, and after correcting according to the local generation probability of the candidate action, output the external support compensation weight for adjusting the local generation probability of the output of the large language model.

[0018] Preferably, in step S3, when calculating the comprehensive selection value of the candidate action through the upper confidence bound algorithm, where a dynamic exploration constant is generated according to the local generation probability and / or external support compensation weight of the candidate action to correct the exploration term in real time, the following steps are adopted. Calculate the comprehensive selection value of the candidate action through the upper confidence bound algorithm, set the dynamic exploration constant, calculate and generate the dynamic exploration constant according to the local generation probability and / or external support compensation weight of the candidate action, and increase the dynamic exploration constant when the local generation probability of the candidate action is lower than a preset threshold or the external support compensation weight is lower than a preset threshold.

[0019] When the present invention works, by setting a dynamically changing dynamic exploration constant when calculating the comprehensive selection value of the candidate action through the upper confidence bound algorithm, further automatic tuning of the reinforcement learning method is achieved. When facing candidate actions with insufficient local confidence or insufficient external support, the exploration probability can be automatically increased, thereby avoiding the situation of falling into a local optimum and being unable to globally search for the best inference chain.

[0020] Preferably, in step S4, when selecting the candidate action with the highest comprehensive selection value, updating the current state, and transferring to step S1 until the global optimal inference chain is output, the following steps are adopted. Select the candidate action with the highest comprehensive selection value, update the current state, and transfer to step S1. Repeatedly call Monte Carlo tree search, obtain the cumulative reward of each iteration to update the statistical data of each path under the root node. After the iteration ends, extract and output the global optimal inference chain according to the search tree statistics.

[0021] To solve the above technical problems, the second aspect of the present invention adopts the following technical solution. An intelligent reasoning system based on global search combined with external information applies an intelligent reasoning method based on global search combined with external information as described above, and includes: A large language model module for deploying a large language model to generate candidate actions according to the current state; A Monte Carlo tree search module for constructing an inference state tree and outputting a globally optimal inference chain by balancing local generation and global exploration; An external information module for providing external support data; The large language model module, the Monte Carlo tree search module, and the external information module are connected by data. The large language model generates a number of candidate actions to the Monte Carlo tree search module. The Monte Carlo tree search module outputs a globally optimal inference chain by balancing local generation and global exploration. When the information generated inside the large language model module is insufficient, the external information module outputs external support data to the large language model module. When the Monte Carlo tree search module balances local generation and global search, the external information module outputs external support data to the Monte Carlo tree search module.

[0022] When the present invention works, a closed-loop feedback system is constructed through the mutual cooperation of the large language model module, the Monte Carlo tree search module, and the external information module, which efficiently links the generation of candidate actions by the large language model, the global search of the Monte Carlo tree, and the automatic invocation of external support data. During each step of generation and search, the resource invocation and parameter settings are automatically and dynamically adjusted, realizing the real-time nature of information transmission and the optimization of the collaborative work of each module, greatly improving the overall inference accuracy, coherence, and robustness in multi-task scenarios.

[0023] The beneficial technical effects of the present invention include: 1. The present invention can achieve a deep integration of local generation and global search, with significant improvements in candidate action evaluation and dynamic integration of external information, greatly improving the accuracy, coherence, and system robustness of multi-step inference tasks. By combining the fusion mechanism of external support compensation weights, it overcomes the problems of insufficient information and error accumulation caused by solely relying on the local confidence of the large language model, and can automatically adjust the exploration intensity according to the generation quality of the current candidate step and the support of external information, ensuring the efficiency and quality during the search process.

[0024] 2. By dynamically calculating the uncertainty index of candidate actions, the present invention can provide real-time feedback and adjustment for the invocation of external information, ensuring the supplementation of external information during the generation process. When the confidence level generated by the large language model is low, the semantics are ambiguous, or there is a shortage of cross-domain knowledge, it can timely and accurately supplement the necessary external knowledge and correct the generation confidence, thereby enhancing the global consistency and stability of the entire inference chain.

[0025] 3. The present invention calculates the external support compensation weight of external information according to the statistical distribution of the candidate set through a non-linear scaling model, thereby dynamically adjusting the locally generated information. Through a multi-dimensional evaluation method, it can effectively overcome the problems of error accumulation and insufficient robustness that may occur due to single reliance on local information.

[0026] 4. By constructing a dependency tree, the present invention can supplement keywords not extracted by frequency statistics, improving the comprehensiveness and accuracy of keyword extraction. At the same time, by parallel processing to query external support data, the query process of external support data is accelerated, which can further improve the generation efficiency and generation quality of the large language model.

[0027] 5. The present invention uses a multi-head attention mechanism to perform fine-grained analysis on candidate actions, combining global feature matching and importance assignment in the local subspace, enabling the large language model to focus on the similarities in different semantic spaces. Through context-aware similarity enhancement, the robustness of similarity calculation is improved.

[0028] 6. When calculating the comprehensive selection value of candidate actions by the upper confidence interval algorithm, the present invention sets a dynamically changing dynamic exploration constant, realizing further automatic tuning of the reinforcement learning method. When facing candidate actions with insufficient local confidence or insufficient external support, it can automatically increase their exploration probability, thus avoiding the situation of being trapped in a local optimum and unable to globally search for the best inference chain.

[0029] 7. The present invention constructs a closed-loop feedback system through the mutual cooperation of the large language model module, Monte Carlo tree search module, and external information module, efficiently linking the generation of candidate actions by the large language model, global search of the Monte Carlo tree, and automatic invocation of external support data. In each step of the generation and search process, it automatically and dynamically adjusts resource invocation and parameter settings, realizing the real-time nature of information transmission and the optimization of the collaborative work of each module, greatly improving the overall inference accuracy, coherence, and robustness in multi-task scenarios.

[0030] Other features and advantages of the present invention will be disclosed in detail in the following specific embodiments and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The following further describes the present invention with reference to the drawings: Figure 1 It is a flowchart of the working process of an intelligent inference method based on global search combined with external information; Figure 2 It is a flowchart of the working process of step S1 in an intelligent inference method based on global search combined with external information; Figure 3It is a flowchart of step A3 in an intelligent reasoning method based on global search combined with external information; Figure 4 It is a schematic structural diagram of an intelligent reasoning system based on global search combined with external information. Detailed implementation manners

[0032] The technical solutions of the embodiments of the present invention will be explained and described below with reference to the accompanying drawings of the embodiments of the present invention. However, the following embodiments are only the preferred embodiments of the present invention, not all of them. Based on the embodiments in the implementation manners, other embodiments obtained by those skilled in the art without creative efforts all fall within the protection scope of the present invention.

[0033] In the following description, terms such as "inner", "outer", "upper", "lower", "left", "right", etc. indicating orientation or positional relationship are only for the convenience of describing the embodiments and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention. Embodiment 1:

[0034] Please refer to Figure 1 , this embodiment discloses an intelligent reasoning method based on global search combined with external information, including the following steps: S1: Use a pre-trained large language model to generate candidate actions in the current state, synchronously output local generation probabilities, retrieve external support data according to the current state, and calculate the external support compensation weights of the candidate actions; S2: Calculate the immediate rewards of the candidate actions through the local generation probabilities and external support compensation weights of the candidate actions; S3: Calculate the comprehensive selection values of the candidate actions through the upper confidence bound algorithm, where a dynamic exploration constant is generated according to the local generation probabilities and / or external support compensation weights of the candidate actions to correct the exploration term in real time; S4: Select the candidate action with the highest comprehensive selection value, update the current state, and transfer to step S1 until the global optimal inference chain is output.

[0035] When this embodiment works, it can achieve a deep integration of local generation and global search, with significant improvements in candidate action evaluation and dynamic integration of external information, greatly improving the accuracy, coherence, and system robustness of multi-step reasoning tasks. By combining the fusion mechanism of external support compensation weights, it overcomes the information insufficiency and error accumulation problems caused by solely relying on the local confidence of the large language model, and can automatically adjust the exploration intensity according to the generation quality of the current candidate step and the support situation of external information, ensuring the efficiency and quality during the search process.

[0036] Please refer to Figure 2, Preferably, in the step S1, when retrieving external support data according to the current state and calculating the external support compensation weight of the candidate action, the following steps are specifically adopted: A1: Use the encoder to map the current state and the candidate action into embedding vectors, retrieve external support data according to the current state, and obtain support vectors; A2: Calculate the similarity between the embedding vector and the support vector, and normalize the similarity; A3: Obtain the statistical characteristics of the normalized similarity in the candidate set, combine its mean and standard deviation, perform non-linear scaling on the normalized similarity, calculate the external support compensation weight, and output the external support compensation weight for adjusting the local generation probability of the output of the large language model.

[0037] When this embodiment works, the external support compensation weight of the external information is calculated according to the statistical distribution of the candidate set through the non-linear scaling model, so as to dynamically adjust the locally generated information. Through the multi-dimensional evaluation method, the problem of error accumulation and insufficient robustness that may occur due to single dependence on local information can be effectively overcome.

[0038] In this embodiment, in the step A1, when retrieving external support data according to the current state and obtaining support vectors, the following steps are adopted: measure the frequency of the words and sentences appearing in the current state, and correct it by measuring the rarity of the words and sentences in the external support data, extract the keywords in the current state. In specific implementation, the important keywords in the current state can be extracted by the hierarchical keyword extraction method. At the same time, by constructing the dependency tree of the current state, the keywords that have not been extracted before are supplemented, and the keyword set is integrated and output. Query the external support data through parallel processing, and merge the query results to output the final support vector.

[0039] When this embodiment works, by constructing the dependency tree, the keywords that have not been extracted by frequency statistics can be supplemented, improving the comprehensiveness and accuracy of keyword extraction. At the same time, by querying the external support data through parallel processing, the query process of the external support data is accelerated, which can further improve the generation efficiency and generation quality of the large language model.

[0040] As a further improvement of this embodiment, in the step A1, the following steps are also included: cooperate with the multi-head attention mechanism to perform fine-grained analysis on the candidate action, and improve the matching degree between the output candidate action and the actual semantic requirements.

[0041] Preferably, in the step A2, when calculating the similarity between the embedding vector and the support vector, the following steps are adopted: calculate the cosine similarity between the embedding vector and the support vector on several attention heads, calculate the multi-head attention weight through weighted summation, and output the similarity between the embedding vector and the support vector.

[0042] When this embodiment works, the multi-head attention mechanism is used to perform fine-grained analysis on candidate actions, combining global feature matching and importance assignment in the local subspace, enabling the large language model to focus on the similarities in different semantic spaces. Through context-aware similarity enhancement, the robustness of similarity calculation is improved.

[0043] In specific implementation, in the step A2, when calculating the similarity between the embedding vector and the support vector and normalizing the similarity, the following calculation formula is used: ; Where, is the cosine similarity between the embedding vector and the support vector, is the embedding vector obtained by mapping the state s and the candidate action a through the encoder, is the support vector retrieved from external support data, is the normalized similarity between the embedding vector and the support vector.

[0044] In specific implementation, in the step S1, the following steps are further included: dynamically calculating the uncertainty index of the candidate action. When the uncertainty index of the candidate action is greater than the preset threshold, it is determined that the internal generated information of the large language model is insufficient, and external support data is actively called to supplement the semantic information of the current state.

[0045] When this embodiment works, by dynamically calculating the uncertainty index of the candidate action, real-time feedback and adjustment of the call of external information can be carried out to ensure the supplement of external information during the generation process. When the confidence of the generation by the large language model is low, the semantics are ambiguous, or there is a shortage of cross-domain knowledge, necessary external knowledge can be supplemented in a timely and accurate manner and the generation confidence can be corrected, thereby improving the global consistency and stability of the entire inference chain.

[0046] As a further improvement of this example, the following formula is used to calculate the uncertainty index of the candidate action: ; Where: is the uncertainty index of the candidate action. When the uncertainty index of the candidate action exceeds the preset threshold, the re-call and supplement of external information are triggered to ensure the integrity of information during the inference process. Preferably, this threshold can also be further dynamically adjusted according to the generation effect to increase or suppress the call frequency of external support data, thereby maintaining the balance between generation quality and generation efficiency. is the local generation probability of the candidate action, is the logits of the large language model generating the action a in the state s. Embodiment 2:

[0047] Please refer toFigure 3 , this embodiment provides an intelligent reasoning method based on global search combined with external information. The same parts as other embodiments will not be elaborated here, and the differences will be described in detail below.

[0048] In this embodiment, in step A3, when obtaining the statistical characteristics of the normalized similarity in the candidate set, combining its mean and standard deviation, performing non-linear scaling on the normalized similarity, and calculating the external support compensation weight for adjusting the local generation probability of the large language model output, the following steps are adopted: B1: Obtain the normalized similarity of the candidate actions in the candidate set, and calculate the mean and standard deviation of the normalized similarity; B2: Convert the normalized similarity of each candidate action into a standardized score; B3: Combine the local generation probability of the candidate action and a preset hyperparameter to set a dynamic adjustment coefficient, and correct the standardized score of the candidate action through the dynamic adjustment coefficient; B4: Map the standardized score of the candidate action to a smooth non-linear interval, and after correcting according to the local generation probability of the candidate action, output the external support compensation weight for adjusting the local generation probability of the large language model output.

[0049] In step B1, when obtaining the normalized similarity of the candidate actions in the candidate set and calculating the mean and standard deviation of the normalized similarity, the following formulas are adopted: ; Where: is the mean of the candidate set, used to measure the similarity central tendency of all candidate actions, N is the number of candidate actions in the candidate set, is the standard deviation of the candidate set, used to quantify the dispersion degree of the similarity distribution.

[0050] In step B2, when converting the normalized similarity of each candidate action into a standardized score, the following calculation formula is adopted: ; Where: is the standardized score, is the minimum value, used to prevent the denominator from being zero.

[0051] In step B3, when combining the local generation probability of the candidate action and a preset hyperparameter to set a dynamic adjustment coefficient, the following calculation formula is adopted: ; Where: is a dynamic adjustment coefficient, base_gamma is a preset hyperparameter used to control the initial intensity. By setting the adjustment of the local generation probability with the dynamic adjustment coefficient, the influence degree of the normalized similarity on the external support compensation weight is amplified or suppressed. When the large language model is more confident in the candidate action and the local generation probability of the candidate action is higher, the reward for high similarity is weakened to avoid over-reliance on a single indicator. On the contrary, when the local generation probability of the candidate action becomes smaller, it is necessary to rely on similarity judgment to increase the invocation frequency of external support data, so as to dynamically balance the confidence of the large language model and the matching degree of external support data in complex scenarios.

[0052] In the step B4, when mapping the normalized score of the candidate action to a smooth non-linear interval and correcting it according to the local generation probability of the candidate action, and then outputting the external support compensation weight for adjusting the local generation probability of the large language model output, the following formula is used: ; Where: is the external support compensation weight, is a non-linear scaling model, is a decay factor. When the local generation probability increases, the amplitude of weight adjustment is reduced, and the invocation of external support data is reduced through conservative adjustment to improve the generation efficiency.

[0053] Preferably, in the step S2, when calculating the immediate benefit of the candidate action through the local generation probability of the candidate action and the external support compensation weight, the following calculation formula is used: ; Where: is the immediate benefit of the candidate action. Embodiment 3:

[0054] This embodiment provides an intelligent reasoning method based on global search combined with external information. The same parts as other embodiments will not be described in detail, and the differences will be described in detail below.

[0055] In this embodiment, in the step S3, the comprehensive selection value of the candidate action is calculated by the upper confidence bound algorithm. When generating a dynamic exploration constant according to the local generation probability of the candidate action and / or the external support compensation weight to correct the exploration term in real time, the following steps are adopted. The comprehensive selection value of the candidate action is calculated by the upper confidence bound algorithm, a dynamic exploration constant is set, and the dynamic exploration constant is calculated according to the local generation probability of the candidate action and / or the external support compensation weight. When the local generation probability of the candidate action is lower than a preset threshold or the external support compensation weight is lower than a preset threshold, the dynamic exploration constant is increased.

[0056] When this embodiment works, by setting a dynamically changing dynamic exploration constant when calculating the comprehensive selection value of candidate actions through the upper confidence bound algorithm, further automatic tuning of the reinforcement learning method is achieved. When facing candidate actions with insufficient local confidence or insufficient external support, the exploration probability can be automatically increased, thereby avoiding the situation of being trapped in a local optimum and unable to globally search for the best inference chain.

[0057] In specific implementation, when calculating the comprehensive selection value of candidate actions through the upper confidence bound algorithm, and generating a dynamic exploration constant according to the local generation probability and / or external support compensation weight of the candidate action to correct the exploration term in real time, the following formula is adopted: ; Where: is the comprehensive selection value, is the average reward, is the number of visits, is the number of visits under candidate action a, is the dynamic exploration constant. When the local generation probability of the candidate action is low or the external information compensation is insufficient, the dynamic exploration constant automatically increases, promoting more exploration of this candidate action, thereby avoiding being trapped in a local optimum. Embodiment Four:

[0058] This embodiment provides an intelligent inference method based on global search combined with external information. The same parts as other embodiments will not be elaborated here, and the differences will be described in detail below.

[0059] In this embodiment, in step S4, when selecting the candidate action with the highest comprehensive selection value, updating the current state, and transferring to step S1 until the global optimal inference chain is output, the following steps are adopted: select the candidate action with the highest comprehensive selection value, update the current state, transfer to step S1, repeatedly call Monte Carlo tree search, obtain the cumulative reward of each iteration to update the statistical data of each path under the root node, and after the iteration ends, extract and output the global optimal inference chain according to the search tree statistical results.

[0060] In specific implementation, the main process of an intelligent inference method based on global search combined with external information is shown by the following pseudocode: Input: Problem P, pre-trained large model LLM, external information module ExternalModule; Output: Optimal inference chain S*; / / Initialize the root state s0 (i.e., problem P); s0 ← P; N(s0) ← 1; Function MCTS(s): if Terminal(s) or depth(s) ≥ T_max: return Evaluate(s); for each candidate action a ∈ GenerateCandidates(LLM, s): / / 1. Use the LLM module to generate candidate actions and calculate the local generation probability pLLM; pLLM ← Softmax(LLM output logits(s, a)); / / 2. Use the encoder to generate the state-action embedding E_sa: E_sa ← Encoder(s, a); / / 3. Obtain the support vector r_ext from the external information module through hierarchical keyword extraction and cache optimization: keywords ← TF-IDF weighted extraction(s) + dependency parsing extraction(s); r_ext ← ExternalModule(keywords, using parallel processing); / / 4. Calculate the multi-head attention weighted cosine similarity: attn_weights ← MultiHeadAttention(E_sa, r_ext); cos_sim ← dot(E_sa, r_ext) / (||E_sa|| * ||r_ext||); weighted_cos_sim ← Σ(attn_weights_i * partial_cos_sim_i); cos_sim_norm ← clip((weighted_cos_sim + 1) / 2, 0, 1); / / 5. Adaptive non-linear scaling to calculate the compensation weight: mean_sim ← calculate the mean of all candidate cos_sim_norm; std_sim ← calculate the standard deviation of all candidate cos_sim_norm; z_score ← (cos_sim_norm - mean_sim) / (std_sim + ε); gamma ← base_gamma * (1 - (1 - pLLM)); beta ← sigmoid(gamma * z_score) * (1 - log(1 + pLLM) / log(2)) + baseline; / / 6. Calculate the immediate reward R of the candidate action (combined by the LLM confidence and the external compensation weight): R ← pLLM + beta; / / 7. Calculate the dynamic exploration constant alpha: Combine the square root of (1 - pLLM) and (1 - cos_sim_norm), and tune it through reinforcement learning: alpha ← sqrt((1 - pLLM) + (1 - cos_sim_norm)); / / 8. Calculate the exploration term U and the comprehensive selection value F according to the improved UCT strategy: U ← alpha * sqrt(ln(N(s)) / (N(s, a) + 1)); F ← Q(s, a) + U; Select a* = argmax_a F(s, a); / / Expand to generate the next state s' = s ∪ {a*}: if s' is not in the tree: Initialize s', N(s') ← 1, Q(s') ← R(s, a*); else: s' ← the existing node; / / Simulation stage (Rollout): Perform a complete inference starting from state s', and accumulate the reward as R_total: R_total ← Rollout(s'); / / Backtracking update: Update the statistical data and Q values of each node and action along the path from s0 to s'; for each state-action pair (s, a) in the path from s0 to s': Q(s, a) ← (Q(s, a) * N(s, a) + R_total) / (N(s, a) + 1); N(s, a) ← N(s, a) + 1, N(s) ← N(s) + 1; return R_total; function Rollout(s): R_sum ← 0, t ← 0; while not Terminal(s) and t < T_max: a ← GenerateCandidate(LLM, s); pLLM ← Softmax(logits(s, a) output by LLM); E_sa ← Encoder(s, a); r_ext ← ExternalModule(RetrieveKeywords(s)); cos_sim ← CosineSimilarity(E_sa, r_ext); cos_sim_norm ← (cos_sim + 1) / 2; / / Calculate the uncertainty metric delta (combining pLLM and cos_sim_norm): delta ← 1 - (pLLM + cos_sim_norm) / 2; if delta > τ: r_ext ← ExternalModule(RetrieveKeywords(s)); cos_sim ← CosineSimilarity(Encoder(s, a), r_ext); cos_sim_norm ← (cos_sim + 1) / 2; beta ← NonLinearScaling(cos_sim_norm, current candidate average similarity, standard deviation, ε); R ← pLLM + beta; R_sum ← R_sum + R; / / Update the state: add the currently selected action a: s ← s ∪ {a}; t ← t + 1; return R_sum; / / Main loop: repeatedly call MCTS(s0) to continuously update the search tree; for iteration = 1 to N_iter:; R_iter ← MCTS(s0); / / The cumulative reward R_iter for each iteration updates the statistics of each path under the root node; / / After the iteration, extract the optimal inference chain according to the search tree statistics; S* ← ExtractBestPath(s0); return S*。 Example Five:

[0061] Please refer to Figure 4 , this example provides an intelligent inference system based on global search combined with external information, applying an intelligent inference method based on global search combined with external information as described in the above example, including: Large language model module, used to deploy a large language model to generate candidate actions according to the current state; Monte Carlo tree search module, used to construct an inference state tree and output a globally optimal inference chain by balancing local generation and global exploration; External information module, used to provide external support data; The large language model module, Monte Carlo tree search module, and external information module are connected by data. The large language model generates several candidate actions to the Monte Carlo tree search module. The Monte Carlo tree search module outputs a globally optimal inference chain by balancing local generation and global exploration. When the information generated inside the large language model module is insufficient, the external information module outputs external support data to the large language model module. When the Monte Carlo tree search module balances local generation and global search, the external information module outputs external support data to the Monte Carlo tree search module.

[0062] When this embodiment works, a closed-loop feedback system is constructed through the mutual cooperation of the large language model module, Monte Carlo tree search module, and external information module, efficiently linking the generation of candidate actions by the large language model, global search of the Monte Carlo tree, and automatic invocation of external support data. During each step of generation and search, the resource invocation and parameter settings are automatically and dynamically adjusted, realizing the real-time nature of information transmission and the optimization of the collaborative work of each module, greatly improving the overall inference accuracy, coherence, and robustness in multi-task scenarios.

[0063] The beneficial technical effects of this embodiment include: The present invention can achieve a deep integration of local generation and global search, with significant improvements in candidate action evaluation and dynamic integration of external information, greatly improving the accuracy, coherence, and system robustness of multi-step inference tasks. By combining the fusion mechanism of external support compensation weights, it overcomes the problems of insufficient information and error accumulation caused by solely relying on the local confidence of the large language model, and can automatically adjust the exploration intensity according to the generation quality of the current candidate step and the support of external information, ensuring the efficiency and quality during the search process.

[0064] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that the present invention includes but is not limited to the content described in the drawings and the above specific embodiments. Any modification that does not deviate from the functional and structural principles of the present invention will be included in the scope of the claims.

Claims

1. An intelligent reasoning method based on global search combined with external information, characterized in that: The following steps are involved: S1: Use the pre-trained large language model to generate candidate actions in the current state, synchronously output the local generation probability, retrieve external support data according to the current state, and calculate the external support compensation weight of the candidate action; S2: Calculate the immediate benefit of the candidate action through the local generation probability of the candidate action and the external support compensation weight; S3: calculating the comprehensive selection value of the candidate action by an upper confidence interval algorithm, wherein a dynamic exploration constant is generated according to the local generation probability of the candidate action and / or the external support compensation weight to modify the exploration item in real time; S4: Select the candidate action with the highest comprehensive selection value, update the current state, and go to step S1 until the global optimal reasoning chain is output.

2. According to claim 1, the intelligent reasoning method based on global search combined with external information is characterized by: In the step S1, the following steps are also included: dynamically calculating the uncertainty index of the candidate action; when the uncertainty index of the candidate action is greater than a preset threshold, it is determined that insufficient information is generated internally in the large language model, and external support data is actively called to supplement the semantic information of the current state.

3. The intelligent reasoning method based on global search combined with external information according to claim 1, characterized in that: In step S1, when retrieving external support data according to the current state and calculating the external support compensation weight of the candidate action, the following steps are specifically adopted: A1: Use the encoder to map the current state and candidate actions into embedding vectors, retrieve external support data based on the current state, and obtain the support vector; A2: Calculate the similarity between the embedding vector and the support vector, and normalize the similarity; A3: Obtain the statistical characteristics of the normalized similarity in the candidate set, combine its mean and standard deviation, perform nonlinear scaling on the normalized similarity, calculate the external support compensation weight, and output the external support compensation weight used to adjust the local generation probability output by the large language model.

4. The intelligent reasoning method based on global search combined with external information according to claim 3 is characterized by: In step A1, external support data is retrieved according to the current state. When the support vector is obtained, the following steps are adopted to measure the frequency of occurrence of words and sentences in the current state, and to make corrections by measuring the rarity of words and sentences in the external support data, to extract the keywords in the current state, and at the same time, by constructing a dependency tree of the current state, to supplement the keywords that have not been extracted before, to integrate and output the keyword set, to query the external support data through parallel processing, and to merge the query results to output the final support vector.

5. The intelligent reasoning method based on global search combined with external information according to claim 3 is characterized by: In the step A1, the following steps are also included, in which a fine-grained analysis of candidate actions is performed in conjunction with a multi-head attention mechanism to improve the matching degree between the output candidate actions and actual semantic requirements.

6. The intelligent reasoning method based on global search combined with external information according to claim 5, characterized in that: In step A2, when calculating the similarity between the embedding vector and the support vector, the following steps are adopted to calculate the cosine similarity between the embedding vector and the support vector on several attention heads, calculate the multi-head attention weights by weighted summation, and output the similarity between the embedding vector and the support vector.

7. The intelligent reasoning method based on global search combined with external information according to claim 3 is characterized by: In step A3, the statistical characteristics of the normalized similarity in the candidate set are obtained, and the normalized similarity is nonlinearly scaled in combination with its mean and standard deviation, and the external support compensation weight is calculated. When the external support compensation weight for adjusting the local generation probability output by the large language model is output, the following steps are adopted: B1: Obtain the normalized similarity of the candidate actions in the candidate set, and calculate the mean and standard deviation of the normalized similarity; B2: Convert the normalized similarity of each candidate action into a standardized score; B3: Combine the local generation probability of the candidate action and the preset hyperparameters to set the dynamic adjustment coefficient, and correct the standardized score of the candidate action through the dynamic adjustment coefficient; B4: Map the standardized scores of the candidate actions to a smooth nonlinear interval, and after correcting them according to the local generation probability of the candidate actions, output the external support compensation weight used to adjust the local generation probability output by the large language model.

8. The intelligent reasoning method based on global search combined with external information according to claim 1 is characterized by: In step S3, the comprehensive selection value of the candidate action is calculated by an upper confidence interval algorithm, wherein a dynamic exploration constant is generated based on the local generation probability of the candidate action and / or the external support compensation weight to correct the exploration item in real time. The following steps are adopted: the comprehensive selection value of the candidate action is calculated by an upper confidence interval algorithm, the dynamic exploration constant is set, and the dynamic exploration constant is generated based on the local generation probability of the candidate action and / or the external support compensation weight. When the local generation probability of the candidate action is lower than a preset threshold or the external support compensation weight is lower than a preset threshold, the dynamic exploration constant is increased.

9. The intelligent reasoning method based on global search combined with external information according to claim 1, characterized in that: In step S4, the candidate action with the highest comprehensive selection value is selected, the current state is updated, and the process goes to step S1. When the global optimal reasoning chain is output, the following steps are adopted to select the candidate action with the highest comprehensive selection value, update the current state, and go to step S1. The Monte Carlo tree search is repeatedly called to obtain the cumulative reward of each iteration to update the statistical data of each path under the root node. After the iteration is completed, the global optimal reasoning chain is extracted and output according to the statistical results of the search tree.

10. An intelligent reasoning system based on global search combined with external information, using an intelligent reasoning method based on global search combined with external information as described in any one of claims 1 to 9 above, characterized in that: include: Large language model module, used to deploy a large language model to generate candidate actions based on the current state; Monte Carlo tree search module, used to build the inference state tree and output the global optimal inference chain by balancing local generation and global exploration; External information module, used to provide external support data; The large language model module, the Monte Carlo tree search module and the external information module are data connected. The large language model generates a plurality of candidate actions to the Monte Carlo tree search module. The Monte Carlo tree search module outputs a global optimal reasoning chain by balancing local generation and global exploration. When the information generated inside the large language model module is insufficient, the external information module outputs external support data to the large language module. When the Monte Carlo tree search module balances local generation and global search, the external information module outputs external support data to the Monte Carlo tree search module.

Citation Information

Patent Citations

  • Multi-hop inference knowledge editing method based on model cognitive verification

    CN118364912A

  • Knowledge intensive question reasoning and generating method based on LLM

    CN118798367A

  • Multi-role digital human construction method based on multi-modal graph retrieval enhancement generation

    CN119003746A

  • Intranet and extranet knowledge association retrieval method based on RAG technology

    CN119415711A

  • Revision of and attribution for output of text generation models

    WO2024073087A1