Intelligent agent decision-making method and device, computer readable storage medium and equipment

By introducing a hierarchical recursive memory network and an adaptive termination mechanism, the intelligent agent decision-making method solves the problems of insufficient reasoning depth and low efficiency in complex abstract tasks of the Agentic RAG system, and achieves efficient and accurate decision output.

CN121303373APending Publication Date: 2026-01-09CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511431755.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing Agentic Retrieval Augmentation (AAG) systems suffer from insufficient reasoning depth, shallow semantic understanding, low efficiency, and excessive computational resource consumption when dealing with complex abstract tasks, making it difficult to effectively handle complex multi-step reasoning and deep domain knowledge reasoning.

Method used

A hierarchical recursive memory network (HRM) is used as the inference engine for the intelligent agent. Through the coupling of low-level and high-level modules, hierarchical recursive state evolution is achieved. Combined with an adaptive termination mechanism, the inference depth is dynamically adjusted, and the final internal hidden state is generated before decoding and decision output.

Benefits of technology

It improves the accuracy and efficiency of intelligent agents in making decisions in complex tasks, optimizes the consumption of computing resources, enhances the interpretability and adaptability of the system, and can optimize the use of computing resources while ensuring output quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303373A_ABST
    Figure CN121303373A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an intelligent agent decision-making method, an intelligent agent decision-making device, a computer readable storage medium and electronic device.The intelligent agent decision-making method comprises the steps that a user query is received, and a retrieval document set related to the user query is obtained; performing joint embedding on the user query and retrieval document set to generate context vector representation; inputting the context vector representation into a hierarchical recursive memory network, and performing hierarchical recursive state evolution in a plurality of calculation segments through the hierarchical recursive memory network to generate a final internal hidden state; and decoding the final internal hidden state to generate decision output of the intelligent agent. According to the invention, a deep and adaptive reasoning process can be realized through the hierarchical recursive memory network, and the decision accuracy and reasoning efficiency of an intelligent agent under a complex task are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to an intelligent agent decision-making method, an intelligent agent decision-making device, a computer-readable storage medium, and an electronic device. Background Technology

[0002] Retrieval augmented generation (RAG) technology has become a key technique for improving the performance of large language models (LLMs) when handling knowledge-intensive tasks. By introducing the concept of "agent," agentic retrieval augmented generation (AAG) further expands the capabilities of RAG systems, enabling them to perform more complex planning, multi-round iterative retrieval, and interaction with external tools.

[0003] However, existing proxy-based RAG systems often exhibit insufficient reasoning depth and superficial semantic understanding when faced with highly abstract tasks involving complex abstract characters, deep psychological motivations, or philosophical concepts. Their decision-making processes rely heavily on immediate retrieval and forward generation, lacking a mechanism for continuous integration of contextual information and hierarchical, progressive "deep thinking," making it difficult to capture the multi-dimensional connotations and implicit connections behind abstract entities. Furthermore, to compensate for insufficient reasoning ability, the systems frequently rely on multiple rounds of repeated retrieval and trial-and-error, resulting in low reasoning efficiency and excessive computational resource consumption.

[0004] Therefore, there is an urgent need for an intelligent agent decision-making mechanism that can simulate the hierarchical and recursive understanding of abstract concepts by humans, in order to improve the depth of reasoning and the quality of decision-making in complex abstract tasks.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure. Summary of the Invention

[0006] The purpose of this disclosure is to provide an intelligent agent decision-making method, an intelligent agent decision-making device, a computer-readable storage medium, and an electronic device, thereby overcoming, to at least a certain extent, the technical problems of insufficient reasoning depth and shallow semantic understanding caused by the limitations of related technologies.

[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0008] According to a first aspect of this disclosure, an intelligent agent decision-making method is provided, comprising: Receive user queries and obtain a set of search documents related to the user queries; The user query and retrieved document set are jointly embedded to generate a context vector representation; The context vector representation is input into a hierarchical recursive memory network, and the hierarchical recursive state evolution is performed in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; The final hidden internal state is decoded to generate the decision output of the intelligent agent.

[0009] In an exemplary embodiment of this disclosure, the hierarchical recursive memory network includes coupled low-level modules and high-level modules; The lower-level module updates its state at a first frequency, and the higher-level module updates its state at a second frequency, wherein the second frequency is lower than the first frequency.

[0010] In an exemplary embodiment of this disclosure, the step of generating a final hidden state by performing hierarchical recursive state evolution within multiple computational segments through the hierarchical recursive memory network includes: In each computation segment, the context vector representation is updated in a fine-grained state at the first frequency by the lower-level module; When the low-level updates accumulate to a preset number of times, the high-level module receives the current state of the low-level module at the second frequency and updates the high-level hidden state. The high-level hidden state is recursively passed as a memory carrier across computational segments, and when a preset reasoning termination condition is met, it is determined as the final internal hidden state.

[0011] In an exemplary embodiment of this disclosure, the step of recursively passing the high-level hidden state as a memory carrier across computational segments, and determining it as the final inner hidden state when a preset inference termination condition is met, includes: At the end of each computation segment, obtain the current high-level hidden state corresponding to the high-level module; An action value score is generated based on the current high-level hidden state, and the preset reasoning termination condition is determined based on the action value score. If the preset reasoning termination condition is met, then the current high-level hidden state is determined as the final internal hidden state; If the preset reasoning termination condition is not met, the current high-level hidden state will be used as the initial hidden state of the next computation segment to start a new round of hierarchical recursive state evolution.

[0012] In an exemplary embodiment of this disclosure, generating an action value score based on the current high-level hidden state includes: The current high-level hidden state is input into the decision evaluation head network; The decision evaluation head network outputs predicted value scores for both the continue action and the terminate action; The preset reasoning termination conditions include: the predicted value score corresponding to the termination action is greater than the predicted value score corresponding to the continuation action, or the predicted value score corresponding to the termination action is greater than a preset threshold.

[0013] In an exemplary embodiment of this disclosure, decoding the final hidden state to generate the decision output of the intelligent agent includes: The final hidden internal state is input into one or more output head networks for decoding, and the output head networks generate decision results that are adapted to the current task requirements. The decision result is used to drive the intelligent agent to perform at least one of the following operations: generate a natural language response, invoke a specified external tool, initiate a new round of information retrieval, or generate a structured analysis report.

[0014] In an exemplary embodiment of this disclosure, the structured analysis report includes at least one of the following: The core conclusion represents the optimal decision result derived from the final hidden state when the reasoning process terminates; Confidence assessment is used to quantify the credibility of the reasoning regarding the core conclusions. Supporting evidence and contradictory evidence are respectively identified as information fragments from the retrieved documents obtained in this reasoning that support or refute the core conclusion. Unresolved conflict points are used to indicate semantic contradictions in the current set of retrieved documents that cannot be resolved through evidence fusion.

[0015] According to a second aspect of this disclosure, an intelligent agent decision-making device is provided, comprising: The receiving module is used to receive user queries and obtain a set of search documents related to the user queries; The encoding module is used to jointly embed the user query and retrieved document set to generate a context vector representation; The hierarchical recursive processing module is used to input the context vector representation into the hierarchical recursive memory network, and perform hierarchical recursive state evolution in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; The decision output module is used to decode the final internal hidden state and generate the decision output of the intelligent agent.

[0016] According to a third aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the intelligent agent decision-making method described in the first aspect above.

[0017] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the intelligent agent decision-making method described in the first aspect by executing the executable instructions.

[0018] As can be seen from the above technical solutions, the intelligent agent decision-making method, intelligent agent decision-making device, computer-readable storage medium, and electronic device in the exemplary embodiments of this disclosure have at least the following advantages and positive effects: In some embodiments of this disclosure, the technical solutions involve receiving user queries and obtaining a set of retrieved documents related to the user query. The user query and the retrieved document set are then jointly embedded to generate a context vector representation. This context vector representation is input into a hierarchical recursive memory network. The network performs hierarchical recursive state evolution across multiple computational segments to generate a final hidden state. This final hidden state is then decoded to generate the decision output of the intelligent agent. On one hand, this achieves effective synergy between external knowledge retrieval and deep internal reasoning, improving the decision accuracy of the intelligent agent in complex tasks (such as multi-hop reasoning and symbolic solving). Furthermore, through the hierarchical recursive structure and adaptive termination mechanism, the model can dynamically adjust the reasoning depth, optimizing computational resource consumption while ensuring output quality, thereby enhancing the system's efficiency and interpretability.

[0019] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0021] Figure 1 A flowchart illustrating the intelligent agent decision-making method in an embodiment of this disclosure is shown. Figure 2 This diagram illustrates how a hierarchical recursive memory network is used in an embodiment of this disclosure to perform hierarchical recursive state evolution across multiple computational segments to generate the final internal hidden state. Figure 3 This is a flowchart illustrating how, in an embodiment of the present disclosure, a high-level hidden state is recursively passed as a memory carrier across computational segments, and when a preset inference termination condition is met, it is determined as the final internal hidden state. Figure 4This diagram illustrates how the final hidden internal state is decoded to generate the decision output of the intelligent agent in an embodiment of this disclosure. Figure 5 A schematic diagram of the HRM processing procedure in an embodiment of this disclosure is shown; Figure 6 This diagram illustrates the overall flow of the intelligent agent decision-making method in an embodiment of this disclosure. Figure 7 This diagram illustrates the structure of the intelligent agent decision-making device in an exemplary embodiment of this disclosure. Figure 8 A schematic diagram of the structure of an electronic device in an exemplary embodiment of this disclosure is shown. Detailed Implementation

[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0023] The terms “a,” “an,” “the,” and “the” are used in this specification to indicate the presence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended inclusion and to mean that there may be other elements / components / etc. in addition to the listed elements / components / etc.; the terms “first” and “second” are used only as markings and are not a limitation on the number of objects.

[0024] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0025] First, the technical terms used in this disclosure will be explained: Retrieval-Augmented Generation (RAG): An artificial intelligence framework that enhances the accuracy and timeliness of content generated by large language models by retrieving relevant information from external knowledge bases.

[0026] Agentic Retrieval Enhanced Generation (AAG): This introduces one or more "agents" into the RAG framework. These agents can simulate human thinking, planning, actions (such as retrieving, using tools, and generating content), and reflection to handle more complex tasks.

[0027] Hierarchical Recurrent Memory (HRM) is a neural network architecture inspired by the brain's hierarchical and multi-timescale processing mechanisms. It enables the model to perform deep, coherent reasoning in its internal hidden state space without relying on explicit chaining of thought by constructing a recurrent network that couples a high-level slow module (responsible for abstract planning) and a low-level fast module (responsible for detailed computation).

[0028] Latent Reasoning: refers to the process by which a model performs complex computations and reasoning directly in its internal hidden state space without generating intermediate language descriptions.

[0029] Adaptive Halting Strategy: A mechanism that allows the model to dynamically determine the depth of the inference process (i.e., the number of recursive steps) based on the complexity of the task, achieving a balance between "fast thinking" and "slow thinking" to optimize the use of computational resources.

[0030] The Abstraction and Reasoning Corpus (ARC-AGI) is a benchmark dataset used to measure the abstraction and reasoning abilities of artificial intelligence systems and is considered a key challenge in measuring the potential of artificial general intelligence.

[0031] Retrieval augmented generation (RAG) technology has become a key technique for improving the performance of large language models (LLMs) when handling knowledge-intensive tasks. By introducing the concept of "agent," agentic retrieval augmented generation (AAG) further expands the capabilities of RAG systems, enabling them to perform more complex planning, multi-round iterative retrieval, and interaction with external tools.

[0032] While Agentic Retrieval Augmented Generative Architectures (AAGs) can retrieve massive amounts of information, their ability to process and integrate this information depends entirely on the reasoning quality of their core language model. Currently, this reasoning is primarily achieved through Chained Thinking (CoT), but the following core bottlenecks severely limit the performance ceiling of Agentic RAG systems: First, insufficient inference depth and the fragility of planning: When faced with complex queries, the Agentic RAG system needs to perform in-depth synthesis, verification, and inference on multiple, even potentially contradictory, information sources. However, the fixed computational depth of the standard Transformer architecture fundamentally limits its ability to perform such complex, multi-step inference. The CoT inference chain on which the system relies is extremely fragile; when processing multi-step logic extracted from documents or resolving conflicts, even a small error in any link can lead to the failure of the entire task planning (such as multi-round retrieval and tool calls).

[0033] Second, the efficiency bottleneck of retrieval and reasoning: The strength of Agentic RAG lies in its iterative "retrieval-thinking-action" cycle. However, when each "thinking" step relies on generating a large number of intermediate tokens (CoT), the response speed of the entire system becomes extremely slow, and the computational cost increases dramatically. This significantly reduces the practicality of RAG systems for complex problems that require multiple iterations to solve, making them unsuitable for real-time or high-concurrency application scenarios.

[0034] Third, there is the challenge of generalizing reasoning capabilities to specific domains: the effectiveness of CoT reasoning highly depends on specific patterns learned from massive amounts of data. This makes it difficult for Agentic RAG systems to be efficiently adapted to professional scenarios requiring deep domain knowledge reasoning (such as law, medicine, and scientific research). This is because labeled data (SFT data) used to train models for deep and accurate reasoning of retrieved content is often scarce in these domains, limiting the depth and reliability of RAG systems in key vertical fields.

[0035] Therefore, there is an urgent need to develop a novel Agentic RAG framework that integrates a more efficient and deeper "Latent Reasoning" mechanism. This mechanism should be able to simulate the efficient "silent thinking" process in the brain, performing deep and coherent computations within the model's hidden state space without relying on generating large amounts of intermediate language descriptions. Only by deeply integrating this advanced latent reasoning capability with RAG's knowledge retrieval capabilities can the performance bottlenecks of existing methods be overcome, enabling Agentic RAG systems to potentially solve highly complex and abstract problems.

[0036] To address the core challenges of existing Agentic Retrieval Augmented Generation (AAG) techniques in handling complex symbolic reasoning and pattern recognition—namely, low reasoning efficiency, insufficient depth, and excessive computational resource consumption—this disclosure proposes a novel Agentic Retrieval Augmented Generation method based on Hierarchical Recursive Memory Network (HRM). The core objective is to deeply integrate the efficient internal state reasoning architecture of a hierarchical recursive memory network into a dedicated reasoning engine within the Agentic RAG system, thereby breaking through the current mainstream methods' reliance on explicit chained thinking.

[0037] In the embodiments of this disclosure, an intelligent agent decision-making method is first provided, which at least to some extent overcomes the shortcomings of insufficient reasoning depth and shallow semantic understanding in related technologies.

[0038] Figure 1 The diagram shows a flowchart of the intelligent agent decision-making method in an embodiment of this disclosure. The executing entity of the intelligent agent decision-making method can be an intelligent agent.

[0039] refer to Figure 1 The intelligent agent decision-making method according to an embodiment of the present disclosure includes the following steps: Step S110: Receive user query and obtain a set of search documents related to user query; Step S120: Perform joint embedding on the user query and the retrieved document set to generate a context vector representation; Step S130: Input the context vector representation into the hierarchical recursive memory network, and perform hierarchical recursive state evolution in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; Step S140: Decode the final internal hidden state to generate the decision output of the intelligent agent.

[0040] exist Figure 1 In the technical solution provided by the illustrated embodiment, by receiving user queries and obtaining a set of search documents related to the user queries, the user queries and the set of search documents are jointly embedded to generate a context vector representation. The context vector representation is then input into a hierarchical recursive memory network. Through the hierarchical recursive memory network, hierarchical recursive state evolution is performed in multiple computational segments to generate a final internal hidden state. The final internal hidden state is then decoded to generate the decision output of the intelligent agent. On the one hand, this achieves effective synergy between external knowledge retrieval and deep internal reasoning, improving the decision accuracy of the intelligent agent in complex tasks (such as multi-hop reasoning and symbolic solving). Furthermore, through the hierarchical recursive structure and adaptive termination mechanism, the model can dynamically adjust the reasoning depth, optimizing computational resource consumption while ensuring output quality, thereby enhancing the efficiency and interpretability of the system.

[0041] The following are Figure 1 The specific implementation process of each step in the process will be explained in detail: Before step S110, it should be noted that this disclosure can pre-train a hierarchical recursive memory model. The hierarchical recursive memory model in this disclosure includes not only low-level modules (L-modules) for state evolution and information integration, but also... ) and high-level modules (H-module, It also includes a decision evaluation head network and an output head network that work together with it, and the three together constitute a complete decision generation and evaluation architecture. Among them, the decision evaluation head network is used to perform meta-level evaluation of the current reasoning state and dynamically determine whether to continue deep reasoning or terminate the current thinking process. The output head network is responsible for decoding the final internal hidden state of HRM into the decision output (such as text response, action instruction or classification result) required for specific tasks.

[0042] For example, this disclosure can employ an efficient training mechanism that combines single-step gradient approximation with a deep supervision strategy to reduce training overhead and improve stability. Specifically, to avoid the high memory consumption of traditional backpropagation over time (BPTT), this disclosure only calculates the gradient in the last step of each inference segment, treating other intermediate states as constants, thereby reducing the memory overhead of backpropagation from O(T) to O(1). Simultaneously, the training process is divided into multiple computational segments. At the end of each segment, the loss is calculated, the model parameters are updated, and the current hidden state is separated from the computation graph and passed to the next segment, achieving denser supervision and stable gradient updates. This training method supports end-to-end joint optimization of the HRM and its decision evaluation head and output head, laying the foundation for subsequent inference.

[0043] Furthermore, this disclosure introduces a reinforcement learning-based dynamic inference control mechanism after each inference segment, which evaluates the current inference state through a dedicated Q-value head network. This Q-value head network, as a specific implementation of the decision evaluation head network, receives the final hidden state from the higher-level HRM modules. As input, the output is the predicted Q-value for the two actions: "continue" and "halt". and It is used to characterize the long-term expected benefit (also known as value score) of performing the corresponding action in the current reasoning state.

[0044] This Q-value head network is trained using a Q-learning framework, aiming to enable the model to proactively terminate inference when semantics are sufficient or computational gains diminish, thereby achieving adaptive allocation of computational resources. At the end of each segment, a supervision signal can be constructed based on current task performance and future expectations: if the current prediction is close to the true label or there is limited room for further improvement, then "terminating" is the superior action; otherwise, "continuing" is better. This signal is used to calculate the target Q-value. and the predicted Q value We construct a binary cross-entropy loss or mean squared error loss as the optimization objective of Q-learning.

[0045] For example, this disclosure may define an overall loss function. This is a weighted sum of the task performance loss and the Q-learning loss. Specifically, please refer to the following formula 1:

[0046] in, y represents the model's predicted output for the m-th computation segment; y is the corresponding true label. The main task loss measures the difference between the model output and the actual result; The value of the "continue / termination" action predicted by the Q-value head network; The target Q-value is calculated by combining environmental feedback (such as the potential for improving task accuracy) and the estimated Q-value of the next state, forming the temporal difference (TD) objective of reinforcement learning.

[0047] By minimizing This disclosure enables joint optimization of the HRM body, output head network and Q-value head network, so that the model can not only improve task performance, but also learn the metacognitive strategy of "when to stop thinking", which significantly enhances its autonomous decision-making ability in complex and abstract tasks.

[0048] Based on the above mechanism, this disclosure can achieve dynamic decision control during the inference phase. Specifically, after each round of inference, the model compares the "termination" and "continue" Q-values ​​output by the Q-head network. If they match... Greater than ,or If the preset threshold is exceeded, the preset reasoning termination condition is met, and the multi-step reasoning process immediately ends and enters the output stage; otherwise, the next round of state evolution continues. This mechanism endows the model with human-like "thinking rhythm" control capabilities, ensuring reasoning depth while avoiding redundant calculations and improving overall efficiency.

[0049] The following is the Chinese pseudocode implementation of the training process combining deep supervision and a dual-head structure: # Deep supervised training main loop for each (x, y_true, eval_label) in training_data_loader: # Initialize the hidden state of the samples in this round. z = z_init # For each supervised segment: for each segment in N_supervision_steps: # Execute a reasoning fragment to obtain the current state z and the two-headed output. z, y_hat, s = hrm_inference_with_heads(z, x) # Calculate the main task loss (output header) loss_output = cross_entropy_loss(y_hat, y_true) # Calculate and evaluate task losses (Decision Evaluation Header) loss_eval = bce_loss(s, eval_label) # Joint Losses total_loss = alpha loss_output + beta loss_eval # Backpropagation and parameter update total_loss.backward() optimizer.step() optimizer.zero_grad() # Separate the hidden state and pass in the next fragment. z = z.detach() # HRM inference fragment with dual-head output function hrm_inference_with_heads(z, x, N, T): x_vector = input_encoding_network(x) zH, zL = z # Unpacking status # Top N T-1 step: Gradient-free evolution with torch.no_grad(): for i in range(N T - 1): zL = L_module(zL, zH, x_vector) if (i + 1) % T == 0: zH = H_module(zH, zL) # Final step: Enable gradients zL = L_module(zL, zH, x_vector) zH = H_module(zH, zL) # Dual-head output y_hat = output_head(zH) # Main decision output s = decision_eval_head(zH, zL) # Decision evaluation score return (zH, zL), y_hat, s Based on the aforementioned pseudocode, this disclosure employs an efficient and scalable training mechanism that decomposes the long sequence inference process into multiple supervised segments. Within each segment, single-step gradient approximation is achieved by "freezing intermediate states and retaining only the final gradient," significantly reducing memory usage. Simultaneously, at the end of each segment, the output head network calculates the main task loss, and the decision evaluation head network calculates the evaluation task loss. Backpropagation and parameter updates are then performed using a joint loss function, achieving end-to-end collaborative optimization of the HRM main body and the dual-head structure. After the update, the current hidden state is separated from the computation graph and passed to the next segment to block gradient backpropagation and prevent training instability caused by long-term dependencies. This method significantly improves training efficiency and convergence stability while maintaining the model's deep inference capabilities, making it suitable for multi-level recursive modeling of complex abstract tasks.

[0050] After obtaining the trained hierarchical recursive memory model, the efficient internal reasoning architecture of the above model can be integrated into the proxy-based retrieval enhancement generation system as a dedicated neural reasoning engine to realize the dynamic, deep, and interpretable intelligent decision-making method proposed in this disclosure.

[0051] Therefore, we can proceed to step S110, where we receive the user query and obtain a set of search documents related to the user query.

[0052] In this step, user queries can be received. And obtain the collection of search documents related to the user's query. .

[0053] Among them, the set of retrieved documents related to the user query refers to a group of candidate documents selected by the information retrieval system from a preset knowledge base or document library based on semantic similarity or keyword matching degree, whose content has potential relevance to the user's query request in terms of topic, intent or fact.

[0054] Specifically, the process of obtaining this collection of documents typically includes the following steps: Query encoding: Converting user queries into vector representations, for example, through dense retrieval models (such as DPR, Contriever) or sparse retrieval models (such as BM25); Similarity matching: In large-scale document databases (such as Wikipedia, enterprise knowledge bases, web indexes, etc.), calculate the similarity between the query vector and each document vector (such as cosine similarity or word overlap score). Document filtering: Select several documents with the highest similarity ranking (e.g., Top-K) to form an initial candidate document set; Optional re-ranking: To further improve relevance, candidate documents can be re-ranked, for example, by using a more refined cross-encoder model to evaluate the matching degree between the query and the document; Final output: The set of documents obtained after filtering and sorting is the "set of retrieved documents related to the user query", which can be used to support multi-hop reasoning, fact verification or generate enhanced answers.

[0055] The core function of this document collection is to provide external knowledge support for intelligent agents, compensating for the limitations of parameterized knowledge in the model, which is especially crucial when dealing with factual, time-sensitive, or specialized domain issues. In the hierarchical recursive memory model framework disclosed herein, this collection will be used as part of the input, progressively read, analyzed, and integrated across multiple reasoning segments to achieve a deep and interpretable reasoning process.

[0056] In step S120, the user query and the retrieved document set are jointly embedded to generate a context vector representation.

[0057] In this step, user queries can be... and retrieve document collections As input, it is processed through a science department's input encoding network. Transform into a unified context vector representation This input encoding network is used to fuse user query intent information with external knowledge from the retrieved documents, generating a semantically aligned and structurally consistent vector sequence or context embedding, which serves as the external input signal for subsequent hierarchical recursive memory (HRM) multi-step reasoning.

[0058] Specifically, the above-mentioned collection of retrieved documents It contains one or more candidate documents related to the user query, whose content is preprocessed (e.g., sentence segmentation, chunking, or entity annotation) and then concatenated or jointly encoded with the user query. Input encoding network. It can be implemented based on the Transformer architecture, a dual-tower model, or a multimodal fusion structure, and its parameters... Optimization is performed during training using end-to-end backpropagation. During encoding, the model not only captures keyword matching relationships between queries and documents but also models deep semantic connections, such as referential resolution, logical implication, or causal reasoning clues, thereby generating vector representations rich in contextual information. .

[0059] Furthermore, this disclosure relates to the initial hidden state of the HRM model. Perform initialization. This is the initial state of the lower-level modules. This is the initial state of the high-level modules. This initial state is sampled from a fixed, learnable truncated normal distribution and remains unchanged throughout the training and inference process, independent of input samples. This design ensures the model has a stable "initial mental state," avoiding inference fluctuations caused by random initialization, while guaranteeing comparability and consistency in state evolution across different inference segments.

[0060] In summary, the output of step S120 can include two parts: Unified context vector representation : Serves as the external driving signal for each step of the HRM network calculation; Fixed initial hidden state : Serves as the initial internal state of both low-level and high-level HRM modules, and participates in subsequent recursive updates.

[0061] This joint embedding mechanism ensures the effective integration of user intent and external knowledge, providing a high-quality input foundation for HRM models to conduct deep, multi-hop inference in complex abstract tasks.

[0062] In step S130, the context vector representation is input into the hierarchical recursive memory network, and the hierarchical recursive state evolution is performed in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state.

[0063] In this step, the aforementioned context vector representation can be input into a hierarchical recursive memory network. The hierarchical recursive memory network performs hierarchical recursive state evolution within multiple computational segments to generate the final internal hidden state.

[0064] The hierarchical recursive memory network disclosed herein includes coupled low-level modules (L-modules) and high-level modules (H-modules), wherein the low-level modules update their states at a first frequency, and the high-level modules update their states at a second frequency, wherein the second frequency is lower than the first frequency. Specifically, the low-level modules... Within each high-level cycle, fine-grained state updates are performed T times at a high frequency to capture local features of the input signal, process detailed information, and perform fast computations; while the high-level module... An update is performed only once every T low-level time steps to integrate the accumulated local results from the low level, and to perform abstract semantic modeling, long-term strategy planning, and macro-state adjustment.

[0065] refer to Figure 2 , Figure 2 This embodiment of the present disclosure illustrates how a hierarchical recursive memory network performs hierarchical recursive state evolution across multiple computational segments to generate the final hidden state, including steps S201-S203: In step S201, in each computation segment, the context vector representation is updated in a fine-grained state at a first frequency by the low-level module.

[0066] In this step, the aforementioned context vector As an external input, it can drive lower-level modules. Guided by the initial hidden state, state evolution is performed continuously at a relatively high time frequency. Specifically, within each higher-level update cycle, the lower-level module sequentially performs T rapid state updates with no gradient or the first T frozen gradients, to capture local patterns in the input signal, process detailed information, and complete preliminary semantic extraction. Each update depends on the lower-level hidden state of the previous time step. Current fixed high-rise hidden state and shared context vectors Through recursive functions To achieve dynamic modeling of the input content, where, These are the learnable parameters for lower-level modules, and this process reflects the model's perception and computational capabilities at fine-grained time scales.

[0067] In step S202, when the low-level update accumulates to a preset number of times, the high-level module receives the current state of the low-level module at a second frequency and updates the hidden state of the high-level module.

[0068] In this step, after the lower-level module completes the Tth update, it can be determined that the conditions for updating the higher-level state are met, and thus the latest hidden state of the lower-level module is revealed. As input, it is passed to the higher-level module. High-level modules perform an abstraction and integration operation at a low frequency, based on their previous state. With the received underlying state Through recursive functions Update its own high-level hidden state, among which... These are learnable parameters for higher-level modules. This update process summarizes and abstracts the results of multi-step computations at lower levels, forming higher-level semantic representations or policy judgments to guide lower-level reasoning in subsequent computational segments.

[0069] By alternating between steps S201 and S202, this disclosure constructs a hierarchical recursive computation mechanism across time scales, realizing the co-evolution of fast low-level perception and slow high-level abstraction. Specifically, within each high-level update cycle, the low-level module, guided by the current high-level hidden state, continuously performs T fine-grained state updates at a high frequency, focusing on processing the local dynamics of the input signal, capturing semantic details, and completing sub-task-level computations, gradually approaching local convergence. When the T-th update is completed, the high-level module is activated, receives the final state of the low-level module, integrates and abstracts its accumulated information, performs a macro-state transition (cognitive leap), updates its global representation, and uses this new state as the initial guiding condition for the low-level computation in the next cycle.

[0070] This mechanism achieves the core characteristic of "Hierarchical Convergence": lower-level modules iterate rapidly under the constraint of fixed higher-level states, forming locally stable intermediate representations; higher-level modules periodically absorb these results, completing a cognitive upgrade from "execution" to "reflection," and thus adjusting the direction of reasoning. This process effectively simulates the progressive cognitive cycle of "focused execution - summarization and induction - strategy adjustment" in human thinking, endowing the model with human-level progressive reasoning ability.

[0071] In step S203, the high-level hidden state is recursively passed as a memory carrier across computational segments, and when the preset inference termination condition is met, it is determined as the final internal hidden state.

[0072] In this step, the high-level hidden state can be recursively passed as a memory carrier across computational segments, and when a preset inference termination condition is met, it is determined as the final internal hidden state. For details, please refer to... Figure 3 , Figure 3 This illustration shows a flowchart of how, in an embodiment of this disclosure, a high-level hidden state is recursively passed as a memory carrier across computational segments, and when a preset inference termination condition is met, it is determined as the final internal hidden state, including steps S301-S304: In step S301, at the end of each calculation segment, the current hidden state of the high-level module is obtained.

[0073] In this step, after T low-level updates and one high-level update are completed within a computation segment (e.g., the m-th computation segment), the current high-level hidden state output by the high-level module can be extracted as the global semantic representation at the end of the segment. This state integrates all low-level fine-grained processing results and high-level abstract integration information in the current segment, reflecting the model's understanding of the input context and the current inference process, and is the core basis for subsequent decision-making.

[0074] In step S302, an action value score is generated based on the current high-level hidden state, and it is determined whether the preset reasoning termination condition is met based on the action value score.

[0075] In this step, the current high-level hidden state can be input into the decision evaluation head network (or Q-value head network), which outputs the predictive value scores for the two actions: "continue" and "halt," denoted as... and Then, it can be determined whether to end the multi-step reasoning process based on the preset reasoning termination condition. The preset reasoning termination condition can include any of the following conditions being met: The value score for the terminating action is higher than that for the continuing action, i.e. Greater than ; The value score of the termination action exceeds the preset threshold. ,Right now ; The current inference segment number has reached the preset maximum number of segments. ; This mechanism enables the model to have dynamic reasoning depth control capabilities, and can autonomously decide "when to stop thinking" based on semantic integrity, avoiding redundant calculations.

[0076] In step S303, if the preset reasoning termination condition is met, the current high-level hidden state is determined as the final internal hidden state.

[0077] In this step, once any of the above termination conditions is met, the execution of subsequent inference segments can be stopped immediately. The current high-level hidden state output by the current segment is then used as the final internal hidden state of the HRM network and passed to the output head network to generate the final decision output. This state represents the most complete and stable semantic understanding formed by the model after multiple hierarchical recursive evolutions, and is the cornerstone for generating high-quality responses.

[0078] In step S304, if the preset reasoning termination condition is not met, the current high-level hidden state is used as the initial hidden state of the next computation segment to start a new round of hierarchical recursive state evolution.

[0079] In this step, if the aforementioned preset reasoning termination condition is not met, the current high-level hidden state can be retained and passed to the next computation segment as its initial high-level state; meanwhile, the low-level state can be reset or continued (depending on the specific implementation). This recursive passing mechanism realizes cross-segment knowledge accumulation and state continuation, enabling the model to gradually deepen its understanding at multiple stages and complete complex multi-hop reasoning tasks.

[0080] This disclosure constructs a dynamic reasoning architecture with metacognitive capabilities by using high-level hidden states as cross-segment memory carriers and combining them with a value-based adaptive termination mechanism. This design breaks through the limitations of traditional fixed-step reasoning, achieves intelligent allocation of computing resources, and significantly improves efficiency and flexibility while ensuring reasoning depth.

[0081] Next, refer to Figure 1 In step S140, the final internal hidden state is decoded to generate the decision output of the intelligent agent.

[0082] In this step, after obtaining the final hidden state, the final hidden state can be decoded to generate the decision output of the intelligent agent.

[0083] For details, please refer to Figure 4 , Figure 4 This illustration shows a flowchart of how the final hidden internal state is decoded to generate the decision output of the intelligent agent in an embodiment of this disclosure, including step S401: In step S401, the final hidden internal state is input to one or more output head networks for decoding, and a decision result adapted to the current task requirements is generated through the output head networks.

[0084] In this step, the final hidden internal state can be fed into one or more dedicated output header networks. The system decodes the abstract semantics inherent in the state to generate structured or unstructured decision results. This process transforms the model's "internal thinking" into "external behavior," and is a key step for the intelligent agent to complete the task loop.

[0085] The above decision results are used to drive the intelligent agent to perform at least one of the following operations: generate a natural language response, invoke a specified external tool, initiate a new round of information retrieval, or generate a structured analysis report, the specific format of which is dynamically adapted according to the current task type and context requirements: First, direct answer generation: For question-answering tasks, the output head network acts as a sequence generator (such as a Transformer-based decoder), decoding the final hidden state into a fluent and accurate natural language response. This model is suitable for scenarios where the reasoning process has converged and the conclusion is clear, such as "Einstein won the Nobel Prize in Physics in 1921"; Second, generating structured action instructions: In complex agentic processes, not only is the final answer included, but the decision-making intent for subsequent actions is also implicit. In this case, the output header network, acting as a policy network, generates standardized structured instructions to trigger external execution modules, including but not limited to: SEARCH["more precise keywords"]: Indicates that an optimized query should be constructed based on the current inference bottleneck and a new round of retrieval should be initiated; TOOL[Calculator,"3 5": Use the calculator tool to perform mathematical calculations; TASK_DECOMPOSE["Subtask Goals"]: Decomposes complex problems into a sequence of executable subtasks, supporting multi-stage planning.

[0086] Third, generate a structured report with confidence level and conflict analysis: When there are contradictory information, insufficient evidence, or high uncertainty in the reasoning results in the retrieved documents, the output head network generates a structured analysis report to enhance the transparency and interpretability of decision-making. This report may include at least one of the following: Key conclusion: The representation is the optimal decision result or the most likely answer derived from the final hidden state when the reasoning process terminates; Confidence assessment: used to quantify the credibility of reasoning about core conclusions, and can be expressed as a probability value, score range or grade label (such as high / medium / low). Its calculation is based on model uncertainty, strength of evidence support or multi-expert voting mechanism. Supporting evidence and contradictory evidence: These are key information fragments from the retrieved documents obtained in this reasoning that support or refute the core conclusion, and may include the source and citation. Unresolved conflict points: These indicate contradictory statements in the current collection of retrieved documents that cannot be resolved through semantic fusion or logical reasoning, suggesting the need for further verification or human intervention.

[0087] Through the aforementioned multimodal output mechanism, this disclosure achieves a flexible mapping from a unified internal representation to diverse external behaviors, enabling the HRM model to be seamlessly integrated into the core components of the Agentic RAG framework, such as planning, execution, reflection, and iteration. This design not only enhances the task adaptability of the intelligent agent but also strengthens its robustness and reliability in open domains, high-risk scenarios, or situations with uncertain information.

[0088] The solution disclosed herein will be explained in the following specific application scenarios: Based on the technical solution disclosed herein, an intelligent problem-solving agent system can be constructed for efficiently solving complex symbolic reasoning tasks, such as "solving a difficult Sudoku puzzle." When a user submits the problem, the system first retrieves rules related to Sudoku solving (such as "the numbers 1-9 in each row, column, and cell must not be repeated") and typical solution examples from an external knowledge base through the Retrieval Enhanced Generation (RAG) module, thereby obtaining structured or semi-structured prior knowledge.

[0089] Subsequently, an information encoding and structuring module jointly parses the retrieval results and the initial Sudoku board input by the user, transforming the unstructured text rules and graphical board into a unified, structured task representation that can be processed by the HRM model. This representation typically encodes the initial numerical distribution, constraints, and inference objectives in the form of vector sequences or tensors, serving as external input to the HRM inference engine.

[0090] HRM module receives Then, the hierarchical recursive reasoning process is initiated: the lower-level modules, guided by the higher-level states, attempt to locally fill candidate numbers and propagate constraints; the higher-level modules periodically integrate the lower-level results, evaluate global consistency, and plan the next reasoning strategy. The entire process proceeds step by step in multiple computational segments. After each segment, the decision evaluation head network determines whether the current state has converged to a complete and valid solution. Once the preset termination condition is met (such as all cells being filled and no conflicts), the HRM stops reasoning and outputs the structured solution decomposition result decoded from the final internal hidden state. It is a complete 9×9 digital matrix.

[0091] The structured solution can be presented to the user directly in tabular form, or it can be further input into a lightweight generation module to be transformed into solution steps described in natural language, such as "Only 7 can be filled in the 5th row and 3rd column because other numbers have already appeared in the same row, column or palace".

[0092] Compared to traditional Chain-of-Thought (CoT) methods that rely on autoregressive text generation to represent intermediate reasoning processes, this approach utilizes Hidden Regression Metrics (HRM) to perform compact and efficient symbolic operations in the hidden state space. This avoids lengthy and error-prone natural language intermediate representations, significantly improving the efficiency and accuracy of solving complex, structured, and rule-based reasoning tasks such as Sudoku. Furthermore, the adaptive reasoning termination mechanism ensures that the model performs depth searches only when necessary, achieving dynamic optimization of computational resources.

[0093] This application scenario fully demonstrates the comprehensive advantages of the disclosed technical solution in combining external knowledge, performing multi-step logical reasoning, and outputting structured decisions, showcasing its broad application potential in fields such as educational assistance, logic game solving, and formal verification.

[0094] refer to Figure 5 , Figure 5 This diagram illustrates the HRM processing procedure in an embodiment of the present disclosure, including steps S501-S507: In step S501, a query from the user is received. 1. Search document collection and historical hidden state (This is the contextual memory accumulated by the intelligent agent during multiple rounds of historical interactions. It is usually encoded by the final high-level hidden state after each round of reasoning and stored in the external memory module. It is then retrieved and reused in subsequent tasks to support coherent reasoning and decision-making across rounds.) In step S502, the input information is jointly encoded, and the internal states of the low-level module (L-module) and the high-level module (H-module) are initialized to prepare for subsequent reasoning; In step S503, the hierarchical recursive reasoning process is initiated. This process dynamically schedules two modules based on task complexity: on the one hand, the low-level module L-module performs rapid state updates to achieve efficient perception and local processing of input information; on the other hand, the high-level module H-module performs slow abstract thinking to complete deep modeling of semantic structure and logical relationships. In step S504, the lower-level module L-module and the higher-level module H-module run in parallel, outputting their current state evolution results respectively, and then merging the information of the two and passing it to the next judgment step; In step S505, based on the fused state information, an adaptive termination judgment is performed to assess whether the current inference state has reached convergence or a decision threshold. The judgment is based on a Q-value comparison: if the "continue" value is greater than the "terminate" value, the next inference cycle is entered (i.e., the process returns to step S503); otherwise, the inference is terminated. In step S506, when the adaptive termination judgment result is termination, the output generation and decision-making stage is entered, and the final high-level hidden state is used for decoding to generate the response action or decision content of the intelligent agent. In step S507, the final decision result is output, completing a full HRM inference process.

[0095] refer to Figure 6 , Figure 6This diagram illustrates the overall flow of the intelligent agent decision-making method in this embodiment, including steps S601-S606: In step S601, the intelligent agent receives user queries. ; In step S602, the intelligent agent will process the user query. The query is submitted to the RAG retrieval module, which combines semantic matching and information retrieval with an external knowledge base to obtain a collection of documents related to the current query. ; In step S603, the user query is... Document collection Interaction with historical states Input is fed into the HRM inference module to initiate a hierarchical recursive inference process, generating the hidden state and preliminary decision of the current task; In step S604, the HRM reasoning module performs multiple rounds of iterative reasoning based on the current context and outputs two possible results: one is to directly generate the final answer or action instruction; the other is to generate an "iteration and reflection" signal, indicating that the reasoning path needs to be further optimized. In step S605, if the HRM determines that further optimization is needed, the system will analyze the current inference state, generate new retrieval requirements or adjust the strategy, and update the historical interaction state. This triggers the RAG module to perform a new round of retrieval, and then the new results are re-inputted into the HRM to continue inference until the termination condition is met. In step S606, the final answer, decision, or action instruction is output and submitted to the user, completing a full intelligent agent decision-making process.

[0096] Based on the above technical solutions, this disclosure has at least the following technical effects: First, it achieves efficient deep solving of complex symbolic reasoning tasks. By introducing a hierarchical recursive memory (HRM) module as a dedicated reasoning engine, this disclosure enables intelligent agents to perform non-linguistic multi-step logical reasoning and symbolic operations in the hidden state space. This significantly reduces the computational overhead and error accumulation risk caused by the autoregressive generation of a large number of intermediate tokens in traditional chain-of-thought (CoT) methods when dealing with tasks such as ARC-AGI and complex Sudoku, thus achieving a better balance between reasoning efficiency and accuracy.

[0097] Second, a modular intelligent architecture that separates and coordinates retrieval and reasoning is constructed. This disclosure clearly defines the functional boundaries between the RAG module and the HRM module: RAG is responsible for retrieving relevant information from external knowledge bases and completing preliminary semantic integration, while HRM focuses on deep, recursive reasoning computation on structured task representations. By designing a standardized "information conversion" interface, unstructured text is automatically parsed into structured inputs (such as constraint matrices and rule encodings) that HRM can recognize, achieving decoupling and efficient collaboration between knowledge acquisition and logical reasoning, and improving the modularity and maintainability of the system.

[0098] Third, a dynamic computational resource allocation mechanism based on cognitive state is implemented. Leveraging the adaptive termination strategy integrated within the HRM (such as the ACT mechanism based on action value scoring), this disclosure enables the intelligent agent to autonomously determine whether to continue "thinking" or terminate output based on the convergence degree of the current inference state. This mechanism endows the system with human-like "metacognitive" capabilities: it can respond quickly to simple problems and automatically extend the inference depth for complex tasks, thereby effectively avoiding redundant computation and improving overall operational efficiency and resource utilization while ensuring decision-making quality.

[0099] In summary, by using HRM as a pluggable "intelligent reasoning unit" in the Agentic RAG system, this disclosure not only enhances the system's ability to handle highly complex abstract tasks, but also promotes the paradigm upgrade of intelligent agents from "retrieval-assembly" to "understanding-reasoning-decision-making," significantly expanding its application boundaries in complex scenarios such as educational assistance, logic solving, and program synthesis, and improving the system's intelligence level and practical value.

[0100] This disclosure also provides an intelligent agent decision-making device. Figure 7 This diagram illustrates the structure of the intelligent agent decision-making device in an exemplary embodiment of this disclosure; as shown... Figure 7 As shown, the intelligent agent decision-making device 700 may include a receiving module 710, an encoding module 720, a hierarchical recursive processing module 730, and a decision output module 740. Wherein: The receiving module 710 is used to receive user queries and obtain a set of search documents related to the user queries; Encoding module 720 is used to jointly embed the user query and retrieved document set to generate a context vector representation; The hierarchical recursive processing module 730 is used to input the context vector representation into the hierarchical recursive memory network, and perform hierarchical recursive state evolution in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; The decision output module 740 is used to decode the final internal hidden state and generate the decision output of the intelligent agent.

[0101] In an exemplary embodiment of this disclosure, the hierarchical recursive memory network includes coupled low-level modules and high-level modules; The lower-level module updates its state at a first frequency, and the higher-level module updates its state at a second frequency, wherein the second frequency is lower than the first frequency.

[0102] In an exemplary embodiment of this disclosure, the hierarchical recursive processing module 730 performs hierarchical recursive state evolution within multiple computational segments through the hierarchical recursive memory network to generate a final internal hidden state, including: In each computation segment, the context vector representation is updated in a fine-grained state at the first frequency by the lower-level module; When the low-level updates accumulate to a preset number of times, the high-level module receives the current state of the low-level module at the second frequency and updates the high-level hidden state. The high-level hidden state is recursively passed as a memory carrier across computational segments, and when a preset reasoning termination condition is met, it is determined as the final internal hidden state.

[0103] In an exemplary embodiment of this disclosure, the hierarchical recursive processing module 730 recursively passes the high-level hidden state as a memory carrier across computational segments, and determines it as the final inner hidden state when a preset inference termination condition is met, including: At the end of each computation segment, obtain the current high-level hidden state corresponding to the high-level module; An action value score is generated based on the current high-level hidden state, and the preset reasoning termination condition is determined based on the action value score. If the preset reasoning termination condition is met, then the current high-level hidden state is determined as the final internal hidden state; If the preset reasoning termination condition is not met, the current high-level hidden state will be used as the initial hidden state of the next computation segment to start a new round of hierarchical recursive state evolution.

[0104] In an exemplary embodiment of this disclosure, the hierarchical recursive processing module 730 generates an action value score based on the current high-level hidden state, including: The current high-level hidden state is input into the decision evaluation head network; The decision evaluation head network outputs predicted value scores for both the continue action and the terminate action; The preset reasoning termination conditions include: the predicted value score corresponding to the termination action is greater than the predicted value score corresponding to the continuation action, or the predicted value score corresponding to the termination action is greater than a preset threshold.

[0105] In an exemplary embodiment of this disclosure, the decision output module 740 decodes the final internal hidden state to generate a decision output for the intelligent agent, including: The final hidden internal state is input into one or more output head networks for decoding, and the output head networks generate decision results that are adapted to the current task requirements. The decision result is used to drive the intelligent agent to perform at least one of the following operations: generate a natural language response, invoke a specified external tool, initiate a new round of information retrieval, or generate a structured analysis report.

[0106] In an exemplary embodiment of this disclosure, the structured analysis report includes at least one of the following: The core conclusion represents the optimal decision result derived from the final hidden state when the reasoning process terminates; Confidence assessment is used to quantify the credibility of the reasoning regarding the core conclusions. Supporting evidence and contradictory evidence are respectively identified as information fragments from the retrieved documents obtained in this reasoning that support or refute the core conclusion. Unresolved conflict points are used to indicate semantic contradictions in the current set of retrieved documents that cannot be resolved through evidence fusion.

[0107] The specific details of each module in the aforementioned intelligent agent decision-making device have been described in detail in the corresponding intelligent agent decision-making method, so they will not be repeated here.

[0108] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0109] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0110] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0111] This disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device.

[0112] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0113] Computer-readable storage media can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0114] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0115] Furthermore, this disclosure also provides an electronic device capable of implementing the above-described method.

[0116] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0117] The following reference Figure 8 To describe an electronic device 800 according to such an embodiment of the present disclosure. Figure 8 The electronic device 800 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0118] like Figure 8 As shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one processor 810, at least one memory 820, a bus 830 connecting different system components (including memory 820 and processor 810), and a display 840.

[0119] The memory stores program code that can be executed by the processor 810, causing the processor 810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processor 810 can perform actions such as... Figure 1 As shown: Step S110, receive the user query and obtain the set of retrieved documents related to the user query; Step S120, jointly embed the user query and the set of retrieved documents to generate a context vector representation; Step S130, input the context vector representation into a hierarchical recursive memory network, and perform hierarchical recursive state evolution in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; Step S140, decode the final internal hidden state to generate the decision output of the intelligent agent.

[0120] The memory 820 may include a readable medium in the form of volatile storage, such as random access memory (RAM) 8201 and / or cache memory 8202, and may further include read-only memory (ROM) 8203.

[0121] The memory 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0122] Bus 830 can represent one or more of several types of bus structures, including a memory bus or memory controller, peripheral bus, graphics acceleration port, processor, or a local bus using any of the various bus structures.

[0123] Electronic device 800 can also communicate with one or more external devices 900 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 800, and / or with any device that enables electronic device 800 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 850. Furthermore, electronic device 800 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 860. As shown, network adapter 860 communicates with other modules of electronic device 800 via bus 830. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 800, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0124] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

Claims

1. A smart agent decision-making method, characterized in that, include: Receive user queries and obtain a set of search documents related to the user queries; The user query and retrieved document set are jointly embedded to generate a context vector representation; The context vector representation is input into a hierarchical recursive memory network, and the hierarchical recursive state evolution is performed in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; The final hidden internal state is decoded to generate the decision output of the intelligent agent.

2. The method according to claim 1, characterized in that, The hierarchical recursive memory network includes coupled low-level modules and high-level modules; The lower-level module updates its state at a first frequency, and the higher-level module updates its state at a second frequency, wherein the second frequency is lower than the first frequency.

3. The method according to claim 2, characterized in that, The process of generating the final hidden state through hierarchical recursive state evolution within multiple computational segments via the hierarchical recursive memory network includes: In each computation segment, the context vector representation is updated in a fine-grained state at the first frequency by the lower-level module; When the low-level updates accumulate to a preset number of times, the high-level module receives the current state of the low-level module at the second frequency and updates the high-level hidden state. The high-level hidden state is recursively passed as a memory carrier across computational segments, and when a preset reasoning termination condition is met, it is determined as the final internal hidden state.

4. The method according to claim 3, characterized in that, The step of recursively passing the high-level hidden state as a memory carrier across computational segments, and determining it as the final internal hidden state when a preset inference termination condition is met, includes: At the end of each computation segment, obtain the current high-level hidden state corresponding to the high-level module; An action value score is generated based on the current high-level hidden state, and the preset reasoning termination condition is determined based on the action value score. If the preset reasoning termination condition is met, then the current high-level hidden state is determined as the final internal hidden state; If the preset reasoning termination condition is not met, the current high-level hidden state will be used as the initial hidden state of the next computation segment to start a new round of hierarchical recursive state evolution.

5. The method according to claim 4, characterized in that, The generation of action value scores based on the current high-level hidden state includes: The current high-level hidden state is input into the decision evaluation head network; The decision evaluation head network outputs predicted value scores for both the continue action and the terminate action; The preset reasoning termination conditions include: the predicted value score corresponding to the termination action is greater than the predicted value score corresponding to the continuation action, or the predicted value score corresponding to the termination action is greater than a preset threshold.

6. The method according to any one of claims 1 to 5, characterized in that, Decoding the final hidden state to generate the intelligent agent's decision output includes: The final hidden internal state is input into one or more output head networks for decoding, and the output head networks generate decision results that are adapted to the current task requirements. The decision result is used to drive the intelligent agent to perform at least one of the following operations: generate a natural language response, invoke a specified external tool, initiate a new round of information retrieval, or generate a structured analysis report.

7. The method according to claim 6, characterized in that, The structured analysis report includes at least one of the following: The core conclusion represents the optimal decision result derived from the final hidden state when the reasoning process terminates; Confidence assessment is used to quantify the credibility of the reasoning regarding the core conclusions. Supporting evidence and contradictory evidence are respectively identified as information fragments from the retrieved documents obtained in this reasoning that support or refute the core conclusion. Unresolved conflict points are used to indicate semantic contradictions in the current set of retrieved documents that cannot be resolved through evidence fusion.

8. An intelligent agent decision-making device, characterized in that, include: The receiving module is used to receive user queries and obtain a set of search documents related to the user queries; The encoding module is used to jointly embed the user query and retrieved document set to generate a context vector representation; The hierarchical recursive processing module is used to input the context vector representation into the hierarchical recursive memory network, and perform hierarchical recursive state evolution in multiple computation segments through the hierarchical recursive memory network to generate the final internal hidden state; The decision output module is used to decode the final internal hidden state and generate the decision output of the intelligent agent.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the intelligent agent decision-making method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the intelligent agent decision-making method according to any one of claims 1 to 7 by executing the executable instructions.