Multi-level memory collaborative decision-making method, device, equipment and medium
By employing a multi-level memory-based collaborative decision-making method, environmental observation data is acquired to update the memory module, generating and analyzing communication content. This solves the problem of low collaboration efficiency in multi-agent systems in communication-constrained environments, achieving efficient task execution and decision consistency.
Patent Information
- Application Number
- CN202511090379.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multi-agent collaborative systems suffer from low collaboration efficiency, communication redundancy, and poor task continuity in decentralized scenarios where communication is limited and the environment is partially observable, due to the lack of dynamic context awareness and efficient communication decision-making mechanisms.
A multi-level memory collaborative decision-making method is adopted. By acquiring environmental observation data to update the multi-level memory module, information is retrieved to generate candidate communication content, and action plans are analyzed based on the thought chain reasoning of language models. The communication content is determined and included in the set of actions to be executed, and finally, basic operation instructions are generated and executed.
It enhances the collaborative optimization of task planning and communication behavior among agents in complex environments, thereby improving response efficiency, decision consistency, and task execution capabilities.
Smart Images

Figure CN120975119A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multi-level memory collaborative decision-making method, apparatus, device, and storage medium. Background Technology
[0002] In the development of embodied agents and multi-agent collaborative systems, although existing technologies have made some progress in single-agent perception and control, they still have significant limitations when faced with situations where multiple agents collaborate to complete complex tasks. Many traditional systems rely on centralized scheduling and control methods or assume that each agent can obtain complete and synchronous environmental observation information. These assumptions are difficult to hold in dynamic and non-ideal environments, especially when perception capabilities are limited, communication incurs physical costs, and control mechanisms cannot be unified. This can easily lead to a lack of contextual coordination and global consistency in individual behaviors, ultimately affecting the overall collaborative efficiency and task completion quality of the system.
[0003] Furthermore, existing multi-agent systems typically allow agents to initiate communication at any time without fully considering practical constraints such as limited bandwidth, time latency, and energy costs. This near-idealized communication mechanism ignores the resource consumption inherent in the communication itself, potentially leading to information overload among agents, increased response latency, and difficulties in conflict resolution in real-world deployments. Especially in tasks requiring semantic understanding of communication content and adjustments to execution plans accordingly, the lack of modeling and control over communication costs often results in redundant interactions and unnecessary information synchronization, thereby weakening the system's execution efficiency and robustness.
[0004] In the fintech business, embodied intelligent agents often need to perceive the physical environment and collaborate with other execution units when performing tasks such as counter interaction, on-site verification, and data verification. Existing systems lack sufficient integration and sharing of locally perceived data, resulting in ineffective collaboration between different agents and difficulty in flexibly handling complex scenarios such as high-concurrency customer processing and sudden changes in transaction processes, further limiting the depth and breadth of task automation.
[0005] In the healthcare field, multi-agent systems are applied to complex tasks such as intelligent ward rounds, surgical assistance, and material delivery. However, due to constantly changing environmental conditions and the need for real-time collaboration among multiple agents, traditional methods struggle to simultaneously ensure both the rationality of task planning and the continuity of cross-agent information coordination. Instability in communication, insufficient semantic understanding, or delayed information updates can easily lead to interruptions or misjudgments in critical tasks, impacting the safety and timeliness of medical procedures. Summary of the Invention
[0006] The main objective of this invention is to provide a multi-level memory collaborative decision-making method, apparatus, device, and storage medium, which aims to solve the technical problems of low collaboration efficiency, communication redundancy, and poor task continuity in existing multi-agent collaborative systems in decentralized scenarios where communication is limited and the environment is partially observable, due to the lack of dynamic context awareness and efficient communication decision-making mechanisms.
[0007] To achieve the above objectives, the present invention provides a multi-level memory collaborative decision-making method, comprising:
[0008] Acquire raw environmental observation data and update the multi-level memory module used to store environmental information, interaction history and operation instructions based on the raw observation data;
[0009] Retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information;
[0010] Based on the current task objective, executable high-level actions are extracted from the operation instructions of the updated multi-level memory module to form a list of candidate action plans;
[0011] The language model-based thought chain reasoning analysis is used to determine whether to send the candidate communication content, based on the list of alternative action plans and the candidate communication content.
[0012] When it is determined that the candidate communication content will be sent, the candidate communication content will be included in the set of actions to be executed.
[0013] A final action plan is generated based on the list of candidate action plans and the set of actions to be executed, and the final action plan is decomposed into at least one basic operation instruction, which is then executed to interact with the environment.
[0014] Furthermore, to achieve the above objectives, the present invention provides a multi-level memory collaborative decision-making device, comprising:
[0015] The environmental perception and memory update module is used to acquire raw environmental observation data and update the multi-level memory module for storing environmental information, interaction history and operation instructions based on the raw observation data.
[0016] The memory retrieval and communication construction module is used to retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information.
[0017] The target analysis and action extraction module is used to extract executable high-level actions from the operation instructions of the updated multi-level memory module based on the current task target, and form a list of candidate action plans.
[0018] The reasoning analysis and communication decision module is used to perform reasoning analysis based on the language model's thought chain to determine whether to send the candidate communication content;
[0019] The communication action generation module is used to include the candidate communication content to be sent into the set of actions to be executed when it is determined that the candidate communication content should be sent.
[0020] The action orchestration and interaction execution module is used to generate a final action plan based on the list of candidate action plans and the set of actions to be executed, and to decompose the final action plan into at least one basic operation instruction, and to execute the basic operation instruction to interact with the environment.
[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multi-level memory cooperative decision-making program stored in the memory and executable on the processor, wherein when the multi-level memory cooperative decision-making program is executed by the processor, it implements the steps of the multi-level memory cooperative decision-making method as described above.
[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multi-level memory cooperative decision-making program, wherein the multi-level memory cooperative decision-making program, when executed by a processor, implements the steps of the multi-level memory cooperative decision-making method as described above.
[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a multi-level memory collaborative decision-making method, apparatus, device, and medium, comprising: acquiring raw observation data of the environment; updating a multi-level memory module for storing environmental information, interaction history, and operation instructions; retrieving information from the updated multi-level memory module and generating candidate communication content; extracting high-level actions from operation instructions based on the current task objective to form a list of candidate action plans; analyzing the list of candidate action plans and candidate communication content through the thought chain reasoning of a language model to determine whether to send the communication content; if sending is determined, including the communication content in a set of actions to be executed; generating a final action plan based on the list of candidate action plans and the set of actions to be executed, and decomposing the final action plan into basic operation instructions to achieve interaction with the environment. This invention integrates the environment, history, and task instructions through the reasoning ability of a language model combined with a multi-level memory structure. In complex environments with communication costs and perceptual incompleteness constraints, it achieves collaborative optimization of task planning and communication behavior among agents, thereby improving the response efficiency, decision consistency, and task execution capability of multi-agent systems in real-world scenarios. Attached Figure Description
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0025] Figure 1 This is a schematic diagram of an application environment for a multi-level memory collaborative decision-making method according to an embodiment of the present invention;
[0026] Figure 2 This is a flowchart illustrating an embodiment of the multi-level memory collaborative decision-making method of the present invention;
[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multi-level memory collaborative decision-making device of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0029] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0030] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0031] The multi-level memory collaborative decision-making method provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain raw environmental observation data from the user terminal and update the multi-level memory module used to store environmental information, interaction history, and operation instructions. It retrieves information from the updated multi-level memory module and generates candidate communication content. Based on the current task objective, it extracts high-level actions from the operation instructions to form a list of candidate action plans. Through the reasoning chain of a language model, it analyzes the list of candidate action plans and the candidate communication content to determine whether to send the communication content. If sending is determined, the communication content is included in the set of actions to be executed. Based on the list of candidate action plans and the set of actions to be executed, a final action plan is generated, and the final action plan is decomposed into basic operation instructions to achieve interaction with the environment. This invention integrates the environment, history, and task instructions through the reasoning ability of a language model combined with a multi-level memory structure. In complex environments with communication costs and perceptual incompleteness constraints, it achieves collaborative optimization of task planning and communication behavior among agents, thereby improving the response efficiency, decision consistency, and task execution capability of multi-agent systems in real-world scenarios. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multi-level memory collaborative decision-making method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0033] like Figure 2 As shown, the multi-level memory collaborative decision-making method proposed in this invention includes the following steps:
[0034] S10, acquire raw environmental observation data, and update the multi-level memory module used to store environmental information, interaction history and operation instructions based on the raw observation data;
[0035] In this embodiment, the actuator collects observational information from the current physical environment through a perception module. This observational information may include image data, depth information, radar scan results, voice input, temperature and humidity sensor values, tactile feedback, or other forms of raw multimodal input. The entry point for environmental perception typically consists of a set of front-end acquisition hardware, including a camera array, LiDAR, inertial navigation device, microphone array, or robot tactile sensors. Each perception source has independent information dimensions for different scenarios. After the input is converted into digital form by the signal acquisition circuit, it is processed through image decoding, filtering, and normalization before being transmitted to the perception processing unit.
[0036] The perception processing unit needs to perform feature extraction on the acquired environmental images or other modal data. It uses image segmentation networks, object detection networks, or semantic recognition models to label key elements in the physical environment. The extracted content includes spatial layout information, object category labels, relative positional relationships, changing states, pose parameters, and texture features. The processing results are structured and encapsulated according to spatial hierarchy and temporal dimensions, forming well-structured, indexable state vectors or semantic label sets. For non-visual information, such as speech signals, changes in environmental parameters, vibrations, and other event-type inputs, a dedicated modal adaptation module needs to be built to perform feature encoding, unifying it into a high-order state representation.
[0037] After the observation data is processed, an environmental information mapping relationship needs to be constructed. The current observation state is compared and updated with the environmental state stored in historical memory to assess which entities have been added, moved, or deleted, thereby maintaining a continuously evolving spatial semantic model. This model is stored in a structured memory architecture in the form of a hierarchical cache, divided into three logical partitions: semantic memory, episodic memory, and procedural memory. Semantic memory is used to store static or slowly changing scene elements and spatial layout information for a long time. Episodic memory records the interaction history organized chronologically, including communication segments, operation feedback, state transition sequences, etc., while procedural memory contains operation parameters, process templates, and historical execution paths related to task execution.
[0038] The interaction history is obtained by the control module's structured recording of each operation task, including instruction content, execution time, execution feedback, and communication semantic intent. After encapsulation, this information is written into the plot memory and updated chronologically. Operation instructions are stored in the program memory, containing action type, call parameters, operation target, expected result, and some instructions also include intermediate variables and control flow state generated during inference. The coordination of the three partitions relies on a dynamic indexing mechanism, which automatically associates the task context in state recognition, behavior recording, and policy invocation.
[0039] The update process of multi-level memory modules is typically controlled by state change triggers. Data writing is triggered when a significant change in the environmental state, task phase advancement, or communication response completion is detected. The update process includes operations such as structure compression, keyframe extraction, redundancy removal, and hierarchical fusion to ensure that long-term memory does not become excessively bloated and to maintain retrieval efficiency. The module structure supports asynchronous updates and partial refreshes to avoid global deadlock.
[0040] One implementation involves acquiring RGB image and depth map data via an embedded perception module, combining this with YOLO series detectors for object recognition, and using a SLAM system to estimate relative coordinates, extract entity positions and motion trajectories. The recognition results are then fused with a prior environment map to construct a local scene map, and a vector database is used to record labels and location information. This scene map is input into a Transformer structure to embed semantic features, ultimately transforming it into a unified state vector, which is then written into semantic memory.
[0041] Another approach is to use a speech recognition model to extract natural language descriptions in a speech-aware scenario. This description is then input into a language model and combined with the current context to extract key objects and behaviors, serving as a structured expression of the current plot. A behavior controller then records the operation sequence and execution feedback, packaging and writing this information into the plot memory structure. The operation flow is then handled by the execution engine, which retrieves the execution record template according to a dynamic scheduling strategy, loads parameters, forms a control flow diagram, and writes it into the program memory.
[0042] In real-time scenarios where the system needs to support high-frequency updates, a sliding window mechanism can be used to partially refresh the semantic memory content; a state checkpoint mechanism can be used for program memory, with snapshot updates only performed during process switching or parameter adjustments. The three can be dynamically bound together through hash indexes and entity relationship graphs built from graph databases.
[0043] Example: In the healthcare sector, the executor is deployed in a hospital triage robot. Using cameras and microphones, it captures images of the waiting area and patient inquiries, identifying patient location, queue order, and intent, and storing this structured information in a multi-level memory structure. In subsequent interactions, the robot can combine its current spatial state with past inquiry records to recommend registration windows or provide guidance routes.
[0044] In the fintech business, the execution unit is deployed in branch service terminals to collect customer counter operation behavior and voice communication content, identify customer identity, intent, and historical operation trajectory, and embed these states into multi-level memory. When customers repeatedly visit or perform complex business applications, the system can automatically link past records, improve interaction efficiency, and avoid repetitive operations.
[0045] This embodiment constructs a unified state representation based on perceptual input and organizes it into a multi-level structured memory system, enabling the complete reconstruction of scenarios, behaviors, and planned instructions in multi-agent environments with incomplete information, spatiotemporal heterogeneity, and complex tasks. This design solves the problems of information fragmentation and memory loss caused by existing single-state caches and the lack of context maintenance mechanisms, providing a continuous, interpretable, and efficient data foundation for subsequent communication content generation and behavioral decision-making.
[0046] S20, retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information;
[0047] In this embodiment, after being updated, the multi-level memory module structures and stores the environment state, interaction history, and operation instruction information. This information is indexed and organized according to semantic and temporal dimensions to form a storage system that supports semantic retrieval. The core of information retrieval lies in constructing a retrieval query vector or retrieval intent description based on the current interaction context, and filtering content relevant to the current context from different memory layers through vector matching or structural pattern comparison.
[0048] Retrieval operations typically construct retrieval targets based on the following types of inputs: current task objectives, dynamic changes in the environmental state, and previous communication intent or operation execution feedback. For the semantic memory layer, the primary retrieval focuses on object attributes, spatial relationships, or fixed rules related to the current scene; for the plot memory layer, the focus is on retrieving past interactive language fragments, task interruption locations, and contextual intent expressions; and for the program memory layer, the focus is on existing operation patterns, conditional judgments, and preset behavior sequences within the task execution logic.
[0049] To improve retrieval accuracy and semantic relevance, an embedding coding model can be used to perform a unified semantic mapping between the current input information and the memory content. An attention mechanism is then used to select the set of information that best matches the current task from a high-dimensional vector. The retrieval results do not directly form the response output; instead, they are used to construct the semantic basis for communication candidate content.
[0050] The generation of candidate communication content requires structurally fusing retrieved environmental information, historical language interaction records, program paths, and other content to form a natural language expression with complete semantic intent. Content generation typically relies on language models for contextual modeling and semantic completion. During this process, the current task objective and environmental state are introduced as cue signals to drive the model to construct multiple communication expressions that meet the current needs. The generated results include multiple candidate response texts, each with interpretable intent labels, triggering conditions, and potential behavioral consequences.
[0051] The generated communication content also needs to undergo compliance checks and... Figure 1 Consistency checks and semantic conflict screening with existing candidate actions ensure the effectiveness and minimal redundancy of candidate communication content in the current multi-agent collaborative scenario.
[0052] In one approach, a dual-tower encoding structure is used to bidirectionally encode the current state information and the memory content. Inner product matching is used to filter out several closest interaction history fragments from the episodic memory, and these fragments, along with the current semantic embedding, are input into a large language model. Multiple candidate communication sentences are then generated using structural cue templates. The template structure used includes a target entity description, a summary of the preceding dialogue, and a current task label to enhance the consistency of communication generation.
[0053] Another approach is to use a graph neural network to represent the environmental relationship graph in memory, propagate the similarity between the current observation state and the nodes in the graph, find the historical event node that is most closely related to the current node, then extract language content, behavioral intention and target object from the node, construct communication candidate fragments, and then use the language generation module to complete the natural language expression.
[0054] In scenarios with high-density multi-agent interactions, a priority mechanism can be introduced to score and rank the generated candidate communication content according to urgency, policy relevance, and context coverage, and select the top-ranked communication content to input into the downstream inference process.
[0055] Example: In healthcare scenarios, a triage robot can retrieve a patient's historical consultation records, previously visited departments, and consultation content from multi-level memory, and generate candidate response statements based on the patient's current location and purpose of visit, such as: "You consulted the cardiology department last time. If you need a follow-up visit, you can go to Section B on the second floor."
[0056] In fintech business scenarios, service terminals generate communication content with reminder and guidance functions based on customers' past loan application history, recent credit score changes, and current account activity. For example, "Based on your recent credit activity, the system suggests that you postpone the submission of your current loan application." This content is derived from deep information retrieval and contextual understanding in the memory module, supporting the intelligent agent to accurately complete high-value task collaboration.
[0057] This embodiment improves the efficiency, coherence, and operational guidance of language interaction in multi-agent systems by selectively retrieving semantic association information from the updated multi-level memory module and constructing communication content that conforms to the current task context. This ensures that communication generation is not only context-consistent but also task-oriented, thus avoiding the problems of redundant language output that traditional communication generation relies on static templates or lacks contextual connections, significantly reducing communication latency and improving the success rate of collaboration.
[0058] S30, based on the current task objective, extract executable high-level actions from the operation instructions of the updated multi-level memory module to form a list of candidate action plans;
[0059] In this embodiment, the current task objective is typically input into the system in the form of natural language expression, task label vectors, or task parameter configuration to guide the embodied agent's intention planning and behavior selection. Upon receiving the current task objective, the system needs to parse the objective content and extract key operational intentions, target objects, expected states, or constraints. This parsing process can be combined with the task understanding module, using structured template matching, task type classification models, or target state prediction modules to ensure that the objective information is transformed into a semantic representation suitable for structured reasoning.
[0060] The operation instructions in the updated multi-level memory module refer to instruction units stored in the system history that have clear semantic labels and action execution boundaries. They are usually derived from the accumulation of experience from previous tasks or expert-preset behavior templates. These instructions are associated with environmental states, object attributes, and execution contexts during storage, and have contextual dependencies and action sequence logic.
[0061] To extract executable high-level actions, the semantic structure of the current task target needs to be matched with the stored operation instructions. The matching process can employ a bidirectional encoder structure, encoding the task target as a query vector and the operation instructions as memory vectors. High-matching operation instructions are then filtered out in the embedding space through similarity calculation. Alternatively, a rule-based action adapter can be used to directly compare the operation verbs and target entities in the task target with the behavior labels and object descriptions in the action instructions, filtering for high-level actions with behavioral adaptability.
[0062] High-level actions refer to complex behavioral units with a high level of abstraction that can be expanded into multiple basic operation sequences, such as "completing identity verification," "entering the treatment process," and "initiating transaction confirmation." These high-level actions are characterized by complete structure, clear goal orientation, and strong context adaptability, making them suitable for plan construction and strategy deployment. The extracted high-level actions are prioritized according to their relevance to the current task objective, contextual consistency, and feasibility, forming a structured list of candidate action plans, which serves as input for subsequent behavior planning processes.
[0063] Each high-level action in this list includes an action name, triggering conditions, expected target state, contextual dependency description, and a range of optional control parameters, forming a task execution candidate structure with a unified semantic format. The process of forming the list has a logical causal relationship and operational dependencies, supporting subsequent conflict analysis and plan optimization.
[0064] In one implementation, a bidirectional semantic matching model based on the Transformer structure is used to transform the current task target into a target vector representation. At the same time, operation instructions are extracted in batches from multi-level memory modules, an action embedding vector is constructed for each instruction, the semantic distance between the two is calculated, instructions with a distance less than a preset threshold are selected as executable high-level actions, and structured items are added to the candidate list.
[0065] In another implementation, an action classifier can be built based on the action tag tree. The task target is input into the classifier, which outputs a list of possible matching high-level action categories. Then, the corresponding instructions are indexed from the memory module by tag and the current environment is checked to see if the instructions can be executed, such as whether the target object exists or whether the preset resources are available. The high-level actions that can be executed in the current context are further filtered out, and finally the list is built.
[0066] The task objective and system state can also be input into the conditional generation network to output a candidate set of potential high-level actions. This set is then aligned and matched with known instructions in the multi-level memory module to form a list after confirming structural integrity and contextual relevance.
[0067] Example description: In a healthcare business scenario, when the task objective is "to assist in completing the initial diagnosis of a patient", the system can extract high-level actions related to the initial diagnosis from the multi-level memory module, such as "calling the symptom consultation process", "guiding the filling of electronic medical records", and "confirming registration information", and build a list of pending action plans for subsequent behavior scheduling and process guidance.
[0068] In fintech business scenarios, if the task objective is to "issue transaction alerts to users with medium risk", then high-level actions such as "loading user risk tags", "matching historical transaction patterns" and "sending warning-level communication instructions" can be extracted and stored in a structured list format. This facilitates further integration with the communication judgment and plan generation modules for operation deployment and behavioral decision-making.
[0069] This embodiment combines the current task objective with the operation instructions in the multi-level memory module for semantic matching and adaptability analysis. This allows for the rapid selection of high-level actions executable in the current environment from a complex historical behavior database, constructing a structured list of action plan candidates. This process not only improves the adaptability and contextual consistency of task planning but also significantly shortens the response time from task understanding to action deployment, reduces the task failure rate, and improves the accuracy of behavior execution.
[0070] S40, using language model-based thought chain reasoning to analyze the list of candidate action plans and the candidate communication content, determine whether to send the candidate communication content;
[0071] In this embodiment, the language model not only needs to understand the content during the reasoning process, but also needs to have the ability to make comprehensive judgments about the task objectives, execution plans, and environmental states. During the pre-execution judgment process, the semantic dependencies, causal relationships, and task suitability between the action plan and the communication content need to be analyzed to determine whether to put the generated communication content into execution.
[0072] The list of potential action plans includes multiple high-level actions, each with its corresponding operational intent, triggering conditions, target state, and contextual requirements. Candidate communication content, derived from environmental understanding and historical interaction information, consists of expressive units with linguistic structure generated by the system based on the current state and semantic context. The correlation between the two reflects the causal logic and goal consistency between action and expression.
[0073] The thought chain reasoning structure, which incorporates a language model, aims to enhance multi-turn logical judgment capabilities. This allows the judgment process to move beyond single-step matching or superficial semantic alignment, enabling multi-layered deduction and causal analysis. The thought chain structure is typically built upon multi-hop retrieval and internal activation paths, including: task objective → action semantics → communication content → historical behavior → expected feedback → reasoning conclusion. In each hop, the system evaluates the consistency and necessity of the content based on the current input and intermediate states, and determines whether candidate communication content has positive benefits or collaborative value for the execution of subsequent actions.
[0074] During the reasoning process, the language model first encodes the goal-oriented semantics of the task objective and the high-level actions in the action plan. Then, candidate communication content is input into the language model's context window, forming a semantic intersection with the planned actions. Based on a language understanding model (such as chained prompt templates or structured logical units), the system outputs a thought process. The system can construct chained conclusions including reasons, background, impact, and expected feedback, ultimately outputting a Boolean decision or confidence score to determine whether to send the communication content.
[0075] This process can simultaneously consider factors such as the response effects in historical communications, the resource status of the current context, the identity or role of the communication object, and the priority and cost weight of the communication, so as to achieve controllable, robust, and context-sensitive communication judgment.
[0076] A chain-like reasoning template can be built using a language model, and a multi-round Prompt can be used to trigger the logical analysis process. First, the task objective, a list of alternative action plans, and candidate communication content are input into the language model in a structured manner, for example: "The current task objective is X, the planned actions include A, B, and C, and the communication candidate is Y. Should Y be sent as part of the current task?" Based on its pre-trained knowledge and context alignment capabilities, the language model will output a reasoning chain containing explanations and suggestions. The system then reads the final suggestion as the basis for its judgment.
[0077] Alternatively, in engineering implementation, conditional control logic can be embedded into the language model prompt chain, such as setting a control template: "If the target action contains a human collaboration request, or the target state depends on an external response, then the communication candidate will be included in the sending sequence." This type of structure facilitates the introduction of a rule engine to dynamically constrain the generation judgment, thereby improving the consistency of the judgment.
[0078] Furthermore, by training a dedicated fine-tuned model with task labels and action semantic labels, it can accept task descriptions, action plans, and communication candidates as joint inputs and output judgment results and corresponding thought paths, thereby improving the generalization ability under multi-task conditions.
[0079] Example: In a healthcare scenario, when an action plan includes a high-level action such as "requesting the patient to provide their past medical history," the system uses language model reasoning to identify that the candidate communication content "Have you recently taken any antihypertensive medication?" is related to the target action and is irreplaceable. Therefore, it decides to include it in the communication execution sequence.
[0080] In fintech scenarios, when the task objective involves "judging the compliance of transaction behavior", the action plan includes "confirming customer information for suspected abnormal transactions", and the candidate communication content is "Have you authorized this cross-border transfer?" The language model can determine the necessity of the communication content based on the action plan structure and provide sending suggestions, thereby ensuring the consistency between the legality of the process and the user experience.
[0081] This embodiment utilizes the thought chain reasoning mechanism of a language model to enable the system to possess multi-layered logical judgment and contextual semantic inference capabilities, moving beyond superficial matching or template-based rule judgment. By combining the structured expression of task objectives and planned actions for semantic joint analysis, the decision-making process for initiating communication becomes more intelligent, rational, and possesses a cause-and-effect inference chain, significantly reducing the frequency of invalid communication and improving the efficiency of communication resource utilization and the consistency of interactive responses. Simultaneously, a semantic mapping path between communication intent and action plans is established, forming a closed-loop control structure between the behavioral chain and language output.
[0082] S50, when it is determined that the candidate communication content will be sent, the candidate communication content will be included in the set of actions to be executed;
[0083] In this embodiment, during the task execution preparation process, communication behaviors and physical operations should be managed uniformly as clearly schedulable execution units. The "sending candidate communication content" in this stage is a language interaction behavior with instruction attributes. After determining that its sending conditions are met, it should be included in the current task execution set as an action entity with a structured identifier, thereby ensuring that it is processed and executed together in subsequent action scheduling stages.
[0084] Action sets are a structured management mechanism used to organize and maintain all pending actions within the current cycle, including heterogeneous behavior types such as navigation, grasping, perception, and communication. Including language communication actions in this set means the system treats them as parallel, schedulable, and quantifiable task units. When generating action sets, each action unit includes key fields such as type flag, trigger condition, priority, and expected result. Language communication actions should explicitly identify their target object, communication method, language content, and execution window to facilitate unified resource allocation and execution by the scheduling module.
[0085] Transforming candidate communication content into actions to be executed requires two steps: semantic encapsulation and behavioral intent modeling. Semantic encapsulation transforms the communication content into structured language instructions, typically including: language output, target recipient, contextual background labels, and execution triggering conditions. Behavioral intent modeling then associates these instructions with the current task state, generating behavioral dependencies and scheduling conditions. Finally, this communication behavior is added to the set of actions to be executed, allowing subsequent scheduling modules to complete it according to sequential or concurrent strategies.
[0086] This approach ensures the continuity between the decision-making process and the execution process, so that language behavior is no longer a byproduct of the system's external output, but becomes part of the planned control mechanism, thereby achieving a unified scheduling management and resource allocation strategy.
[0087] By constructing a unified action representation structure, the content to be sent can be defined as an action type, and a dedicated communication action template structure can be defined in the system. For example, the action type field can be defined as "communication" in the structured representation, along with the text to be sent, the target object ID, context information, and priority level. When the inference module outputs a sending decision, the communication behavior is encapsulated as an action unit and written to the current action set queue.
[0088] Furthermore, communication node types can be defined in the behavior graph and connected to the high-level task graph structure, thereby enabling communication behaviors to have schedulable paths in the task planning graph. When updating the action set, the system can parse the task graph, automatically identify newly added communication nodes, and establish logical constraints between behaviors, such as waiting for upstream actions to complete before executing communication, or triggering downstream operations only after successful communication.
[0089] In implementation, the triggering conditions for communication behaviors can be bound to the current resource status through a behavior plan manager or scheduling engine. When the resource is available and the receiver is ready, the communication output module is triggered to complete the language generation and expression process. For example, the actual communication behavior can be completed by scheduling a TTS (Text-to-Speech) module or sending messages to a remote interface.
[0090] Example Explanation: In a healthcare scenario, after identifying the patient's basic information and current symptoms, the system plans "obtaining medication history" as a sub-objective within a higher-level task. The language model infers that it needs to ask the patient, "Have you recently taken any antihypertensive medication?" The system then incorporates this verbal inquiry into the current execution set as a communication action, uniformly scheduling it to the speech module for output. Simultaneously, it binds this to the "waiting for a response" action logic, achieving task closure.
[0091] In fintech scenarios, after identifying abnormal user transactions and planning the "verify transfer intent" action, the system generates the communication content "Did you authorize yesterday's cross-border transfer?" through a language model. The system then packages this communication action into a structured action unit and adds it to the execution set, setting the trigger conditions to "account status is normal" and "time window is working hours" to ensure that the communication behavior is executed smoothly under the premise that resources are available and the context is appropriate, thus ensuring the compliance of the task process and the accuracy of the communication effect.
[0092] This embodiment integrates confirmed communication content to be sent into a structured action set for execution, achieving unified management of language interaction and physical operation. This overcomes the limitation of independent execution of language output in traditional task planning, enabling language actions to be schedulable, traceable, and consistent with strategies. This mechanism enhances the system's ability to uniformly allocate resources and arrange action execution in complex tasks. Particularly in multimodal and multi-objective environments, it effectively reduces the conflict rate between language actions and action execution, improving the stability and consistency of the overall execution loop.
[0093] S60, based on the list of candidate action plans and the set of actions to be executed, generate a final action plan, decompose the final action plan into at least one basic operation instruction, and execute the basic operation instruction to interact with the environment.
[0094] In this embodiment, in a scenario where an embodied agent performs multi-objective tasks, the action plan not only needs to express high-level intentions but also must be refined into schedulable low-level actions and orchestrated with resources. This stage involves three consecutive and structurally coupled processes: action fusion, instruction decomposition, and action execution, whose logical order is strictly dependent on the output of preceding modules.
[0095] The list of potential action plans typically originates from the preceding action reasoning stage. It represents the high-level planned actions that the agent can consider executing under the current environmental conditions. Formally, it can be a structured set of action nodes, including fields such as action type, target state, dependencies, and logical priority. The set of actions to be executed consists of specific operations related to external interactions, such as language output, feedback reception, or environmental response triggering. These two sets of actions need to be integrated in this stage to generate a final action plan that is executable, resource-free, and has a reasonable temporal order.
[0096] The process of generating the final action plan includes resource conflict detection, action dependency graph construction, and time scheduling sequence generation. In resource conflict detection, the system needs to detect whether actions request the same sensor, actuator, or communication channel at the same time, and avoid this by adjusting the timing or replacing action paths. In dependency graph construction, all candidate actions are constructed into a directed graph structure, with edges representing dependencies, and the optimal execution order is generated through topological sorting or reinforcement learning methods. Scheduling sequence generation combines the current system state, task urgency, and historical success rate to generate a deterministic, ordered chain of actions.
[0097] After completing the above process, the system obtains the final action plan and decomposes the high-level actions into actions. The action decomposition module maps the high-level actions into multiple basic operation instructions based on program memory or action libraries. These instructions are feasible for low-level execution and typically include execution module call instructions, control parameter structures, and expected feedback conditions. For example, navigation-related high-level actions can be decomposed into a sequence of actions such as moving the starting point, calculating the path, adjusting the angle, and advancing the step size; language-related high-level actions can be decomposed into atomic instructions such as text generation, speech synthesis, and interface calls.
[0098] Ultimately, basic operational instructions will be scheduled and executed in the order outlined in the action plan, triggering environmental interactions. This process involves the coordinated operation of the instruction scheduling engine, resource management module, and perception feedback mechanism. The execution status of each basic instruction will be recorded, and the feedback results will be synchronized to multi-level memory modules, forming a closed-loop decision-making chain.
[0099] One approach is to use a task graph fusion mechanism, which integrates the list of candidate actions and the set of actions to be executed into a graph structure. Nodes in the graph represent action units, and edges represent dependencies or conflicts. The system can then generate conflict-free and shortest-duration action paths based on graph traversal strategies, such as minimum-cost priority traversal, and construct the final action plan.
[0100] Furthermore, execution priority strategy functions can be constructed to assign dynamic weights to different types of actions. For example, in medical tasks, treatment-related actions can be prioritized, while in financial transaction tasks, high-value verification operations can be prioritized. These strategy functions are dynamically adjusted based on historical task success rates, task urgency, and environmental response efficiency, making the final action plan adaptable to different scenarios.
[0101] During instruction decomposition, action templates stored in the program's memory can be used for deconstruction, with different high-level actions calling different template structures. The templates define the execution parameter range, instruction order, and intermediate state verification rules for the underlying operation instructions. After system invocation, an instruction sequence can be automatically generated. After decomposition, instructions can be queued for execution according to priority and resource occupancy status by the instruction scheduler.
[0102] It can also be combined with a parallel execution framework to enable a parallel scheduling mechanism between operation instructions that have no dependencies, and accelerate execution efficiency through thread pools or asynchronous calls, thereby improving the system's resource utilization and response speed.
[0103] Example Explanation: In a healthcare scenario, when faced with multiple parallel tasks, such as obtaining patient vital signs, retrieving historical examination records, and notifying a doctor, the system first generates high-level representations of these tasks and adds reminder-type language output actions to the execution set. Then, it integrates all actions to generate a unified action plan. The plan schedules notifying the doctor before obtaining basic information. Based on this plan, the system decomposes the language reminders into "generate statement," "select target doctor," and "execute announcement," triggering these operations in the planned sequence.
[0104] In fintech scenarios, customer identification, high-risk operation verification, and user alerts may occur simultaneously. The system will integrate various actions based on priority functions and resource scheduling graphs, generating a plan that proceeds "first verify identity, then confirm high-risk activity, and finally notify customer feedback." High-risk confirmation includes two actions: sending communication content and waiting for feedback. The system breaks down the communication behavior into basic instructions such as language generation, text push, and result recording, and calls these instructions in the execution module according to the scheduling order to complete the actual output of the customer interaction.
[0105] This embodiment constructs a final action plan that integrates high-level actions and specific behaviors, and decomposes high-level actions into basic operation instructions. The system achieves a closed-loop process from language reasoning generation to physical operation execution. In this mechanism, the agent can automatically adjust task order, allocate limited resources, and avoid concurrent conflicts, thereby fulfilling multi-objective, high-concurrency task execution requirements in complex environments and improving the continuity, coordination, and success rate of task responses.
[0106] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a multi-level memory collaborative decision-making method, apparatus, device, and medium, comprising: acquiring raw observation data of the environment; updating a multi-level memory module for storing environmental information, interaction history, and operation instructions; retrieving information from the updated multi-level memory module and generating candidate communication content; extracting high-level actions from operation instructions based on the current task objective to form a list of candidate action plans; analyzing the list of candidate action plans and candidate communication content through the thought chain reasoning of a language model to determine whether to send the communication content; if sending is determined, including the communication content in a set of actions to be executed; generating a final action plan based on the list of candidate action plans and the set of actions to be executed, and decomposing the final action plan into basic operation instructions to achieve interaction with the environment. This invention integrates the environment, history, and task instructions through the reasoning ability of a language model combined with a multi-level memory structure. In complex environments with communication costs and perceptual incompleteness constraints, it achieves collaborative optimization of task planning and communication behavior among agents, thereby improving the response efficiency, decision consistency, and task execution capability of multi-agent systems in real-world scenarios.
[0107] In one embodiment, step S10 includes:
[0108] S101 captures environmental image data as raw observation data using a visual sensor;
[0109] S102, perform object recognition on the environmental image data to determine object category information, and perform spatial positioning to determine object location information;
[0110] S103, construct a local semantic map based on the object category information and object location information, and use the local semantic map as environmental information;
[0111] S104, record the basic operation instruction types and execution results executed by the intelligent agent, and generate the current behavior sequence in combination with the actual communication content sent;
[0112] S105, load the initial operation process parameters from the predefined task template library, and dynamically update the operation process parameters based on the optimization instructions of the planning module;
[0113] S106, the environmental information is stored in the semantic memory partition of the multi-level memory module, the current behavior sequence is recorded in the plot memory partition of the multi-level memory module, and the operation process parameters are stored in the program memory partition of the multi-level memory module.
[0114] S107, Establish a dynamic index relationship between the semantic memory partition, the plot memory partition and the program memory partition based on the current task context.
[0115] In this embodiment, to achieve stable behavior generation and semantic alignment perception processing of the intelligent agent in a dynamic and complex environment, it is first necessary to acquire the raw observation data of the environment as the perception basis for updating the memory system. The raw observation data is collected through visual sensors, such as RGB cameras, depth cameras, and LiDAR with image acquisition units. Their output is mainly structured image frames or point cloud frames, possessing high spatial resolution and temporal continuity, suitable for the input requirements of real-time perception tasks. The raw image data is a two-dimensional or three-dimensional data array, which needs to undergo decoding and preprocessing operations to form an input tensor before entering the subsequent recognition module.
[0116] The processing flow for environmental image data comprises two parts: object recognition and spatial localization. Object recognition utilizes convolutional neural network-based target detection models, such as YOLO and Faster R-CNN, to output object category, confidence score, and bounding box information. The bounding boxes further participate in spatial localization, projecting depth information into 3D space and combining this with camera intrinsic parameters to calculate the object's center point coordinates and orientation angle, resulting in a stable spatial localization result. In this step, the recognition process is not limited to static objects but also encompasses dynamic individuals and environmental boundaries, thereby enhancing the agent's understanding of scene structure and task area.
[0117] After obtaining object category and location information, the system constructs a local semantic map based on the current time window. This map maps the spatial coordinates and category labels of identified objects into a structured graph or tensor matrix, forming semantic entity nodes and spatial topological relationships. Graph nodes represent specific targets, such as "hospital bed," "counter," and "checkout terminal," while edges represent spatial adjacency or functional paths, such as "navigation reachable" and "visual occlusion." The local semantic map preserves the structural characteristics of the environment in the current frame, providing spatial support for subsequent task planning, language generation, and behavior prediction. The map can be structured using a searchable graph, sparse tensor, or key-value dictionary, facilitating storage in semantic memory partitions.
[0118] After the perception mapping is completed, the behavioral history data is updated synchronously. The behavioral history includes two parts: basic operation instructions and execution results. Instructions are issued by the task scheduling module, such as navigation, pickup, and language generation; execution results are generated through sensor or system status feedback, representing whether the action was successful, failed, or an abnormal state. Simultaneously, communication content is also recorded as part of the behavioral history, including fields such as language output content, semantic tags, and target receiver. These behavioral elements are combined chronologically to form the current behavioral sequence, reflecting the agent's policy output trajectory in the current task phase.
[0119] Dynamic management of task flows relies on the organization and updating of operational flow parameters. Initial parameters can be loaded from a predefined task template library, indexed by task type, such as "identity verification process" or "medical order execution process." Each template embeds configurations such as flow nodes, judgment conditions, and optional branches. The optimization module can dynamically adjust existing flows based on the current environment, task progress, and feedback results, updating judgment logic, jump nodes, or target parameters. The data structure of flow parameters should be scalable and hierarchical, supporting staged nesting and insertion operations.
[0120] The multi-level memory module comprises three partitions: a semantic memory partition for storing local semantic maps, recording entities identified in the current environment and their semantic and spatial information; a plot memory partition for storing the current action sequence, reflecting changes in the action chain and communication process; and a program memory partition for storing operation flow parameters, carrying the evolution trajectory of task logic and state conditions. These three partitions are interconnected, requiring the construction of a dynamic index structure.
[0121] Dynamic index relationships are established through the current task context, which consists of task objectives, time status, and current process nodes. Based on a context matching mechanism, the system retrieves target entities from semantic memory, similar behavioral segments from episodic memory, and historical process fragments from program memory to construct a dynamic linkage path between task perception and execution. This structure can be constructed using vector similarity matching, graph neural network mapping, or semantic label-based matching functions, forming a cross-domain bridging mechanism between semantics, behavior, and program.
[0122] This embodiment enhances the agent's ability to model environmental states and improves the interpretability and traceability of behavioral logic and task execution by constructing a local semantic map, updating behavioral sequences and task flows within each cycle, and storing this information in a structured manner in a multi-level memory module. The introduction of a three-part structure effectively decouples perception, execution, and control information at the semantic layer and enables cross-dimensional linkage through a dynamic indexing mechanism, providing complete and context-consistent support for subsequent action plan generation, communication content reasoning, and collaborative decision-making. The system can achieve continuous contextual understanding, behavioral continuity, and task goal maintenance in unstructured environments, significantly improving execution efficiency and inter-agent collaboration quality under complex tasks.
[0123] In one embodiment, step S20 above includes:
[0124] S201, Select the target retrieval partition of the multi-level memory module based on the current task requirements;
[0125] S202, extract the environmental state description and task target description from the semantic memory partition in the target retrieval partition;
[0126] S203, retrieve historical communication records related to the current task from the plot memory partition in the target retrieval partition;
[0127] S204, Extract key object attributes from the environmental state description;
[0128] S205, Based on the task objective description and historical communication records, the task objective is associated with the collaboration intent to obtain the associated collaboration intent;
[0129] S206, Reassemble the historical communication records in chronological order to form a time series;
[0130] S207, Integrate the key object attributes, associated collaborative intentions, and time series to generate a structured template;
[0131] S208, the structured template is processed by the language model to generate candidate communication content, and the candidate communication content is temporarily stored as an internal deduction result.
[0132] In this embodiment, to generate communication content with contextual consistency and semantic adaptability in a multi-agent collaborative system, the system first determines the target partition to be retrieved in the multi-level memory modules based on the current task requirements. Task requirements can be analyzed in real-time by the task scheduling engine, considering the current environmental state, task type, and stage objectives, and mapped accordingly to semantic memory partitions, plot memory partitions, or program memory partitions. The selection logic for the target partition needs to be mapped and matched with the memory partition index fields through task semantic tags to ensure that the retrieved content has structural and contextual relevance.
[0133] After determining the target retrieval partition, environmental state descriptions and task objective descriptions are extracted from the semantic memory partition. Environmental state descriptions are sets of entity attributes constructed based on local semantic maps, including object categories, spatial locations, relational boundaries, and functional labels. Task objective descriptions originate from the task template or the semantics of task instructions in the current process node, typically represented by structured labels, natural language fragments, or symbolic logic expressions, and are used to identify the execution intent and objective achievement conditions of the current task stage.
[0134] The episode memory partition contains historical communication records arranged in chronological order. The system needs to retrieve communication trajectories related to the current task. This retrieval process is based on semantic similarity or task context alignment mechanisms. For example, it uses embedded semantic representation to calculate the cosine similarity between historical communication statements and the semantics of the current task, or it matches relevant records through task node tags. The retrieved communication content must meet the requirements of diversity and temporality, that is, it includes representative exchanges from multiple different task stages and retains their original chronological order.
[0135] After extracting the environmental state description, the system extracts key object attributes. These key object attributes include target entities and their attribute sets that are directly or potentially related to the task objective, such as "patient tags, examination type, and equipment status" in a medical scenario, and "customer level, transaction intent, and counter location" in a financial scenario. These attributes need to be extracted from the environmental description using entity recognition and contextual attribution algorithms and mapped into structured key-value pairs or vector representations, serving as the core content for subsequent semantic template construction.
[0136] Associating collaborative intent with task objective descriptions and historical communication records is a crucial step in generating the logical structure of linguistic expressions. The system uses a semantic fusion module to analyze the semantic connections between task objective semantics and historical communication semantics, identifying common strategies, behavioral intentions, or historical paths to construct a behavioral mapping relationship between the current task objective and past communication content. This process can enhance the language model's sensitivity to key entity alignment through attention mechanisms or construct a graph-like semantic mapping structure to enhance multi-hop semantic reasoning capabilities, thereby generating a collaborative intent representation that covers the semantic context of the current task.
[0137] To map historical communication behaviors into time-consistent language behavior trajectories, the system reassembles the retrieved communication records in chronological order to form a time-series structure. This time series not only retains the text of the communication content but also labels the initiating entity, the target agent, the behavioral background, and the context state, forming a multi-field, ordered set of interaction trajectories that provides procedural reasoning material for the language model.
[0138] By integrating key object attributes, associated collaborative intents, and time-series structures, a structured prompt template is constructed. This template structure organizes the above three types of elements into sections and uses explicit field naming to achieve content isolation and semantic prompts. For example, the template can contain object attribute field blocks, intent label field blocks, and time-series segment field blocks, guiding the language model to reason and generate according to the given structure through explicit prompts or template instructions.
[0139] After receiving a structured prompt template, the language model, based on its built-in language understanding and generation capabilities, embeds and aligns the template with the context, performing thought-chain reasoning to generate candidate communication content. This process not only generates language based on static semantics but also integrates temporal cues and task context, outputting multi-turn dialogue content or collaborative statements with behavioral coherence and semantic rationality. The generated results are output in the form of vectorized representations or natural language representations and are temporarily stored in an internal inference cache, awaiting processing by subsequent intent verification, feasibility assessment, and execution decision modules.
[0140] This embodiment effectively improves the relevance, contextual fit, and linguistic coherence of communication content in a multi-agent system by introducing a structured prompt template mechanism, combined with the joint retrieval path design of semantic memory and episodic memory, and language model inference processing strategies. Compared to communication methods directly generated based on task instructions, this method can proactively associate historical behaviors with the current goal, enhance the context embedding effect, and guide the language model to focus on key semantic elements through explicit structural organization, reducing the probability of redundant communication. In addition, the temporal reorganization of historical communication and the construction of collaborative intentions enable the communication content to have behavioral chain continuity, which helps to establish stronger semantic consensus among agents, thereby achieving an efficient, low-misunderstanding, and strategy-consistent natural language interaction system in complex task collaboration.
[0141] In one embodiment, step S30 above includes:
[0142] S301, Analyze the current task objective and decompose it into core task elements and execution constraints;
[0143] S302, retrieve the set of operation instructions that match the core task element from the program memory partition of the multi-level memory module;
[0144] S303, verify the executability of each instruction in the set of operation instructions based on the environmental state description, and filter out the initial high-level actions through executability verification;
[0145] S304, Sort the initial high-level actions according to task priority, and filter the final high-level actions that meet the execution constraints.
[0146] S305, organize the final high-level actions according to the logical execution order to form a list of action plans to be selected.
[0147] In this embodiment, to form a sequence of actions with context adaptability and logical coherence in the intelligent agent autonomous decision-making system, the system first needs to parse the current task objective to decompose it into core task elements and execution constraints. The task objective is typically input in the form of natural language commands, semantic tag sets, or structured task descriptions. The system needs to semantically vectorize this input content through an embedded encoding module or semantic parsing network, and combine this with multi-round contextual judgments to determine the functional objectives, constraint boundaries, resource limitations, and behavioral styles required for the current task. This decomposition process not only outputs operation content tags but also extracts the hard constraints and preference-based constraints of task execution, serving as the basis for subsequent selection of action instructions.
[0148] After completing the task semantic parsing, the system retrieves a set of operation instructions that match the extracted core task elements from the program memory partition within the multi-level memory module. The program memory partition stores structured operation instruction sets generated during historical task execution; each instruction includes fields such as operation category, parameter template, context usage boundaries, and expected behavioral effect. Matching can be performed through semantic vector alignment, tag tree mapping, or multimodal retrieval mechanisms. The system constructs a multi-level matching graph between the task semantic vector and the operation instruction semantic vector to obtain a set of candidate instructions covering the task's target semantic space.
[0149] For the retrieved set of operation instructions, the system further verifies the executability of each instruction based on the current environment state description. The environment state description originates from semantic memory partitions and is a structured set of environment information generated through visual analysis and semantic mapping mechanisms. The execution prerequisites for each operation instruction are compared with the current environment state, such as determining whether the target object exists, whether environment variables are satisfied, whether the interaction path is unobstructed, and whether the required resources are available. This verification process is implemented through Boolean logic chains and constraint reasoning modules to output the judgment result of whether each instruction meets the current execution conditions, filtering out the initial set of high-level actions.
[0150] After obtaining the initial set of high-level actions, the system sorts and filters these actions. The sorting process is based on task priority, which can be derived from urgency labels assigned to the task itself, priority tables specified by an external policy management system, or current impact indicators dynamically calculated by the plan generation module. The filtering logic is based on previously extracted execution constraints, such as whether an action will cause a state conflict, whether resources are sufficient to meet concurrent calls, and whether certain steps should be delayed. By establishing constraint verification paths, instructions that do not meet the constraints are excluded, retaining action instructions that can be executed in the current context.
[0151] Finally, the system organizes the selected high-level actions into a list of candidate action plans according to their logical execution order. The logical order is typically constructed based on preconditions, resource dependencies, state transition mappings, or pre-existing task flowcharts within the system. For example, when multiple high-level actions have execution order dependencies, mutual exclusion relationships, or overlapping paths, the system will arrange them in an ordered manner based on causal chain graphs or topological sorting algorithms. The generated list of candidate action plans includes semantic labels, parameter structures, and expected state change labels for the high-level actions, serving as input to subsequent inference and decision-making modules. This supports semantic consistency and behavioral rationality for the agent in communication planning, behavior selection, and task collaboration.
[0152] This embodiment effectively enhances the contextual adaptability of task decision-making and the executability of action plans in multi-agent systems by tightly integrating task objective semantic parsing with operation instruction structure retrieval, environmental state verification, constraint filtering, and logical sequence organization. Compared to the traditional direct task template invocation mode, this process introduces a dynamic environment verification and constraint-driven instruction filtering mechanism, significantly reducing the generation rate of unreasonable actions. Simultaneously, by sorting and organizing the logical sequence, the constructed list of candidate action plans not only possesses contextual consistency and behavioral continuity but also provides structured support for subsequent communication decisions and collaborative path optimization, enhancing the continuity of actions and semantic coordination among agents in complex task environments.
[0153] In one embodiment, step S40 above includes:
[0154] S401, construct a structured multi-select prompt containing the list of candidate action plans and candidate communication content;
[0155] S402, using a language model to perform multi-step thought chain reasoning on the structured multiple-choice prompts, generating correlation analysis results between each action plan and communication content;
[0156] S403, Based on the correlation analysis results, dynamically quantify the task benefits and communication costs of sending candidate communication content;
[0157] S404, when the task benefit exceeds a preset benefit threshold and the communication cost is lower than a preset cost threshold, it is determined that the communication sending condition is met.
[0158] S405, generate a decision result on whether to send the candidate communication content, and generate action plan optimization suggestions based on the correlation analysis results.
[0159] In this embodiment, to achieve intelligent filtering of communication behaviors and strategy optimization driven by task value, it is necessary to construct a structured input compatible with semantic reasoning and benefit evaluation mechanisms. First, a structured multi-choice prompt containing a list of candidate action plans and candidate communication content is constructed. This prompt organizes information from different sources using a unified format, such as using tree-like labels to represent the semantic hierarchy of each high-level action, and is accompanied by context embedding vectors or summary templates of candidate communication content, enabling it to be used as input for unified reasoning in the language model. The structured prompt can be constructed using various methods such as JSON structure, nested vector tables, and sequence templates, depending on the input format supported by the language model. Furthermore, a token mapping table is used to number and semantically align the action tags and communication intentions within the prompt.
[0160] When performing reasoning, the language model does not directly output a single judgment on whether to send communication content, but instead performs multi-step thought chain reasoning. Thought chain reasoning refers to the language model simulating human reasoning paths, progressively evaluating the strategic relevance, synergistic necessity, and semantic consistency between each action plan and the communication content through multiple rounds of self-consistent logical analysis. For example, in the first round of reasoning, the model might identify whether candidate communication responds to the sub-task requirements of the current task objective; in the second round, it further determines whether the communication affects the behavior of other agents; subsequent rounds may also deduce the beneficial effect of the communication content on the overall system efficiency in the task evolution path. The intermediate results of each round of reasoning can be retained as a temporary memory structure for the language model, supporting the generation of complex reasoning paths.
[0161] Based on the intermediate analysis results obtained during the reasoning process, the system needs to further dynamically quantify the expected benefits and costs of communication behaviors. Task benefits can be derived through comprehensive modeling of indicators such as changes in expected task progress, the probability of other agents responding to the communication content, and the reduction in expected completion time. Communication costs include quantifiable parameters such as the number of tokens required for language model generation, bandwidth consumption, and the probability of potential collaborative conflicts. This quantification mechanism does not rely on fixed parameter settings but dynamically adjusts weights according to the current context. For example, it relaxes the benefit threshold in urgent tasks and increases the weight of communication costs in bandwidth-constrained scenarios, reflecting its adaptability to real-world collaborative environments.
[0162] Based on the aforementioned benefit and cost assessment results, the system uses simple logical conditions to determine whether the threshold conditions for sending communication content are met. When the task benefit index exceeds a set threshold and the communication cost is lower than the corresponding cost threshold, the communication sending conditions are considered met, and a decision result is generated marking the communication content as ready to be sent. This result not only drives communication behavior but also guides the fine-tuning of subsequent action plans. To achieve feedforward optimization and knowledge accumulation, the system outputs suggestions for optimizing candidate action plans based on intermediate reasoning conclusions in the thought chain reasoning process, including priority adjustments, content reorganization, or strategy rearrangement, to optimize semantic consistency and communication rhythm in the next round of plan generation.
[0163] This embodiment introduces a structured multi-choice prompt construction mechanism and a multi-turn thought chain reasoning path using a language model to achieve semantic mapping analysis and logical decision-making fusion between action plans and communication content, thereby improving the goal orientation and task benefit sensitivity of communication behavior. This process fully utilizes the reasoning and contextual understanding capabilities of the language model to dynamically assess the necessity of communication, thus avoiding redundant information interfering with the agent collaboration process. Simultaneously, the quantitative evaluation module for communication benefits and costs introduces decision flexibility, supporting adjustments to communication triggering conditions based on environmental conditions and task urgency, improving the decision rationality and resource scheduling efficiency of the multi-agent system in real-world scenarios. The accompanying action plan optimization suggestion output mechanism also prompts agents to continuously self-correct in future actions, achieving high consistency and synergy between behavior and communication in long-term task execution.
[0164] In one embodiment, step S50 above includes:
[0165] S501, when it is determined to send the candidate communication content, extract the receiver identifier and core semantic payload from the candidate communication content;
[0166] S502, construct a sending action entity that includes the receiver identifier and the core semantic payload;
[0167] S503, detect whether the time window of the sending action entity overlaps with the time window of the existing actions in the set of actions to be executed, whether the semantics conflict, and whether the resource consumption is lower than the communication bandwidth threshold configured by the system.
[0168] S504, when the time windows do not overlap, the semantics do not conflict, and the resource consumption is lower than the communication bandwidth threshold configured by the system, the sending action entity is dynamically added to the set of actions to be executed.
[0169] In this embodiment, after candidate communication content is determined to have execution value, it needs to be formally incorporated into the next action execution mechanism to achieve unified scheduling of language communication and entity behavior. First, the receiver identifier and core semantic payload are extracted from the candidate communication content. The receiver identifier is used to uniquely identify the target communication object in the multi-agent system and can be a static ID, IP address, logical role name, or a path reference dynamically generated in the task context. The core semantic payload refers to the smallest semantic unit in the communication content that carries the intention expression, such as action requests, status inquiries, target descriptions, or strategy suggestions. It is often encoded using structured templates to ensure that subsequent processing modules can parse and recognize it.
[0170] After extraction, the system constructs a sending action entity. This entity is considered a timing instruction equivalent to a physical action and contains at least three components: the target address or role name, the encoded semantic content block, and the expected sending time window or priority label. To adapt to asynchronous execution environments, the sending action entity typically has event-triggered attributes or a self-scheduling mechanism, and can be suspended, activated, or replaced based on the system scheduling queue.
[0171] Before being formally added to the action queue, the sending action entity needs to undergo multiple consistency checks to maintain the temporal consistency, semantic coherence, and resource constraint stability of the agent's behavior. Specifically, this includes: first, determining whether the time window overlaps with other actions in the current set of actions to be executed. This time window can be a precise timestamp or a soft boundary interval; if there is an overlap, it is considered a potential concurrent conflict. Next, analyzing whether the semantic content conflicts, such as whether it contradicts the communication goals in the current queue, involves duplicate requests, or causes unnecessary confirmation burdens. Finally, assessing whether the resources required for the communication task meet the system's communication bandwidth limit, including message size, channel contention status, and the number of concurrent communications. The communication bandwidth threshold is defined by the system policy and can be dynamically adjusted based on node hardware capabilities, network topology changes, or external interference, reflecting the system's elastic bandwidth management capabilities.
[0172] Only when all three checks pass—no time conflict, no semantic conflict, and sufficient bandwidth—will the system dynamically add the sending action entity to the set of actions to be executed. Dynamic addition means that the action queue can add the action without interrupting the existing queue structure, supports hot-swapping and priority adjustment mechanisms, and can record the source of addition, triggering reason, and expected impact scope, providing contextual support for future backtracking analysis or optimization.
[0173] This embodiment introduces communication behavior as a manageable action entity into the action scheduling system, achieving a closed-loop fusion of language model inference output and system behavior control. It also introduces refined control logic across three dimensions: time scheduling, semantic consistency, and resource consumption. Compared to existing methods that directly execute communication behavior, this mechanism significantly reduces the risks of timing conflicts and semantic confusion caused by communication, and effectively suppresses performance degradation caused by resource contention through bandwidth occupancy verification. The system can dynamically assess the feasibility of executing communication content based on real-time task status and system resource load, thereby improving the response accuracy and execution efficiency of multi-agent collaborative systems in dynamic environments, and achieving a unified fusion of semantic-driven communication scheduling and entity behavior control mechanisms.
[0174] In one embodiment, step S60 above includes:
[0175] S601, integrate the planned actions in the list of candidate action plans with the communication actions in the set of actions to be executed to form a preliminary action plan sequence;
[0176] S602, perform resource conflict detection and logical conflict resolution on the preliminary action plan sequence to generate an optimized conflict-free action plan;
[0177] S603, based on task priority and action dependency, the conflict-free action plan is time-series arranged to generate the final action plan;
[0178] S604, extract each high-level action from the final action plan and retrieve the corresponding operation flow from the program memory partition of the multi-level memory module;
[0179] S605, according to the operation process, each high-level action is decomposed into basic operation instructions, and the control parameters corresponding to the basic operation instructions stored in the program memory partition are loaded.
[0180] S606, in the environment, execute basic operation instructions with control parameters in the sequence of the final action plan to complete navigation, grasping or communication interaction operations.
[0181] In this embodiment, before forming a behavior path that can be actually executed, it is necessary to fuse the two different dimensions of action sources output by the behavior decision module. First, the planned actions generated by task planning in the candidate action plan list are integrated with the communication or external trigger actions in the set of actions to be executed, and a preliminary action plan sequence is generated. The planned actions usually come from the deduction results after the reasoning module parses the task objectives, while the communication actions come from the output behavior determined to be sent after the candidate communication content is evaluated. This integration needs to retain the semantic category, resource requirements, priority, and expected execution window of each action to support subsequent orchestration logic.
[0182] After generating the initial action plan sequence, the conflict resolution process begins. The system needs to model the resource usage and logical dependencies of all actions to determine if there are overlapping execution times, overlapping resource calls, or conflicting operation goals. Resource conflict detection is achieved by constructing a resource usage matrix, where matrix elements represent whether there is contention for the same physical execution unit within a given time period, such as robotic arm control, navigation paths, and camera field of view. Logical conflicts are resolved using a behavioral graph model to determine if there are semantic mutual exclusions or sequential conflicts, such as failure to complete positioning before grasping or issuing state-dependent communication requests before confirming the state. After conflict detection, conflicts are resolved through mechanisms such as replacement, postponement, decomposition, or merging to generate an executable but conflict-free action plan.
[0183] Based on action priority and dependencies, conflict-free action plans are input into the timing orchestration module. Priorities can be specified based on task urgency, resource scarcity, or task strategy; dependencies include explicit pre-execution requirements and communication prerequisite confirmations. The system generates the final action plan through topology sorting and weighted timeline optimization algorithms, ensuring that high-priority actions are implemented first, the action order on the dependency chain is legal, and resource utilization and execution efficiency are maximized.
[0184] For each high-level action in the final action plan, the operation flow definition stored in the program memory partition needs to be traced back to complete the mapping process from high-level actions to basic operation instructions. The operation flow is the bridge between behavior and execution, typically defined as a state transition diagram or process node diagram, containing the control signals, state judgment conditions, and feedback callback methods required for each node. Basic operation instructions may include navigation instructions, gripping actions, and sending requests, each of which must be bound to specific control parameters, such as movement distance, target point coordinates, gripping force, and communication bandwidth ratio.
[0185] After loading the control parameters, the system executes each instruction sequentially according to the timing sequence. The execution module is responsible for translating the structured instructions into low-level call interfaces and transmitting them to the control hardware or communication module. The operation results can be fed back to the memory module in real time for subsequent state updates and behavior corrections. The execution process involves multimodal interaction with the environment, including exchanging information with other entities in the environment through visual, tactile, or verbal channels, to achieve true semantic closed-loop control.
[0186] Example Description: In the emergency room reception hall, an intelligent triage robot with multimodal perception and multi-level behavior planning capabilities is deployed. Its goal is to assist doctors in quickly triaging patients, assigning test results, broadcasting patient status updates, and navigating patients when multiple patients arrive simultaneously. The system first acquires environmental image data of the hall using built-in vision sensors, including the number of patients, wheelchair trajectories, and the distribution of people at consultation windows. By recognizing elements such as patient tags, wheelchair markings, and doctor uniform colors, the robot generates a semantically annotated local map, marking areas for critically ill patients requiring priority care and available auxiliary resources such as vacant consultation rooms and wheelchairs.
[0187] Simultaneously, the system records previously issued navigation instructions, medical task distribution requests, and broadcast messages in real time, and combines these with currently received collaboration requests, such as "Please complete the pre-blood pressure collection process for patients in Zone B," to form a sequence of actions. The system loads the collaboration instruction parameters corresponding to the current emergency room status into the task template library, such as "Quickly establish patient trajectory," "Allocate medical staff resources," and "Broadcast task status," and writes them to the memory module for dynamic updates.
[0188] Subsequently, based on the triage objective of "completing the status verification of three suspected hypertension patients and assisting in the initial screening," the robot initiated a behavior retrieval request, locking onto the action sequence in its program memory that matched "initial screening verification," such as "identifying patient location," "distributing electronic questionnaires," "collecting vital signs parameters," and "report synchronization." It then determined which instructions could be executed within a short time window based on the current patient location and status, and sorted their execution order according to the doctor's current availability and device usage, generating a list of candidate actions.
[0189] Meanwhile, based on the task requirements, the robot generates candidate communication content for "whether to send a task progress broadcast to the head nurse." To determine whether this broadcast is necessary, the system constructs a structured input containing "the current hypertension screening task progress table + the head nurse's last received time," uses a language model for reasoning, evaluates whether sending this information can improve task completion efficiency, avoid state inconsistencies, and measures its communication cost. If the reasoning results indicate that its benefits exceed a threshold and the current network resource load is acceptable, then a broadcast task is generated and incorporated into the action plan.
[0190] After the broadcast task is confirmed, the robot extracts the head nurse's system identifier and communication semantic payload to form a communication action. The system analyzes whether this action overlaps with current tasks such as "collecting vital signs data" and "uploading reports," or whether it will occupy the same voice channel or visual feedback interface resources. If resource usage is acceptable and there is no semantic conflict, the communication task is added to the set of actions to be executed.
[0191] Afterwards, the system integrates all pending tasks (including navigation data collection, status synchronization, and broadcast transmission) with the planned inspection actions to form the final action plan. The sequence of actions will be adjusted according to patient distance, doctor availability, and channel congestion. For example, the robot may be scheduled to collect patient temperatures in area A before distributing questionnaires in area B. Each high-level action is then mapped into low-level operation instructions through program memory partitions, such as "navigate to bed A3 and activate the temperature sensor module" and "generate a voice inquiry and play it for 5 seconds," along with parameters such as speech rate, path coordinates, and broadcast frequency.
[0192] Ultimately, the robot executes these control commands sequentially, continuously interacting with the environment in a closed loop. During execution, all results are synchronized back to the multi-level memory module for context building and plan re-evaluation of subsequent tasks, thereby achieving a highly efficient, low-interference, and goal-aligned intelligent collaborative assistive behavior network in complex medical settings.
[0193] In the customer service platform of a comprehensive financial institution, an intelligent agent system with multi-level memory management and language understanding capabilities is deployed. Its responsibilities include handling high-frequency customer inquiries, distributing service instructions, coordinating internal operations personnel, and automatically constructing executable transaction plans based on context. During peak trading hours, the system first acquires raw observation data of the current environment through integrated speech recognition and text parsing modules, including customer voice inquiries, service personnel status information, and API response records. This data undergoes semantic annotation and location identification through a unified event-aware channel, and combined with the current API request source, customer tags, and communication window, a structured semantic state mapping is generated, forming an abstract representation of the environment.
[0194] The system combines historical customer service interactions with customer data retrieval records to construct behavioral sequence information. Simultaneously, it loads operation process definitions relevant to the current scenario from an internal business process template library, such as "high-net-worth customer instruction distribution priority strategy," "overdue account agreement negotiation process," and "abnormal transaction freezing process." This information is then written into semantic memory, plot memory, and program memory partitions, forming a complete multi-level contextual expression. The system dynamically establishes a cross-partition index structure to connect specific semantic concepts with past interactions and process entities.
[0195] Upon receiving a service request from a VIP customer regarding a "transaction limit adjustment application," the system automatically parses the request target, identifying its core task elements as "identity verification → transaction history verification → risk control approval → confirmation result feedback," and retrieves all execution paths and service nodes related to "transaction limit" from the program's memory partition. Further, it determines which operation nodes meet the conditions for immediate execution based on real-time environment status, such as "customer has completed authentication via the APP" and "risk control engine load is normal," thus constructing a list of executable high-level actions. Subsequently, the actions are prioritized, and actions that meet the execution constraints are combined to form a list of pending task executions.
[0196] The system synchronously generates candidate communication content for "transaction limit adjustment request push" to internal approval specialists. Based on task objectives, historical approval feedback, current approval load, and event time series, it constructs structured input including key object fields, collaboration intent, and interaction records, and performs multi-round thought chain reasoning through a language model. The reasoning process combines approval response latency statistics, customer sensitivity models, and current business response SLA indicators to evaluate the task benefit of the push, and compares it with the internal communication cost model. If the benefit assessment is higher than the threshold and communication resources have not reached the load limit, the system determines that the push operation is valuable, generates a decision result of "send immediately", and attaches an updated task recommendation path.
[0197] After a communication action is included in the set of actions to be executed, the system constructs a sending action entity by combining the target approval specialist identifier with the semantic payload in the push content. It then further analyzes the resource usage and timing conflicts between this action and actions in the existing task set. When it is confirmed that there is no channel overlap, no semantic conflict, and the message size does not exceed the system's configured communication bandwidth limit, the action entity is formally added to the set of actions to be executed.
[0198] During the integration phase, the system merges the action sequences generated by reasoning with the communication actions to be executed, generating a preliminary execution plan. Conflict detection and logic resolution are then performed in the unified resource scheduler to form an executable sequence. This sequence is time-orchestrated based on action priority and dependency structure, such as "authentication status refresh must be completed before verifying risk control limits" and "approval push should precede confirmation result upload," ultimately generating a structured execution trajectory.
[0199] The system extracts high-level actions from the final plan, calls control templates stored in the program's memory, and decomposes operations such as "querying historical transaction records," "verifying identity credentials," and "sending approval instructions" into basic call instructions, such as database access instructions, interface request formats, and parameter encryption transmission protocols. It loads the corresponding execution parameters and sends them out in real time. Each basic operation is executed in the optimal order, and the system retains contextual paths for subsequent task tracking, auditing, and review.
[0200] Ultimately, the intelligent agent completed a closed-loop task flow from task understanding and communication planning to instruction distribution and action execution. With multi-source environmental perception, complex logical reasoning, and dynamic resource coordination, it efficiently completed key service tasks, improved customer satisfaction, reduced response latency, and lowered the probability of communication conflicts in the approval and execution path.
[0201] This embodiment integrates the actions and communication behaviors generated by task planning into an executable final action plan, and introduces resource conflict detection, logical dependency analysis, and timing orchestration mechanisms. This effectively avoids control anomalies caused by the overlapping execution of multiple actions in complex environments. High-level actions are mapped to basic operation instructions and loaded with matching control parameters, enabling the system to possess not only reasoning capabilities but also actionable behavioral expression capabilities, forming a complete closed loop from semantics to physical execution. Simultaneously, by controlling the order of actions and dependency paths through the orchestration process, the system's execution efficiency and behavioral reliability are improved. In scenarios where dynamic task objectives change frequently or environmental states fluctuate rapidly, this mechanism ensures the stability of action scheduling and the consistency of behavioral results, enhancing the adaptability and semantic execution accuracy of the task-driven system.
[0202] In one embodiment, a multi-level memory collaborative decision-making device is provided, which corresponds one-to-one with the multi-level memory collaborative decision-making method described in the above embodiments. (Refer to...) Figure 3, Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multi-level memory collaborative decision-making device of the present invention. The modules include: environmental perception and memory update module 10, memory retrieval and communication construction module 20, target analysis and action extraction module 30, reasoning analysis and communication decision-making module 40, communication action generation module 50, and action orchestration and interactive execution module 60. Detailed descriptions of each functional module are as follows:
[0203] The environmental perception and memory update module 10 is used to acquire the original observation data of the environment and update the multi-level memory module for storing environmental information, interaction history and operation instructions based on the original observation data.
[0204] The memory retrieval and communication construction module 20 is used to retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information.
[0205] The target analysis and action extraction module 30 is used to extract executable high-level actions from the operation instructions of the updated multi-level memory module based on the current task target, and form a list of candidate action plans.
[0206] The reasoning analysis and communication decision module 40 is used to perform reasoning analysis based on the language model's thought chain to determine whether to send the candidate communication content;
[0207] The communication action generation module 50 is used to include the candidate communication content to be sent into the set of actions to be executed when it is determined that the candidate communication content should be sent.
[0208] The action orchestration and interaction execution module 60 is used to generate a final action plan based on the list of candidate action plans and the set of actions to be executed, and to decompose the final action plan into at least one basic operation instruction, and to execute the basic operation instruction to interact with the environment.
[0209] In one embodiment, the environment perception and memory update module 10 is specifically used for:
[0210] Environmental image data is captured using a visual sensor as raw observation data;
[0211] The environmental image data is subjected to object recognition to determine object category information, and spatial positioning is performed to determine object location information;
[0212] A local semantic map is constructed based on the object category information and object location information, and the local semantic map is used as environmental information.
[0213] Record the basic operation command types and execution results executed by the intelligent agent, and generate the current behavior sequence by combining the actual communication content sent;
[0214] The initial operation process parameters are loaded from the predefined task template library, and the operation process parameters are dynamically updated based on the optimization instructions of the planning module.
[0215] The environmental information is stored in the semantic memory partition of the multi-level memory module, the current behavior sequence is recorded in the plot memory partition of the multi-level memory module, and the operation process parameters are stored in the program memory partition of the multi-level memory module.
[0216] A dynamic index relationship is established between the semantic memory partition, the plot memory partition, and the program memory partition based on the current task context.
[0217] In one embodiment, the memory retrieval and communication construction module 20 is specifically used for:
[0218] Select the target retrieval partition of the multi-level memory module based on the current task requirements;
[0219] Extract the environmental state description and task target description from the semantic memory partition in the target retrieval partition;
[0220] Retrieve historical communication records related to the current task from the plot memory section of the target retrieval section;
[0221] Extract key object attributes from the environmental state description;
[0222] Based on the task objective description and historical communication records, the task objective is associated with the collaboration intent to obtain the associated collaboration intent.
[0223] The historical communication records are reassembled in chronological order to form a time series;
[0224] Integrate the key object attributes, associated collaboration intentions, and time sequences to generate a structured template;
[0225] The structured template is processed by a language model to generate candidate communication content, and the candidate communication content is temporarily stored as an internal deduction result.
[0226] In one embodiment, the target analysis and action extraction module 30 is specifically used for:
[0227] Analyze the current task objective and break it down into core task elements and execution constraints;
[0228] Retrieve the set of operation instructions that match the core task elements from the program memory partition of the multi-level memory module;
[0229] The executableness of each instruction in the set of operation instructions is verified based on the environmental state description, and the initial high-level action is filtered out through the executableness verification.
[0230] The initial high-level actions are sorted according to task priority, and the final high-level actions that meet the execution constraints are selected.
[0231] The final high-level actions are organized in a logical execution order to form a list of action plans to be selected.
[0232] In one embodiment, the reasoning analysis and communication decision module 40 is specifically used for:
[0233] Construct a structured multi-select prompt that includes the list of candidate action plans and candidate communication content;
[0234] The structured multiple-choice prompts are subjected to multi-step thought chain reasoning using a language model to generate correlation analysis results between each action plan and the communication content;
[0235] Based on the correlation analysis results, the task benefits and communication costs of sending candidate communication content are dynamically quantified;
[0236] When the task revenue exceeds a preset revenue threshold and the communication cost is lower than a preset cost threshold, it is determined that the communication sending condition is met.
[0237] A decision result is generated on whether to send the candidate communication content, and an action plan optimization suggestion is generated based on the correlation analysis result.
[0238] In one embodiment, the communication action generation module 50 is specifically used for:
[0239] When it is determined that the candidate communication content will be sent, the receiver identifier and core semantic payload are extracted from the candidate communication content.
[0240] Construct a sending action entity that includes the receiver identifier and the core semantic payload;
[0241] The system detects whether the time window of the sending action entity overlaps with the time window of the existing actions in the set of actions to be executed, whether there is a semantic conflict, and whether the resource consumption is lower than the communication bandwidth threshold configured by the system.
[0242] When there is no overlap in time windows, no semantic conflict, and resource consumption is lower than the communication bandwidth threshold configured by the system, the sending action entity is dynamically added to the set of actions to be executed.
[0243] In one embodiment, the action orchestration and interaction execution module 60 is specifically used for:
[0244] Integrate the planned actions in the list of candidate action plans with the communication actions in the set of actions to be executed to form a preliminary action plan sequence;
[0245] The preliminary action plan sequence is subjected to resource conflict detection and logical conflict resolution to generate an optimized conflict-free action plan;
[0246] Based on task priority and action dependency, the conflict-free action plan is time-series arranged to generate the final action plan;
[0247] Extract each high-level action from the final action plan and retrieve the corresponding operation flow from the program memory partition of the multi-level memory module;
[0248] According to the operation process, each high-level action is decomposed into basic operation instructions, and the control parameters corresponding to the basic operation instructions stored in the program memory partition are loaded.
[0249] In the environment, basic operation instructions with control parameters are executed in the time sequence of the final action plan to complete navigation, grasping, or communication interaction operations.
[0250] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multi-level memory collaborative decision-making method on the server side.
[0251] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides decision-making and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the user-side functions or steps of a multi-level memory collaborative decision-making method.
[0252] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0253] Acquire raw environmental observation data and update the multi-level memory module used to store environmental information, interaction history and operation instructions based on the raw observation data;
[0254] Retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information;
[0255] Based on the current task objective, executable high-level actions are extracted from the operation instructions of the updated multi-level memory module to form a list of candidate action plans;
[0256] The language model-based thought chain reasoning analysis is used to determine whether to send the candidate communication content, based on the list of alternative action plans and the candidate communication content.
[0257] When it is determined that the candidate communication content will be sent, the candidate communication content will be included in the set of actions to be executed.
[0258] A final action plan is generated based on the list of candidate action plans and the set of actions to be executed, and the final action plan is decomposed into at least one basic operation instruction, which is then executed to interact with the environment.
[0259] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0260] Acquire raw environmental observation data and update the multi-level memory module used to store environmental information, interaction history and operation instructions based on the raw observation data;
[0261] Retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information;
[0262] Based on the current task objective, executable high-level actions are extracted from the operation instructions of the updated multi-level memory module to form a list of candidate action plans;
[0263] The language model-based thought chain reasoning analysis is used to determine whether to send the candidate communication content, based on the list of alternative action plans and the candidate communication content.
[0264] When it is determined that the candidate communication content will be sent, the candidate communication content will be included in the set of actions to be executed.
[0265] A final action plan is generated based on the list of candidate action plans and the set of actions to be executed, and the final action plan is decomposed into at least one basic operation instruction, which is then executed to interact with the environment.
[0266] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0267] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0268] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0269] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A multi-level memory-based collaborative decision-making method, characterized in that, Includes the following steps: Acquire raw environmental observation data and update the multi-level memory module used to store environmental information, interaction history and operation instructions based on the raw observation data; Retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information; Based on the current task objective, executable high-level actions are extracted from the operation instructions of the updated multi-level memory module to form a list of candidate action plans; The language model-based thought chain reasoning analysis is used to determine whether to send the candidate communication content, based on the list of alternative action plans and the candidate communication content. When it is determined that the candidate communication content will be sent, the candidate communication content will be included in the set of actions to be executed. A final action plan is generated based on the list of candidate action plans and the set of actions to be executed, and the final action plan is decomposed into at least one basic operation instruction, which is then executed to interact with the environment.
2. The multi-level memory collaborative decision-making method as described in claim 1, characterized in that, Acquire raw environmental observation data and update the multi-level memory module used to store environmental information, interaction history, and operation instructions based on the raw observation data, including: Environmental image data is captured using a visual sensor as raw observation data; The environmental image data is subjected to object recognition to determine object category information, and spatial positioning is performed to determine object location information; A local semantic map is constructed based on the object category information and object location information, and the local semantic map is used as environmental information. Record the basic operation command types and execution results executed by the intelligent agent, and generate the current behavior sequence by combining the actual communication content sent; The initial operation process parameters are loaded from the predefined task template library, and the operation process parameters are dynamically updated based on the optimization instructions of the planning module. The environmental information is stored in the semantic memory partition of the multi-level memory module, the current behavior sequence is recorded in the plot memory partition of the multi-level memory module, and the operation process parameters are stored in the program memory partition of the multi-level memory module. A dynamic index relationship is established between the semantic memory partition, the plot memory partition, and the program memory partition based on the current task context.
3. The multi-level memory collaborative decision-making method as described in claim 1, characterized in that, Information is retrieved from the updated multi-level memory module, and candidate communication content is generated based on the retrieved information, including: Select the target retrieval partition of the multi-level memory module based on the current task requirements; Extract the environmental state description and task target description from the semantic memory partition in the target retrieval partition; Retrieve historical communication records related to the current task from the plot memory section of the target retrieval section; Extract key object attributes from the environmental state description; Based on the task objective description and historical communication records, the task objective is associated with the collaboration intent to obtain the associated collaboration intent. The historical communication records are reassembled in chronological order to form a time series; Integrate the key object attributes, associated collaboration intentions, and time sequences to generate a structured template; The structured template is processed by a language model to generate candidate communication content, and the candidate communication content is temporarily stored as an internal deduction result.
4. The multi-level memory collaborative decision-making method as described in claim 1, characterized in that, Based on the current task objective, executable high-level actions are extracted from the updated multi-level memory module's operation instructions to form a list of candidate action plans, including: Analyze the current task objective and break it down into core task elements and execution constraints; Retrieve the set of operation instructions that match the core task elements from the program memory partition of the multi-level memory module; The executableness of each instruction in the set of operation instructions is verified based on the environmental state description, and the initial high-level action is filtered out through the executableness verification. The initial high-level actions are sorted according to task priority, and the final high-level actions that meet the execution constraints are selected. The final high-level actions are organized in a logical execution order to form a list of action plans to be selected.
5. The multi-level memory collaborative decision-making method as described in claim 1, characterized in that, The language model-based thought chain reasoning analysis, which examines the list of candidate action plans and the candidate communication content, determines whether to send the candidate communication content, including: Construct a structured multi-select prompt that includes the list of candidate action plans and candidate communication content; The structured multiple-choice prompts are subjected to multi-step thought chain reasoning using a language model to generate correlation analysis results between each action plan and the communication content; Based on the correlation analysis results, the task benefits and communication costs of sending candidate communication content are dynamically quantified; When the task revenue exceeds a preset revenue threshold and the communication cost is lower than a preset cost threshold, it is determined that the communication sending condition is met. A decision result is generated on whether to send the candidate communication content, and an action plan optimization suggestion is generated based on the correlation analysis result.
6. The multi-level memory collaborative decision-making method as described in claim 1, characterized in that, When it is determined that the candidate communication content will be sent, the candidate communication content will be included in the set of actions to be executed, including: When it is determined that the candidate communication content will be sent, the receiver identifier and core semantic payload are extracted from the candidate communication content. Construct a sending action entity that includes the receiver identifier and the core semantic payload; The system detects whether the time window of the sending action entity overlaps with the time window of the existing actions in the set of actions to be executed, whether there is a semantic conflict, and whether the resource consumption is lower than the communication bandwidth threshold configured by the system. When there is no overlap in time windows, no semantic conflict, and resource consumption is lower than the communication bandwidth threshold configured by the system, the sending action entity is dynamically added to the set of actions to be executed.
7. The multi-level memory collaborative decision-making method as described in claim 1, characterized in that, A final action plan is generated based on the list of candidate action plans and the set of actions to be executed, and the final action plan is decomposed into at least one basic operation instruction. The basic operation instruction is executed to interact with the environment, including: Integrate the planned actions in the list of candidate action plans with the communication actions in the set of actions to be executed to form a preliminary action plan sequence; The preliminary action plan sequence is subjected to resource conflict detection and logical conflict resolution to generate an optimized conflict-free action plan; Based on task priority and action dependency, the conflict-free action plan is time-series arranged to generate the final action plan; Extract each high-level action from the final action plan and retrieve the corresponding operation flow from the program memory partition of the multi-level memory module; According to the operation process, each high-level action is decomposed into basic operation instructions, and the control parameters corresponding to the basic operation instructions stored in the program memory partition are loaded. In the environment, basic operation instructions with control parameters are executed in the time sequence of the final action plan to complete navigation, grasping, or communication interaction operations.
8. A multi-level memory collaborative decision-making device, characterized in that, The multi-level memory collaborative decision-making device includes: The environmental perception and memory update module is used to acquire raw environmental observation data and update the multi-level memory module for storing environmental information, interaction history and operation instructions based on the raw observation data. The memory retrieval and communication construction module is used to retrieve information from the updated multi-level memory module and generate candidate communication content based on the retrieved information. The target analysis and action extraction module is used to extract executable high-level actions from the operation instructions of the updated multi-level memory module based on the current task target, and form a list of candidate action plans; The reasoning analysis and communication decision module is used to perform reasoning analysis based on the language model's thought chain to determine whether to send the candidate communication content; The communication action generation module is used to include the candidate communication content to be sent into the set of actions to be executed when it is determined that the candidate communication content should be sent. The action orchestration and interaction execution module is used to generate a final action plan based on the list of candidate action plans and the set of actions to be executed, and to decompose the final action plan into at least one basic operation instruction, and to execute the basic operation instruction to interact with the environment.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multi-level memory collaborative decision-making program stored in the memory and executable on the processor, wherein the multi-level memory collaborative decision-making program, when executed by the processor, implements the steps of the multi-level memory collaborative decision-making method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a multi-level memory collaborative decision-making program, which, when executed by a processor, implements the steps of the multi-level memory collaborative decision-making method as described in any one of claims 1-7.
Citation Information
Cited By
Spatial intelligent multi-modal tree planning method and system based on value-driven cutting
CN121477901A
Structured form processing method and system based on intelligent identification and storage medium
CN121503450A
Dynamic compression and efficient recall collaborative end-side robot memory system management method
CN121552393A
A Management Method for End-Side Robot Memory Systems that Combines Dynamic Compression and Efficient Recall
CN121552393B
Memory data processing method and device, storage medium and computer program product
CN121562658A