Multi-agent-based perceptual decision action link construction method
By building a perceptual decision-making action link in a multi-agent system, using agent collaboration and deep reinforcement learning algorithms, the problem of inefficiency in multi-agent systems in complex task environments is solved, and efficient and accurate task execution and emergency responses are achieved.
Patent Information
- Application Number
- CN202510452113.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When facing complex task environments, existing multi-agent systems are difficult to adjust tasks flexibly, resulting in low action efficiency and insufficient coordination of agents.
A method of constructing a perceptual decision-making action link based on multi-agents is proposed. Through the initialization of multi-agent systems, environmental modeling and data collection, target recognition and strategy generation, task execution and status update, effect evaluation and strategy adjustment, collaboration and task optimization between agents are realized.
Through collaboration between agents and optimization of deep reinforcement learning algorithms, dynamic task allocation and execution are achieved, the accuracy and efficiency of task completion are improved, and the ability to respond to emergencies is enhanced.
Smart Images

Figure CN120031341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for constructing a perception, decision-making, and action link based on multiple agents. Background Art
[0002] Traditional task allocation and execution usually rely on the commander's human judgment and teamwork. Although this method can meet the task requirements to a certain extent, as the task requirements become increasingly complex, the traditional method becomes increasingly insufficient. In the existing technology, multi-agent systems have gradually become an important means to solve this problem. Its advantage is that it can achieve rapid response and efficient execution of complex tasks through effective collaboration between agents. However, many existing technologies mainly focus on the execution of single-target tasks and lack a systematic collaborative task allocation mechanism. This makes it difficult for agents to flexibly adjust their respective tasks in the face of a rapidly changing task environment, thereby reducing the overall efficiency of action. Summary of the invention
[0003] In order to solve the above problems in the prior art, the present invention proposes a perception-decision-action chain construction method based on multi-agents to solve the problems existing in the prior art such as low efficiency of perception-decision-action chain construction and insufficient coordination of intelligent agents.
[0004] The present invention provides a method for constructing a multi-agent perception-decision-action link, comprising the following steps: (1) Initialize the multi-agent system, build an agent set, and subdivide each agent into reconnaissance agent, decision agent, action agent, and evaluation agent according to its function to complete the system initialization; (2) Environmental modeling and data collection: collect target information in real time through multimodal sensing devices, build a global environmental state set and update it to the system synchronously; (3) Target identification and strategy generation: The decision-making agent calculates the target value function based on the critic network, identifies high-priority targets and generates the optimal strategy based on resource constraints; (4) Task execution and status update: the agent executes the strategy, provides real-time feedback on the execution results, and updates the environment status; (5) Effect evaluation and strategy adjustment: The evaluation agent analyzes target damage and resource consumption and feeds the evaluation results back to the decision-making agent. The experience pool stores status, action, and reward data to dynamically optimize the strategy.
[0005] Furthermore, step (1) includes the following steps: (11) Collect the agents The roles are divided into four categories: reconnaissance, decision-making, action and assessment, with clear division of tasks; (12) Randomly initialize the Actor network parameters of each agent and Critic network parameters , ensuring exploratory nature in the initial stages.
[0006] Furthermore, step (2) includes the following steps: (21) Obtain the target’s location and motion status data in real time through drones and radar sensors; (22) Use multimodal sensing technology to fuse data and generate the global environment state S t ; (23)Continuously update the environment status.
[0007] Furthermore, step (3) includes the following steps: (31) Use the Critic network to evaluate the target value Q(s, T) and select high-priority targets ; (32) Generate the optimal strategy based on resource constraints .
[0008] Furthermore, step (4) includes the following steps: (41) The agent acts according to the strategy Execute tasks and choose appropriate actions; (42) Update the environment state S t , record target status and resource consumption data.
[0009] Furthermore, step (5) includes the following steps: (51) The evaluation agent calculates the target damage amount T destroyed and resource consumption C used ; (52) Feedback the evaluation results to the decision-making agent to adjust subsequent strategies; (53) Change state S t 、Action a t and Rewards Stored in the shared experience pool R for reinforcement learning optimization.
[0010] The present invention provides a multi-agent based perception, decision-making, action link construction system, comprising: Initialize the multi-agent system: used to build an agent set, subdivide each agent into reconnaissance agent, decision agent, action agent and evaluation agent according to its function, and complete the system initialization; Environmental modeling and data acquisition module: used to collect target information in real time through multimodal perception devices, build a global environmental state set, and update it to the system synchronously; Target identification and strategy generation module: used by decision-making agents to calculate target value functions based on the Critic network, identify high-priority targets, and generate optimal strategies based on resource constraints; Task execution and status update module: used for action agents to execute strategies, provide real-time feedback on execution results and update environmental status; Effect evaluation and strategy adjustment module: used to evaluate the agent's analysis of target damage and resource consumption, feed back the evaluation results to the decision-making agent, store the state, action and reward data through the experience pool, and dynamically optimize the strategy.
[0011] An electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the methods for constructing a multi-agent perception, decision-making, and action link.
[0012] A storage medium described in the present invention stores a computer program, and when the computer program is executed by a processor, it implements a method for constructing a perception, decision-making and action link based on multiple agents according to any one of the above.
[0013] Beneficial effects: Compared with the prior art, the present invention has the following advantages: by utilizing the collaborative ability between intelligent agents and deep reinforcement learning algorithms, the whole process of tasks from reconnaissance, decision-making to action is optimized; intelligent agents can share information and make collaborative decisions in real time in complex task environments, thereby realizing dynamic task allocation and execution; in this way, the system not only improves the accuracy and efficiency of task completion, but also enhances the ability to deal with emergencies. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the intelligent agent training process of the present invention. DETAILED DESCRIPTION
[0015] The present invention will be further explained and illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only used to illustrate and explain the present invention, and do not impose any limitation on the scope of implementation of the present invention.
[0016] like Figure 1 As shown, the embodiment of the present invention provides a method for constructing a perception-decision-action link based on multiple agents, comprising the following steps: Initialize the multi-agent system and initialize the agent set , where each agent is divided into different task roles according to the actual environment perception decision action chain process, including reconnaissance agent, decision agent, action agent and evaluation agent. The initialization of the agent system lays the foundation for subsequent task allocation and collaborative operation. It includes the following steps: Step 1-1: System initialization agent set , the intelligent agent is subdivided into four roles: reconnaissance, decision-making, action, and evaluation, in order to clarify their respective task division.
[0017] Step 1-2: Randomly initialize the Actor network of each agent and Critic Network parameters to ensure that the agent is exploratory in the initial stage.
[0018] Step 2: Task environment modeling and data collection. Modeling task environment as a state set , where each state includes the current environment's task environment, target information, and the agent's current position. The reconnaissance agent detects targets through its multimodal sensing equipment (such as drones, infrared sensors, and radar systems) and obtains the target's location information, type, and trajectory in real time. All reconnaissance data will be synchronized to the global environment state for subsequent task identification and decision-making. The following steps are included: Step 2-1: The reconnaissance agent obtains target information on the battlefield in real time through sensors such as drones and radars, including its location information, motion status, and other relevant data.
[0019] Step 2-2: Use multimodal sensing technology to fuse and analyze the collected data to generate the global mission environment state .
[0020] Step 2-3: The reconnaissance agent continuously updates the environmental status to ensure that the decision-making agent can make accurate strategic judgments based on the latest enemy information.
[0021] Step 3: Target identification and strategy generation: After the decision agent receives the target data from the reconnaissance agent, it uses the Critic network of reinforcement learning to calculate the target value function. , identify high priority targets Through this process, the command intelligent agent can dynamically adjust the action strategy in the task environment and generate the optimal strategy. , ensuring that resource consumption is minimized. The strategy generation process also combines the constraints of actual environmental resources, such as the number of units, to ensure the efficiency and accuracy of task execution. It includes the following steps: Step 3-1: After the decision agent receives the target data provided by the reconnaissance agent, it uses the critic network Evaluate all goals and identify high priority goals .
[0022] Step 3-2: By maximizing the target value, , formulate corresponding action tasks. Consider the task environment and available resources to generate the optimal strategy , ensuring the effectiveness and efficiency of task execution.
[0023] Step 4: Task execution and status update. After receiving the task instructions from the decision-making agent, the action agent performs the target action according to the pre-generated optimal strategy, executes the action and provides real-time feedback on the effect. After each action is completed, the task environment status It will be updated according to the results of the task execution, including the degree of damage to the target, the task completion rate and resource consumption. It includes the following steps: Step 4-1: After receiving the task assigned by the decision-making agent Afterwards, the action agent selects the appropriate action mode and executes the action according to the target characteristics and the task environment status.
[0024] Step 4-2: After the action is completed, the action agent needs to update the task environment state , record the execution effect and related data. Collect and feedback the action effect data, including the target destruction and its own resource consumption, for subsequent effect evaluation and strategy adjustment.
[0025] Step 5: Effect evaluation and strategy adjustment: The evaluation agent cooperates with the decision agent to conduct a comprehensive evaluation of the effect of each action task. The evaluation indicators include the number of targets destroyed. , Resource consumption and its deviation from the expected results. Evaluate the results based on the effects , the system dynamically adjusts the subsequent strategies. All states, actions and reward data during the execution process will be stored in the experience pool R for the strategy optimization of the reinforcement learning algorithm. Through continuous experience accumulation and strategy adjustment, the global goal is finally optimized. It includes the following steps: Step 5-1: The evaluation agent evaluates the action effect of the action agent and analyzes the target destruction situation and resource consumption .
[0026] Step 5-2: Feedback the evaluation results to the decision-making agent for subsequent strategy adjustment and optimization.
[0027] Step 5-3: After completing the task, all agents will ,action ,award The information is stored in the shared experience pool R to achieve strategy optimization based on historical experience.
[0028] The embodiment of the present invention also provides a multi-agent based perception, decision-making, action link construction system, including: Initialize the multi-agent system: used to build an agent set, subdivide each agent into reconnaissance agent, decision agent, action agent and evaluation agent according to its function, and complete the system initialization; Environmental modeling and data acquisition module: used to collect target information in real time through multimodal perception devices, build a global environmental state set, and update it to the system synchronously; Target identification and strategy generation module: used by decision-making agents to calculate target value functions based on the Critic network, identify high-priority targets, and generate optimal strategies based on resource constraints; Task execution and status update module: used for action agents to execute strategies, provide real-time feedback on execution results and update environmental status; Effect evaluation and strategy adjustment module: used to evaluate the agent's analysis of target damage and resource consumption, feed back the evaluation results to the decision-making agent, store the state, action and reward data through the experience pool, and dynamically optimize the strategy.
[0029] An embodiment of the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the computer program implements any one of the methods for constructing a multi-agent perception, decision-making and action link.
[0030] An embodiment of the present invention also provides a storage medium storing a computer program, which, when executed by a processor, implements a method for constructing a perception, decision-making, and action link based on multiple agents according to any one of the items.
Claims
1. A method for constructing a perception-decision-action link based on multi-agents, characterized in that: The following steps are involved: (1) Initialize the multi-agent system, build an agent set, and subdivide each agent into reconnaissance agent, decision agent, action agent, and evaluation agent according to its function to complete the system initialization; (2) Environmental modeling and data collection: collect target information in real time through multimodal sensing devices, build a global environmental state set and update it synchronously to the system; (3) Target identification and strategy generation: The decision-making agent calculates the target value function based on the critic network, identifies high-priority targets and generates the optimal strategy based on resource constraints; (4) Task execution and status update: the agent executes the strategy, provides real-time feedback on the execution results, and updates the environment status; (5) Effect evaluation and strategy adjustment: The evaluation agent analyzes target damage and resource consumption and feeds the evaluation results back to the decision-making agent. The experience pool stores status, action, and reward data to dynamically optimize the strategy.
2. According to claim 1, a method for constructing a multi-agent perception-decision-action link, characterized in that: Step (2) includes the following steps: (21) Obtain the target’s location and motion status data in real time through drones and radar sensors; (22) Use multimodal sensing technology to fuse data and generate the global environment state S t ; (23)Continuously update the environment status.
3. According to claim 1, a method for constructing a multi-agent perception-decision-action link, characterized in that: Step (3) includes the following steps: (31) Use the Critic network to evaluate the target value Q(s, T) and select high-priority targets ; (32) Generate the optimal strategy based on resource constraints .
4. According to claim 1, a method for constructing a perception, decision-making and action link based on multiple agents is characterized in that: Step (4) includes the following steps: (41) The agent acts according to the strategy Execute tasks and choose appropriate actions; (42) Update the environment state S t , record target status and resource consumption data.
5. According to claim 1, a method for constructing a multi-agent perception-decision-action link, characterized in that: Step (5) includes the following steps: (51) The evaluation agent calculates the target damage amount T destroyed and resource consumption C used ; (52) Feedback the evaluation results to the decision-making agent to adjust subsequent strategies; (53) Change state S t 、Action a t and Rewards Stored in the shared experience pool R for reinforcement learning optimization.
6. A multi-agent based perception, decision-making, action link construction system, characterized by: include: Initialize the multi-agent system: used to build an agent set, subdivide each agent into reconnaissance agent, decision agent, action agent and evaluation agent according to its function, and complete the system initialization; Environmental modeling and data acquisition module: used to collect target information in real time through multimodal perception devices, build a global environmental state set, and update it to the system synchronously; Target identification and strategy generation module: used by decision-making agents to calculate target value functions based on the Critic network, identify high-priority targets, and generate optimal strategies based on resource constraints; Task execution and status update module: used for action agents to execute strategies, provide real-time feedback on execution results and update environmental status; Effect evaluation and strategy adjustment module: used to evaluate the agent's analysis of target damage and resource consumption, feed back the evaluation results to the decision-making agent, store the state, action and reward data through the experience pool, and dynamically optimize the strategy.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into the processor, it implements a method for constructing a multi-agent perception, decision-making and action link according to any one of claims 1 to 5.
8. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements a method for constructing a perception, decision-making, and action link based on multiple agents according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-machine collaborative air combat planning method and system based on deep reinforcement learning
CN112861442A
Collaborative operation method combining unmanned aerial vehicle and unmanned vehicle
CN114355900A
DDDPG-based path planning method fused with motion attitude of unmanned aerial vehicle
CN117519231A
Intelligent target allocation method and system based on deep reinforcement learning
CN119130015A
Cited By
Multi-agent collaborative decision-making method and device
CN121635469A
Method and device for collaboratively reconstructing OODA decision closed loop based on multiple agents
CN122226380A