Underwater moving target tracking system based on multi-agent Bayesian reinforcement learning
Through the multi-agent Bayesian reinforcement learning system, the sparse reward and credit allocation problems of the multi-agent underwater moving target tracking system in complex environments are solved, efficient target tracking and collaborative decision-making are achieved, and the autonomy and accuracy of the system are improved.
Patent Information
- Application Number
- CN202510986575.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing multi-agent underwater moving target tracking systems suffer from sparse reward problems, difficult credit allocation and repeated learning in complex and uncertain underwater environments, which lead to long training time and slow convergence, affecting the collaborative effect.
An underwater moving target tracking system based on multi-agent Bayesian reinforcement learning is adopted. By establishing a collaborative network, introducing Bayesian reasoning methods and reinforcement learning algorithms, path planning and task allocation are optimized, and real-time information sharing and decision-making are achieved by combining wireless communication, thereby optimizing the collaborative strategy among multiple agents.
It significantly improves the accuracy, stability and autonomy of target tracking, can effectively respond to dynamic environmental changes and external interference, and improves execution efficiency and adaptability.
Smart Images

Figure CN120468857B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater moving target tracking, and in particular to an underwater moving target tracking system based on multi-agent Bayesian reinforcement learning. Background Art
[0002] Underwater moving target tracking technology has important applications in fields such as marine resource exploration and ecological monitoring. Its primary task is to accurately track and locate dynamic underwater targets through autonomous perception and decision-making. Multi-agent collaborative tracking tasks require not only accurate target positioning in underwater environments but also strong autonomous decision-making and collaborative capabilities. Each agent must not only accurately track targets based on its own perception information but also be able to share information and coordinate actions with other agents in complex underwater environments. Currently, multi-agent reinforcement learning approaches, both domestically and internationally, face several challenges, including sparse rewards, difficulty in credit allocation, and repeated learning. These issues hinder effective learning, extend training time, slow convergence, and compromise overall collaborative effectiveness. Consequently, traditional target tracking methods face challenges in complex and uncertain underwater environments, including the uncertainty of underwater sensors, dynamic changes in target motion, and external interference. Summary of the Invention
[0003] In view of the fact that the current intelligent underwater moving target tracking system cannot make efficient autonomous decisions and accurately track in dynamic and changeable underwater environments, the present invention provides an underwater moving target tracking system based on multi-agent Bayesian reinforcement learning.
[0004] An underwater moving target tracking system based on multi-agent Bayesian reinforcement learning includes the following steps: S1, establishing a target tracking architecture based on multi-agents, so that multiple agents form a collaborative network and share the collected data; S2, establishing a coordinate system with the center of the agent as the origin, and deriving the kinematic equation of the agent in the coordinate system based on Newton's second law and Euler's equation; S3, introducing the Bayesian reasoning method to establish a dynamic estimation model of the target state; S4, developing a reinforcement learning algorithm based on Bayesian reasoning to optimize the path planning among multiple agents; S5, in a multi-dynamic target scenario, determining the target to be tracked through a dynamic task scheduling mechanism and assigning the task to the appropriate agent; S6, optimizing the decision-making of multiple agents in a complex environment; S7, constructing and simulating the collaborative capture task of multiple agents in a complex underwater environment in a simulation platform.
[0005] Furthermore, in S1, multiple agents form a collaborative network via wireless communication modules. Each agent possesses independent perception, decision-making, and execution capabilities, and can share information with other agents in real time. Multiple agents are divided into master agents and auxiliary agents. The master agent receives and integrates information from all other agents, performing data fusion and global optimization. The auxiliary agents dynamically adjust their actions based on the local environment.
[0006] Furthermore, each agent is equipped with a perception module, which is used to sense the underwater environment and the state of the target, such as the target's position, speed, and heading. This perception module includes a sonar sensor, which is used to detect the relative positions of underwater obstacles, other agents, and the target object. Each agent is also equipped with a data preprocessing module, which preprocesses the data collected by the perception module to extract key information and reduce the impact of noise interference.
[0007] Furthermore, in S2, a coordinate system with the center of the agent as its origin contains three mutually perpendicular axes: the x-axis, the y-axis, and the z-axis. The agent's motion is then decomposed into translational motion along the x-, y-, and z-axes and rotational motion around them. Based on the principles of kinematics, the relationships between the agent's linear velocity and angular velocity on the x-, y-, and z-axes and its position, velocity, and acceleration are established. Based on Newton's second law and Euler's equations, the six-degree-of-freedom equations of motion for the agent in the body's coordinate system are derived, encompassing linear motion in three directions and angular motion in three directions. These equations of motion describe the forces and torques acting on the agent in the underwater environment.
[0008] Center of gravity of the agent The position vector expression is:
[0009] ;
[0010] Where, Represented as the center of gravity of the agent frame along the x-axis; Represented as the center of gravity of the agent frame along the y-axis; It is represented as the center of gravity of the agent frame along the z-axis;
[0011] The center of gravity of the agent The position vector expression is:
[0012] ;
[0013] Where, It is represented as the center of buoyancy of the agent frame along the x-axis; It is represented as the buoyancy center of the agent frame along the y-axis direction; It is represented as the center of buoyancy of the agent frame along the z-axis.
[0014] Furthermore, in S3, each agent combines the relative distance and speed of the target obtained by the sonar sensor with the target information shared by other agents, and uses a dynamic estimation model based on Bayesian reasoning to estimate the target state. By fusing prior knowledge and perception data, the target's motion trajectory and behavior pattern are updated in real time.
[0015] Dynamic estimation model expression based on Bayesian inference method:
[0016] ;
[0017] Where, Represented as a global state, that is, the motion state of the target, such as position and velocity; Expressed as Observations obtained by sensors; Represented as a prior distribution of the target state, reflecting the target's initial assumptions or historical behavior; For the The likelihood function of the agent's perception data is expressed in the known target state. When, Agents acquire perception data probability; Represented as the perception data of all given agents After that, the target state The posterior distribution of .
[0018] Furthermore, in S4, a reinforcement learning algorithm based on Bayesian reasoning is developed, and a learning framework including state space, action space and reward function is designed. By combining multi-agent collaboration with the Bayesian reinforcement learning algorithm, the path planning and task allocation between agents are optimized to ensure that there is no conflict between the tasks and paths of the agents. Next, the agent selects an action To maximize the overall reward , which is rewarded by tracking performance , energy efficiency and synergy The three parts are weighted to obtain the path optimization expression of the reinforcement learning algorithm based on Bayesian reasoning:
[0019] ;
[0020] Where, Represented as the agent at time Action choices; Indicates that the status Next, the agent selects an action To maximize the overall reward; It is expressed as a reward for target tracking accuracy, which is related to the target distance or tracking error; It is expressed as an energy consumption optimization reward, which is related to the control of speed and acceleration; It is expressed as a reward for the collaborative efficiency between agents, which is related to the distance and degree of cooperation between agents.
[0021] Furthermore, in S5, in a multi-target dynamic scenario, the dynamic task scheduling mechanism determines the target to be tracked by calculating the priority of the target; when multiple targets enter the sensing range, the dynamic task scheduling mechanism calculates the priority based on the distance, speed and importance factors of the target;
[0022] The calculation expression of target priority is:
[0023] ;
[0024] Where, Represented as a target Priority calculation; Expressed as priority weight; Expressed as velocity weight; Expressed as target weight; Represents the target The Euclidean distance to the agent, in m. The closer the distance, the higher the priority weight. The greater the contribution; Expressed as target speed in m / s. The faster the speed, the higher the tracking urgency. The speed weight Corresponding dynamic adjustment capabilities; Indicates the preset importance level for the target, , target weight response task preference;
[0025] Each weight coefficient must satisfy the normalization constraint:
[0026] ;
[0027] The normalization constraints are satisfied by each weight coefficient to ensure that the priority score is a standardized value.
[0028] Furthermore, after determining the target to be tracked, the system continuously monitors the target and agent status, adjusts task allocation in real time, and assigns the target task to the appropriate agent. When the target position changes or the agent's task load is too high, the system dynamically reallocates tasks. During the entire dynamic task scheduling process, indicators such as task completion rate, energy efficiency, and path conflicts are evaluated in real time, and the scheduling strategy is further optimized based on the feedback.
[0029] Each agent combines path optimization based on Bayesian reasoning reinforcement learning algorithm to develop the trajectory with the lowest energy consumption and the best time. The objective function of path planning is defined as a multi-objective optimization problem to determine the total cost of the appropriate agent. Calculation expression:
[0030] ;
[0031] Where, Expressed as the weight coefficient of time, is expressed as the weight coefficient of energy consumption, and and Need to meet ; Expressed as total task time; Expressed as starting point coordinates; Expressed as the end point coordinates; Expressed as the maximum speed of the agent; Represented as control input, such as acceleration; Expressed as the total cost when the control input is minimized.
[0032] Furthermore, in S6, multi-agents combine target tracking accuracy rewards, collaborative efficiency improvements, and path conflict penalties to achieve real-time collaboration and decision-making in complex environments. The expression for optimizing decision-making is:
[0033] ;
[0034] Where, is the global optimal action combination, which represents the optimal decision of all agents in the current state; Represented as agent state; represents the actions performed by the agent; Represented as for all agents The rewards and penalties are summed up to optimize the overall system performance; Represented as an agent In state Next action The state-action value of is used to reflect the long-term expected return of the current strategy; Denoted as target tracking reward; Expressed as collaborative rewards between agents, it is used to encourage data sharing and task collaboration; Expressed as a penalty for path conflict, it is used to avoid collisions or repeated tasks between agents;
[0035] Among them, target tracking reward , the calculation expression based on the deviation between the target position and the current path is:
[0036] ;
[0037] Where, Represented as the real-time position of the agent; Represented as target position; It is expressed as a measurement tool to quantify the tracking effect;
[0038] Among them, the penalty for path conflict , the calculation expression used to avoid collisions or repeated tasks between agents is:
[0039] ;
[0040] Where, Indicates the strictness of conflict avoidance for control; Expressed as the potential risk of a quantified action.
[0041] Furthermore, in S7, a virtual scene supporting underwater environment simulation was built in the simulation platform, dynamic targets were loaded, and a dynamic scene was constructed in which three friendly intelligent agents cooperated to surround and capture an enemy intelligent agent, simulating the collaborative capture task of multiple intelligent agents in a complex underwater environment, and then further conducting actual sea trials to verify the collaborative capture of multiple intelligent agents.
[0042] The beneficial effects of the present invention are as follows: through the deep integration of multi-agent collaborative strategies and Bayesian reinforcement learning, the present invention can autonomously and in real time optimize tracking strategies in different dynamic environments and multi-target scenarios, significantly improving the accuracy, stability, and autonomy of target tracking. The present invention uses wireless communication technology to achieve real-time information sharing and collaborative decision-making among multiple agents, optimizes task allocation, improves execution efficiency, and enhances the estimation accuracy of target states through Bayesian reasoning. At the same time, the reinforcement learning algorithm dynamically adjusts the action strategy, which can effectively respond to the dynamic changes of the target, the uncertainty of the environment, and various disturbance factors, thereby enhancing the tracking efficiency and adaptability in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Shown is a schematic flow chart of an embodiment of the present invention.
[0044] Figure 2 Shown is a schematic diagram of the model of the UUV intelligent body in the present invention.
[0045] Figure 3 It is a schematic diagram of the encirclement and capture of the target UUV by UUV1, UUV2 and UUV3 in the present invention.
[0046] Figure 4 The figure shows the change of the heading angle of UUV1, UUV2 and UUV3 over time in the actual experiment of the present invention.
[0047] Figure 5 The figure shows the change of the relative distance and relative angle of UUV1, UUV2 and UUV3 over time in the actual experiment of the present invention.
[0048] Figure 6 The figure shows the trajectory of UUV1, UUV2 and UUV3 in the actual experiment of the present invention in cooperating to capture the target UUV. DETAILED DESCRIPTION
[0049] An underwater moving target tracking system based on multi-agent Bayesian reinforcement learning includes the following steps: S1, establishing a target tracking architecture based on multi-agents, so that multiple agents form a collaborative network and share the collected data; S2, establishing a coordinate system with the center of the agent as the origin, and deriving the kinematic equation of the agent in the coordinate system based on Newton's second law and Euler's equation; S3, introducing the Bayesian reasoning method to establish a dynamic estimation model of the target state; S4, developing a reinforcement learning algorithm based on Bayesian reasoning to optimize the path planning among multiple agents; S5, in a multi-dynamic target scenario, determining the target to be tracked through a dynamic task scheduling mechanism and assigning the task to the appropriate agent; S6, optimizing the decision-making of multiple agents in a complex environment; S7, constructing and simulating the collaborative capture task of multiple agents in a complex underwater environment in a simulation platform.
[0050] Furthermore, in S1, multiple agents form a collaborative network via wireless communication modules. Each agent possesses independent perception, decision-making, and execution capabilities, and can share information with other agents in real time. Multiple agents are divided into master agents and auxiliary agents. The master agent receives and integrates information from all other agents, performing data fusion and global optimization. The auxiliary agents dynamically adjust their actions based on the local environment.
[0051] Furthermore, each agent is equipped with a perception module, which is used to sense the underwater environment and the state of the target, such as the target's position, speed, and heading. This perception module includes a sonar sensor, which is used to detect the relative positions of underwater obstacles, other agents, and the target object. Each agent is also equipped with a data preprocessing module, which preprocesses the data collected by the perception module to extract key information and reduce the impact of noise interference.
[0052] Furthermore, in S2, a coordinate system with the center of the agent as its origin contains three mutually perpendicular axes: the x-axis, the y-axis, and the z-axis. The agent's motion is then decomposed into translational motion along the x-, y-, and z-axes and rotational motion around them. Based on the principles of kinematics, the relationships between the agent's linear velocity and angular velocity on the x-, y-, and z-axes and its position, velocity, and acceleration are established. Based on Newton's second law and Euler's equations, the six-degree-of-freedom equations of motion for the agent in the body's coordinate system are derived, encompassing linear motion in three directions and angular motion in three directions. These equations of motion describe the forces and torques acting on the agent in the underwater environment.
[0053] Center of gravity of the agent The position vector expression is:
[0054] ;
[0055] Where, Represented as the center of gravity of the agent frame along the x-axis; Represented as the center of gravity of the agent frame along the y-axis; It is represented as the center of gravity of the agent frame along the z-axis;
[0056] The center of gravity of the agent The position vector expression is:
[0057] ;
[0058] Where, It is represented as the center of buoyancy of the agent frame along the x-axis; It is represented as the buoyancy center of the agent frame along the y-axis direction; It is represented as the center of buoyancy of the agent frame along the z-axis.
[0059] Furthermore, in S3, each agent combines the relative distance and speed of the target obtained by the sonar sensor with the target information shared by other agents, and uses a dynamic estimation model based on Bayesian reasoning to estimate the target state. By fusing prior knowledge and perception data, the target's motion trajectory and behavior pattern are updated in real time.
[0060] Dynamic estimation model expression based on Bayesian inference method:
[0061] ;
[0062] Where, Represented as a global state, that is, the motion state of the target, such as position and velocity; Expressed as Observations obtained by sensors; Represented as a prior distribution of the target state, reflecting the target's initial assumptions or historical behavior; For the The likelihood function of the agent's perception data is expressed in the known target state. When, Agents acquire perception data probability; Represented as the perception data of all given agents After that, the target state The posterior distribution of .
[0063] Furthermore, in S4, a reinforcement learning algorithm based on Bayesian reasoning is developed, and a learning framework including state space, action space and reward function is designed. By combining multi-agent collaboration with the Bayesian reinforcement learning algorithm, the path planning and task allocation between agents are optimized to ensure that there is no conflict between the tasks and paths of the agents. Next, the agent selects an action To maximize the overall reward , which is rewarded by tracking performance , energy efficiency and synergy The three parts are weighted to obtain the path optimization expression of the reinforcement learning algorithm based on Bayesian reasoning:
[0064] ;
[0065] Where, Represented as the agent at time Action choices; Indicates that the status Next, the agent selects an action To maximize the overall reward; It is expressed as a reward for target tracking accuracy, which is related to the target distance or tracking error; It is expressed as an energy consumption optimization reward, which is related to the control of speed and acceleration; It is expressed as a reward for the collaborative efficiency between agents, which is related to the distance and degree of cooperation between agents.
[0066] Furthermore, in S5, in a multi-target dynamic scenario, the dynamic task scheduling mechanism determines the target to be tracked by calculating the priority of the target; when multiple targets enter the sensing range, the dynamic task scheduling mechanism calculates the priority based on the distance, speed and importance factors of the target;
[0067] The calculation expression of target priority is:
[0068] ;
[0069] Where, Represented as a target Priority calculation; Expressed as priority weight; Expressed as velocity weight; Expressed as target weight; Represents the target The Euclidean distance to the agent, in m. The closer the distance, the higher the priority weight. The greater the contribution; Expressed as target speed in m / s. The faster the speed, the higher the tracking urgency. The speed weight Corresponding dynamic adjustment capabilities; Indicates the preset importance level for the target, , target weight response task preference;
[0070] Each weight coefficient must satisfy the normalization constraint:
[0071] ;
[0072] The normalization constraints are satisfied by each weight coefficient to ensure that the priority score is a standardized value.
[0073] Furthermore, after determining the target to be tracked, the system continuously monitors the target and agent status, adjusts task allocation in real time, and assigns the target task to the appropriate agent. When the target position changes or the agent's task load is too high, the system dynamically reallocates tasks. During the entire dynamic task scheduling process, indicators such as task completion rate, energy efficiency, and path conflicts are evaluated in real time, and the scheduling strategy is further optimized based on the feedback.
[0074] Each agent combines path optimization based on Bayesian reasoning reinforcement learning algorithm to develop the trajectory with the lowest energy consumption and the best time. The objective function of path planning is defined as a multi-objective optimization problem to determine the total cost of the appropriate agent. Calculation expression:
[0075] ;
[0076] Where, Expressed as the weight coefficient of time, is expressed as the weight coefficient of energy consumption, and and Need to meet ; Expressed as total task time; Expressed as starting point coordinates; Expressed as the end point coordinates; Expressed as the maximum speed of the agent; Represented as control input, such as acceleration; Expressed as the total cost when the control input is minimized.
[0077] Furthermore, in S6, multi-agents combine target tracking accuracy rewards, collaborative efficiency improvements, and path conflict penalties to achieve real-time collaboration and decision-making in complex environments. The expression for optimizing decision-making is:
[0078] ;
[0079] Where, is the global optimal action combination, which represents the optimal decision of all agents in the current state; Represented as agent state; represents the actions performed by the agent; Represented as for all agents The rewards and penalties are summed up to optimize the overall system performance; Represented as an agent In state Next action The state-action value of is used to reflect the long-term expected return of the current strategy; Denoted as target tracking reward; Expressed as collaborative rewards between agents, it is used to encourage data sharing and task collaboration; Expressed as a penalty for path conflict, it is used to avoid collisions or repeated tasks between agents;
[0080] Among them, target tracking reward , the calculation expression based on the deviation between the target position and the current path is:
[0081] ;
[0082] Where, Represented as the real-time position of the agent; Represented as target position; It is expressed as a measurement tool to quantify the tracking effect;
[0083] Among them, the penalty for path conflict , the calculation expression used to avoid collisions or repeated tasks between agents is:
[0084] ;
[0085] Where, Indicates the strictness of conflict avoidance for control; Expressed as the potential risk of a quantified action.
[0086] Furthermore, in S7, a virtual scene supporting underwater environment simulation was built in the simulation platform, dynamic targets were loaded, and a dynamic scene was constructed in which three friendly intelligent agents cooperated to surround and capture an enemy intelligent agent, simulating the collaborative capture task of multiple intelligent agents in a complex underwater environment, and then further conducting actual sea trials to verify the collaborative capture of multiple intelligent agents.
[0087] The present invention discloses an underwater moving target tracking system based on multi-agent Bayesian reinforcement learning. An embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0088] Combine Figure 1 and Figure 2 As shown in the figure, the underwater moving target tracking system based on multi-agent Bayesian reinforcement learning is applied to multiple UUV agents. UUV agents are underwater vehicles that navigate without a pilot and rely on remote control or automatic control. They are mainly used for high-risk underwater operations such as deep-sea exploration, rescue, and mine clearance. Each UUV agent is equipped with a wireless communication module, a perception module, a data preprocessing module, and a decision-making module. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning specifically includes the following steps:
[0089] The first step is to design a multi-UUV agent architecture. A target tracking architecture based on multiple UUV agents is established, with multiple UUV agents forming a collaborative network via wireless communication modules. Each UUV agent possesses independent perception, decision-making, and execution capabilities, and can share data and information with other UUV agents in real time, ensuring the overall coordination and efficiency of the system. To achieve efficient coordination, the system employs a distributed task allocation mechanism and a multi-UUV agent collaboration strategy. Each UUV agent makes decisions based on the target's location, environmental conditions, and collaboration with other agents to ensure efficient task execution. Multiple UUVs are divided into master and auxiliary UUVs. The master UUV receives and integrates information from all UUVs, performs data fusion, and develops a rough tracking path using a global planning algorithm. The auxiliary UUVs dynamically adjust their action strategies based on the local environment to refine and optimize the task. During the multi-UUV agent collaboration process, information sharing and feedback mechanisms ensure real-time target tracking and path optimization, avoid duplication or conflicting tasks, and improve overall system efficiency.
[0090] The second step is dynamic modeling. Establish a coordinate system with the center of the UUV body as the origin, such as Figure 2As shown in the figure, the coordinate system established with the center of the UUV as the origin contains three mutually perpendicular axes: the x-axis, the y-axis, and the z-axis. The motion of the UUV is then decomposed into translational motion along the x-, y-, and z-axes and rotational motion around them. Based on the principles of kinematics, the relationship between the linear velocity and angular velocity of the UUV on the x-, y-, and z-axes and its position, velocity, and acceleration is established. Based on Newton's second law and Euler's equations, the six-degree-of-freedom motion equations of the UUV in the body coordinate system are derived, including linear motion in three directions and angular motion in three directions. The specific dynamic equations describe in detail the various forces and moments acting on the UUV in the underwater environment, including the center of gravity, center of buoyancy, inertial force, etc.
[0091] Center of gravity of intelligent UUV The position vector expression is:
[0092] ;
[0093] Where, Represented as the center of gravity of the agent frame along the x-axis; Represented as the center of gravity of the agent frame along the y-axis; It is represented as the center of gravity of the agent frame along the z-axis;
[0094] Center of buoyancy of intelligent UUV The position vector expression is:
[0095] ;
[0096] Where, It is represented as the center of buoyancy of the agent frame along the x-axis; It is represented as the buoyancy center of the agent frame along the y-axis direction; It is represented as the center of buoyancy of the agent frame along the z-axis.
[0097] The third step is agent perception and data collection and sharing. In practice, UUV agent perception and data collection and sharing are key components of a multi-UUV agent collaborative system. Each UUV agent uses a perception module to perceive the surrounding underwater environment and target status in real time, including information such as target position, speed, and heading. The perception module includes sonar sensors and other sensors, which the UUV agent uses to detect the relative positions of underwater obstacles, other UUV agents, and the target object. Because sensor data in complex underwater environments is often affected by noise and interference, a data processing module is designed to pre-process the data collected by the perception module, extract key information, reduce noise interference, and input the processed data into the decision module. To achieve efficient information sharing and fusion, the system adopts a centralized or distributed data processing framework. In a centralized framework, the master UUV agent is responsible for receiving and integrating information from all agents, performing data fusion and global optimization. After data sharing, the UUV agents cooperate and coordinate to optimize overall mission execution.
[0098] The fourth step is to build a Bayesian inference model. In response to the uncertainty and complexity of the underwater environment, the Bayesian inference method is introduced to establish a dynamic estimation model for the target state. Each UUV agent uses Bayesian filtering technology to estimate the target state by combining the real-time data obtained by local sensors and the target information shared by other agents. By fusing prior knowledge and perception data, the target's motion trajectory and behavior pattern are updated in real time, improving the reliability of the system's decision-making in underwater conditions. Each UUV agent collects environmental data through local sensors and uses the data from each sensor to update the target state according to a pre-defined sensor model. The relative distance and direction of the target are measured using sonar sensors. At this time, the sensor data is used as observation information and input into the Bayesian filtering model to estimate the true state of the target.
[0099] Dynamic estimation model expression based on Bayesian inference method:
[0100] ;
[0101] Where, It is represented as a global state, i.e., the motion state of the target, i.e., position and velocity; Expressed as Perception data acquired by each UUV agent; Represented as a prior distribution of the target state, reflecting the target's initial assumptions or historical behavior; For the The likelihood function of the UUV agent’s perception data is expressed as When, UUV agents acquire perception data probability; Represented as the perception data of all given UUV agents After that, the target state The posterior distribution of .
[0102] The fifth step is to optimize the reinforcement learning algorithm. A reinforcement learning algorithm based on Bayesian reasoning is developed, and a learning framework consisting of a state space, an action space, and a reward function is designed. By combining multi-UUV agent collaboration with a Bayesian reinforcement learning algorithm, accurate tracking of underwater targets is achieved. Each agent uses a reinforcement learning algorithm to make local decisions based on its local perception data and state estimation results. The system uses collaborative path planning and task allocation algorithms to ensure that the motion trajectories of each agent do not conflict and that the target tracking task can be completed efficiently. The reward function comprehensively considers target tracking accuracy, energy consumption optimization, and inter-agent collaboration efficiency to ensure that the system's reinforcement learning results can adapt to changing environments.
[0103] The UUV agent's own state at this time Next, the UUV agent selects an action To maximize the overall reward , which is rewarded by tracking performance , energy efficiency and synergy The three parts are weighted to obtain the path optimization expression of the reinforcement learning algorithm based on Bayesian reasoning:
[0104] ;
[0105] Where, Represented as the UUV agent at time Action choices; Indicates that the status Next, the UUV agent selects an action To maximize the overall reward; It is expressed as a reward for target tracking accuracy, which is related to the target distance or tracking error; It is expressed as an energy consumption optimization reward, which is related to the control of speed and acceleration; It is expressed as a reward for the collaborative efficiency between UUV agents, which is related to the distance and degree of cooperation between UUV agents.
[0106] The sixth step is task scheduling in multi-target dynamic scenarios. For multi-target tracking scenarios, the system has designed a dynamic task scheduling mechanism. When multiple targets appear within the perception range, the UUV agent can automatically assign tasks based on factors such as the priority, distance, and speed of the target UUV agent, and adjust the task load of each UUV agent in real time to ensure that the entire system can efficiently and collaboratively complete the target tracking task. Through the distributed collaboration of multiple UUV agents, dynamic collaborative strategies are formulated, including target allocation, path optimization, and obstacle avoidance planning. According to local perception information and global target priority, their own behavior is dynamically adjusted to achieve efficient tracking of dynamic targets.
[0107] In practice, in dynamic scenarios with multiple targets, the dynamic task scheduling mechanism needs to determine the priority of the targets. When multiple targets enter the sensing range, the dynamic task scheduling mechanism calculates the priority based on the target's distance, speed, and importance factors.
[0108] The calculation expression of target priority is:
[0109] ;
[0110] Where, Represented as a target Priority calculation; Expressed as priority weight; Expressed as velocity weight; Expressed as target weight; Represents the target The Euclidean distance to the UUV agent, in meters. The closer the distance, the higher the priority weight. The greater the contribution; Expressed as target speed in m / s. The faster the speed, the higher the tracking urgency. The speed weight Corresponding dynamic adjustment capabilities; Indicates the preset importance level for the target, , target weight response task preference;
[0111] Each weight coefficient must satisfy the normalization constraint:
[0112] ;
[0113] The task of tracking the target UUV agent is assigned to the most suitable UUV agent through dynamic task allocation. Each UUV agent combines reinforcement learning and path planning algorithms to develop a trajectory with the lowest energy consumption and the best time. At the same time, path information is shared through wireless communication modules to avoid path conflicts and collisions between UUV agents. During the execution of the task, the system continuously monitors the target and agent status and adjusts the task allocation in real time. When the position of the target UUV changes or the task load of the UUV agent is too high, the system will dynamically reallocate tasks to ensure the continuity and efficiency of the tracking task. In addition, through collaborative strategies, UUV agents can jointly track important targets or divide the work to handle multiple targets, improving the overall system's coordination and tracking accuracy. The entire task scheduling process evaluates indicators such as task completion rate, energy efficiency, and path conflicts in real time, and further optimizes the scheduling strategy based on feedback to ensure the stability and efficiency of the system in complex dynamic environments.
[0114] Each agent combines path optimization based on Bayesian reasoning reinforcement learning algorithm to develop the trajectory with the lowest energy consumption and the best time. The objective function of path planning is defined as a multi-objective optimization problem to determine the total cost of the appropriate agent. Calculation expression:
[0115] ;
[0116] Where, Expressed as the weight coefficient of time, is expressed as the weight coefficient of energy consumption, and and Need to meet ; Expressed as total task time; Expressed as starting point coordinates; Expressed as the end point coordinates; It is expressed as the maximum speed of the UUV agent; Represented as control input, such as acceleration; Expressed as the total cost when the control input is minimized.
[0117] The seventh step is real-time decision-making and control mechanism. The real-time decision-making and control mechanism based on Bayesian reinforcement learning achieves efficient task optimization and execution through comprehensive target tracking, agent collaboration and path conflict avoidance. Each UUV agent perceives the target state in real time through sensors, updates the target position using Bayesian reasoning, and selects the optimal action in combination with reinforcement learning algorithms to maximize long-term rewards. The system ensures that the agent achieves a balance between target tracking accuracy, collaboration efficiency and path conflict avoidance through dynamic task allocation and path planning. At the same time, combined with real-time feedback and model predictive control, the UUV agent can dynamically adjust the path and strategy to adapt to the rapid changes of the target and the interference of the complex environment, ensuring the stability of tracking performance in multi-target scenarios.
[0118] The multi-UUV intelligent agent combines the accuracy reward of target tracking, the improvement of collaborative efficiency, and the penalty of path conflict to achieve real-time collaboration and decision-making in complex environments. The expression for optimizing decision-making is:
[0119] ;
[0120] Where, is the global optimal action combination, which represents the optimal decision of all UUV agents in the current state; Represented as agent state; represents the actions performed by the agent; Represented as for all UUV agents The rewards and penalties are summed up to optimize the overall system performance; Represented as a UUV agent In state Next action The state-action value of is used to reflect the long-term expected return of the current strategy; Denoted as target tracking reward; Expressed as collaborative rewards between UUV agents, it is used to encourage data sharing and task collaboration; Expressed as a penalty for path conflict, it is used to avoid collisions or repeated tasks between agents;
[0121] Among them, target tracking reward , the calculation expression based on the deviation between the target position and the current path is:
[0122] ;
[0123] Where, Expressed as the real-time position of the UUV; Represented as target position; It is expressed as a measurement tool to quantify the tracking effect;
[0124] Among them, the penalty for path conflict , the calculation expression used to avoid collisions or repeated tasks between agents is:
[0125] ;
[0126] Where, Indicates the strictness of conflict avoidance for control; Expressed as the potential risk of a quantified action.
[0127] The eighth step is system verification and actual sea trial verification. By constructing a multi-target dynamic scene and simulating the task execution process in a complex underwater environment, the effectiveness and stability of the algorithm are verified. First, a virtual scene that supports underwater environment simulation is built in the simulation platform, dynamic targets are loaded, and using the Unity3D simulation platform, a dynamic scene is constructed in which three friendly UUV agents cooperate to encircle an enemy UUV agent, simulating the collaborative encirclement task of multiple UUV agents in a complex underwater environment. This verifies the feasibility, stability, and efficiency of the system's encirclement strategy, and optimizes the algorithm strategy and system parameters based on the experimental results. In order to further explore the feasibility and performance of the strategy in actual scenarios and verify whether they can accurately detect targets and effectively cooperate, after passing the system test verification on a high-precision virtual simulation platform, further actual sea trial verification of the collaborative encirclement of multiple UUV agents is carried out.
[0128] Example: This invention conducted a practical sea trial of a multi-UUV swarm in a certain sea area. The scenario involved an enemy UUV infiltrating our waters. Upon observing the enemy, our team dispatched a swarm of three UUVs, numbered UUV1, UUV2, and UUV3. Our team needed to quickly predict the enemy UUV's movements and then, through multi-agent Bayesian reinforcement learning, optimize the multi-agent path planning and encirclement strategy to effectively encircle and intercept the enemy UUV. This experiment is illustrated below with examples.
[0129] The first step is to assemble and debug a multi-UUV cluster consisting of three UUVs after arriving at the site, and identify them as UUV1, UUV2 and UUV3 in the system. Then, a collaborative network is formed between the UUV clusters through the wireless communication module. Then, UUV1, UUV2, UUV3 and multiple target UUVs are placed in the water and arrive at the appropriate location.
[0130] The second step is to determine the basic physical parameters of the UUV. All UUVs in the experiment use the same model. Taking UUV1 as an example, the buoyancy center of the UUV is set as the origin of the fixed coordinate system to obtain the center of gravity of the intelligent UUV. :
[0131] ;
[0132] Get the center of buoyancy of the intelligent UUV :
[0133] ;
[0134] The third step is intelligent perception and data collection and sharing. After completing sensor installation and communication networking for our three UUVs, each UUV activated its sonar for calibration. Sensor data collection confirmed that performance met requirements: target detection probability within 500 meters ≥ 95%, coordinated positioning accuracy of the three UUVs ≤ 3 meters, target trajectory continuity ≥ 99% during the test, and data transmission latency ≤ 1.5 seconds. When any UUV detected a target, the swarm automatically entered advanced tracking mode, increasing data sharing frequency to 1 Hz.
[0135] The fourth step is to build a Bayesian reasoning model. The specific calculation process of the Bayesian reasoning model in multi-UUV collaborative tracking is as follows: Set the position and speed of the target state of the target UUV as ; At this time, the positions of our three UUVs are: UUV1's initial position ; Initial position of UUV2 ; Initial position of UUV3 ; At this time, the sensor detects distance noise , distance noise angle , the covariance matrix , assuming that the prior distribution is Gaussian distribution, the position and velocity of the prior state are obtained from the previous step prediction: , : .
[0136] The observation data is calculated to obtain the observation of UUV1: the true relative position is given by , true distance , true angle , the actual observed distance and angle of UU1 after adding noise: , noise distance: , noise angle: ; The observation data is calculated to obtain the observation of UUV2: the real relative position is given by , true distance: , true angle: , the actual observed distance and angle of UUV2 after adding noise: , noise distance: , noise angle: ; The observation data is calculated to obtain the observation of UUV3: the true relative position is given by , true distance: , true angle: , the actual observed distance and angle of UUV3 after adding noise: , noise distance: , noise angle: .
[0137] Based on the prior state, the relative position of UUV1 is obtained: , predict the distance and angle of UUV1: ; Get the relative position of UUV2: , predict the distance and angle of UUV2: ; Get the relative position of UUV3: , predict the distance and angle of UUV3: .
[0138] According to the two sets of observation data of the three UUVs mentioned above, the observation residuals of the three UUVs are obtained, namely, the observation residual of UUV1: ;
[0139] Observation residuals of UUV2: ;
[0140] Observation residuals of UUV3: .
[0141] Perform three UUV Jacobian matrix calculations, UUV1 Jacobian matrix: , that is, we get ;
[0142] Similarly, the UUV2 Jacobian matrix can be obtained: , UUV3 Jacobian matrix: .
[0143] Multi-sensor data fusion, combined global matrix: , residual .
[0144] Observation noise covariance matrix ;
[0145] Calculate the forecast error covariance matrix , Kalman gain , update the result after calculation, status update , covariance update , position correction vector: .
[0146] Numerical implementation of the Bayesian formula to obtain the state after all observation data are given The posterior probability of : ;in, Expressed as Observation data Likelihood function of ; Expressed as a prior probability.
[0147] Prior probability , which means that the mean is , the covariance matrix is Normal distribution;
[0148] Represents observation data The mean is , the covariance matrix is Normal distribution;
[0149] , represents the observed data The mean is , the covariance matrix is Normal distribution;
[0150] , represents the observed data The mean is , the covariance matrix is Normal distribution.
[0151] By solving the point where the gradient of the logarithmic posterior probability is zero , we can get the peak position of the posterior distribution, , indicating that the maximum value of the posterior distribution occurs at Place.
[0152] Error calculation: True target position: , estimated position: , position error: , this error range meets the requirements of round-up.
[0153] like Figure 4 As shown in the figure, it is a graph of the heading angle and speed of UUV1, UUV2 and UUV3 changing with time. It can be seen from the figure that the blue curve represents the heading angle and the orange curve represents the speed. The three UUVs are actively tracking the target UUV and have reached relatively stable heading angles and speeds at different time points. After that, the heading angles and speeds of the three UUVs have the same change trend. Figure 5 The changes in the distance between the three UUVs and the target UUV and the angle between the three UUVs can be seen from the figure. The three UUVs approach the target UUV from different positions. From 300s on, the distance from the target UUV gradually stabilizes, and a relatively stable capture state is maintained. Finally, a stable equilateral triangle capture structure is formed. The three curves of the angle between the three UUVs are close to 120° for most of the time, with small fluctuations, maintaining a relatively stable shape.
[0154] The fifth step is to optimize the reinforcement learning algorithm. When the target estimated position and speed are , at this time UUV1: position ,speed , UUV2: Position ,speed , UUV3: Position ,speed , time step .
[0155] Total Rewards It consists of three parts, including target tracking reward , energy consumption reward , collaborative rewards , Define the UUV action space and the velocity increment of each UUV: , including Speed change in direction and Speed change in direction ,in , velocity update equation: .
[0156] Target Tracking Reward ;
[0157] The distance between UUV and target , the maximum effective distance of the sensor ,parameter Taking UUV1 as an example, the position ,speed , when choosing the maximum acceleration to move eastward, , , the distance between UUV1 and the target , target tracking reward , energy consumption reward ; Among them, the parameters , , substituting the UUV1 data into the equation: , , , Collaboration Rewards: ; The minimum distance between two UUVs is required to be 10m, where the parameters , , ideal coordination distance ;
[0158] The distance between UUV1 and UUV2 is: , the distance between UUV1 and UUV3 is: , collaborative rewards , the total reward of UUV1 for this action By updating the strategy through the reinforcement learning algorithm, the target tracking error is finally achieved < 6m while ensuring zero collision coordination.
[0159] Depend on Figure 6 As shown, the blue line represents UUV1, the yellow line represents UUV2, the green line represents UUV3, and the red line represents the target UUV. Our three UUVs start from different positions and approach the target UUV, and then drive our UUVs to converge towards the target UUV. Our UUV1, UUV2, and UUV3 work together to approach the target and form an equilateral triangle arrangement. The target UUV is located at the center of the formed equilateral triangle, forming a stable encirclement state.
[0160] Task scheduling in a multi-target dynamic scenario. Set two targets, target 1 (T1) position ,speed , importance level (Highest level), Target 2 (T2) position ,speed , importance level (Normal level), initial position: UUV1 position , UUV2 position , UUV3 position , T1 speed , T2 speed .
[0161] The priority formula of the goal: ;
[0162] Set weight factor: distance weight , speed weight , importance weight , calculate the distances and priorities of the three UUVs to targets T1 and T2 respectively.
[0163] Distance from UUV1 to target T1 , the distance from UUV1 to target T2 , UUV1's priority for target T1 , UUV1's priority for target T2 ;
[0164] Distance from UUV2 to target T1 , the distance from UUV2 to target T2 , UUV2's priority for target T1 , UUV2's priority for target T2 ;
[0165] Distance from UUV3 to target T1 , the distance from UUV3 to target T2 , UUV3's priority for target T1 , UUV3's priority for target T2 .
[0166] To optimize task allocation, tasks need to be allocated based on these priorities so that the total priority is maximized, i.e. the cost is minimized. Cost = 1-priority. Construct a cost matrix , after row and column reduction , and the result is that target T1 is tracked collaboratively by UUV1 and UUV3, and target T2 is tracked alone by UUV2.
[0167] Path planning objective function , where the time weight , energy consumption weight , the maximum speed of the UUV is known to be , exercise time , when the starting position of UUV1 , end position Straight-line distance , time cost , assuming constant acceleration motion, control input , then the equation of motion , we can get , energy consumption cost , total cost , it is verified that the path with the minimum total cost is obtained.
[0168] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.
Claims
1. An underwater moving target tracking system based on multi-agent Bayesian reinforcement learning, characterized by: The following steps are involved: S1, establish a multi-agent-based target tracking architecture, so that multiple agents form a collaborative network and share the collected data; S2, establish a coordinate system with the center of the intelligent agent as the origin, and derive the kinematic equations of the intelligent agent in this coordinate system based on Newton's second law and Euler's equation; S3, introduces the Bayesian reasoning method to establish a dynamic estimation model of the target state; S4, develop reinforcement learning algorithms based on Bayesian reasoning to optimize path planning among multiple agents; S5, in the scenario of multiple dynamic targets, determines the targets to be tracked through the dynamic task scheduling mechanism and assigns the tasks to the appropriate agents; S6, optimizing multi-agent decision making in complex environments; S7, construct and simulate the collaborative capture task of multiple agents in a complex underwater environment in the simulation platform; In S4, the agent is in state s t Next, the agent chooses action a t To maximize the comprehensive reward R total , the reward is determined by tracking performance R tracking , energy efficiency R energy and synergistic effect R collaboration The three parts are weighted to obtain the path optimization expression of the reinforcement learning algorithm based on Bayesian reasoning: Where a t represents the action choice of the agent at time t; In state s t Next, the agent chooses action a t To maximize the comprehensive reward; R tracking represents the reward for target tracking accuracy, which is related to the target distance or tracking error; R energy represents the energy consumption optimization reward, which is related to the control of speed and acceleration; R collaboration The reward for the efficiency of collaboration between agents is related to the distance and degree of cooperation between agents. In S5, in a multi-target dynamic scenario, the dynamic task scheduling mechanism determines the target to be tracked by calculating the target's priority. When multiple targets enter the sensing range, the dynamic task scheduling mechanism calculates the priority based on the target's distance, speed, and importance factors. The calculation expression of target priority is: Where, P i represents the priority calculation of target i; w d represents the priority weight; w v represents the speed weight; w s represents the target weight; d i Represents the Euclidean distance between target i and the agent, in m. The closer the distance, the higher the priority weight w. d The greater the contribution; i Indicates the target speed in m / s. The faster the speed, the higher the tracking urgency. The speed weight w v Corresponding dynamic adjustment capability; s i Indicates the preset importance level of the target, s i ∈{1,2,3}, target weight w s Reflects task preferences.
2. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 1 is characterized in that: In S1, multiple agents form a collaborative network through wireless communication modules. Each agent has independent perception, decision-making and execution functions, and can share information with other agents in real time. Multiple intelligent agents are divided into main intelligent agents and auxiliary intelligent agents. The main intelligent agent is used to receive and integrate information from all intelligent agents, and perform data fusion and global optimization; the auxiliary intelligent agent dynamically adjusts its actions according to the local environment.
3. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 2 is characterized in that: Each agent is equipped with a perception module, which is used to perceive the underwater environment and the state of the target, including the target's position, speed and heading information. The perception module includes a sonar sensor, which is used to detect the relative position of underwater obstacles, other agents and the target object. Each intelligent agent is equipped with a data preprocessing module, which preprocesses the data collected by the perception module to extract key information and reduce the impact of noise interference.
4. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 3 is characterized by: In S2, the coordinate system established with the center of the intelligent body as the origin contains three mutually perpendicular axes, namely the x-axis, y-axis, and z-axis. The motion of the intelligent body is then decomposed into translational motion along the x-, y-, and z-axes and rotational motion around the x-, y-, and z-axes. According to the principles of kinematics, the relationship between the linear velocity, angular velocity, position, velocity, and acceleration of the intelligent body on the x-, y-, and z-axes is established respectively. Based on Newton's second law and Euler's equations, the six-degree-of-freedom motion equations of the intelligent agent in the body coordinate system are derived, including three-dimensional linear motion and three-dimensional angular motion. The motion equations describe the forces and torques acting on the intelligent agent in the underwater environment. The center of gravity r of the agent G The position vector expression is: Where x g represents the center of gravity of the intelligent body frame along the x-axis; y g represents the center of gravity of the intelligent body frame along the y-axis; z g Represents the center of gravity of the agent frame along the z-axis; The center of gravity of the agent r B The position vector expression is: Where x b represents the center of buoyancy of the intelligent body frame along the x-axis; y b represents the buoyancy center of the intelligent body frame along the y-axis direction; z b Represents the center of buoyancy of the agent frame along the z-axis.
5. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 4 is characterized in that: In S3, each agent combines the relative distance and speed of the target obtained by the sonar sensor with the target information shared by other agents, and uses a dynamic estimation model based on Bayesian reasoning to estimate the target state. By fusing prior knowledge and perception data, the target's motion trajectory and behavior pattern are updated in real time. Dynamic estimation model expression based on Bayesian inference method: Where x a Represents the global state, that is, the motion state of the target, including position and speed; z i represents the observation value obtained by the i-th sensor; p(x a ) represents the prior distribution of the target state, reflecting the initial assumption or historical behavior of the target; p(z i |x a ) is the likelihood function of the i-th agent’s perception data, which indicates that when the target state x is known a When the i-th agent obtains the perception data z i The probability of p(x a |z1,z2,…,z N ) represents the perception data z1,z2,…,z given all agents N After that, the target state x a The posterior distribution of .
6. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 5, characterized in that: In S4, a reinforcement learning algorithm based on Bayesian reasoning is developed, and a learning framework including state space, action space and reward function is designed. By combining multi-agent collaboration with the Bayesian reinforcement learning algorithm, path planning and task allocation between agents are optimized to ensure that there is no conflict between the tasks and paths of each agent.
7. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 6, characterized in that: In S5, each weight coefficient must satisfy the normalization constraint: In d +in v +in s =1; The normalization constraints are satisfied by each weight coefficient to ensure that the priority score is a standardized value.
8. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 7, characterized in that: in, After determining the target to be tracked, the system continuously monitors the target and agent status, adjusting task allocation in real time to assign the target task to the appropriate agent. When the target position changes or the agent's task load becomes too high, the system dynamically reallocates tasks. Throughout the dynamic task scheduling process, task completion rates, energy efficiency, and path conflict indicators are evaluated in real time, and scheduling strategies are further optimized based on feedback. Each agent combines path optimization based on Bayesian reasoning reinforcement learning algorithm to develop a trajectory with the lowest energy consumption and the best time. The objective function of path planning is defined as a multi-objective optimization problem. The total cost J of the appropriate agent is calculated as: In the formula, α represents the weight coefficient of time, β represents the weight coefficient of energy consumption, and α and β must satisfy α+β=1; T represents the total task time; p start Indicates the starting point coordinates; p end Indicates the end point coordinates; v max represents the maximum velocity of the agent; u(t) represents the control input, including acceleration; Represents the total cost when the control input is minimized.
9. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 8, characterized in that: In S6, multi-agents combine target tracking accuracy rewards, collaborative efficiency improvements, and path conflict penalties to achieve real-time collaboration and decision-making in complex environments. The expression for optimizing decision-making is: Where a * is the global optimal action combination, which represents the optimal decision of all agents in the current state; Represents the state of the agent; represents the action performed by the agent; Represents the sum of the rewards and penalties of all agents i to optimize the overall system performance; Indicates that agent i is in state Next action The state-action value of is used to reflect the long-term expected return of the current strategy; represents the target tracking reward; Represents collaborative rewards between agents, used to encourage data sharing and task collaboration; The penalty for path conflict is used to avoid collisions or repeated tasks between agents; Among them, target tracking reward The calculation expression based on the deviation between the target position and the current path is: R tracking,i =-distance(x UUV,i ,x target ); Where x UUV,i Indicates the real-time position of the agent; x target Indicates the target location; distance is a measurement tool for quantifying tracking effects; Among them, the penalty for path conflict The calculation expression used to avoid collisions or repeated tasks between agents is: Where γ represents the strictness of conflict avoidance control; collision risk represents the potential risk of the quantified action.
10. The underwater moving target tracking system based on multi-agent Bayesian reinforcement learning according to claim 9, characterized in that: In S7, a virtual scene supporting underwater environment simulation was built in the simulation platform, dynamic targets were loaded, and a dynamic scene was constructed in which three friendly intelligent agents cooperated to surround and capture an enemy intelligent agent. This simulated the collaborative capture task of multiple intelligent agents in a complex underwater environment, and then further actual sea trials and verification of the collaborative capture of multiple intelligent agents were carried out.