Multi-auv bearing-only perception cooperative hunting control method based on reinforcement learning
By employing a reinforcement learning-based multi-AUV pure orientation-aware collaborative encirclement control method, and utilizing time-aware depth geometric inversion and explicit belief guidance mechanisms, the method achieves time-delay-compensated reconstruction of target location and collaborative decision optimization, thereby improving the robustness and success rate of multi-AUV collaborative encirclement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-03-16
- Publication Date
- 2026-05-15
AI Technical Summary
Existing multi-AUV cooperative encirclement control methods are difficult to achieve accurate target position estimation and cooperative encirclement decision-making under pure azimuth observation and time-delay communication conditions, resulting in insufficient robustness of cooperative encirclement.
A reinforcement learning-based multi-AUV pure orientation perception cooperative encirclement control method is adopted. By constructing a parameter-sharing multi-agent deep deterministic policy gradient network, a time-aware deep geometric inversion module, an explicit belief guidance module, and an adaptive context gating fusion module are integrated to perform target belief estimation and cooperative decision optimization.
It improves target positioning accuracy and collaborative capture success rate under pure azimuth observation and time-delay communication conditions, and enhances obstacle avoidance safety and robustness in complex environments.
Smart Images

Figure CN121832631B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater unmanned systems and intelligent control technology, and in particular to a multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning. Background Technology
[0002] With the continuous advancement of marine intelligent equipment, underwater sensing, and underwater acoustic communication technologies, multi-autonomous underwater vehicle (AUV) collaborative operations have been widely applied in numerous marine-related fields. In maritime security scenarios, multi-AUV collaborative operations can construct a comprehensive monitoring network to effectively prevent illegal intrusion; in target interception missions, they can cooperate with each other to quickly intercept threatening targets; in marine inspection, they can achieve efficient patrols of large sea areas and promptly detect potential problems; in the field of resource exploration, multi-AUV collaboration can more comprehensively and accurately detect the distribution of marine resources.
[0003] Cooperative encirclement and control, as a key technology in multi-AUV swarm missions, faces numerous challenges. In the dynamically changing marine environment, each AUV needs to continuously track and cooperatively encircle non-cooperative targets, while ensuring its own safety and maintaining a stable formation under the constraints of complex ocean topography and various obstacles. In actual underwater environments, the performance of cooperative encirclement largely depends on the accurate acquisition of the target's status, especially its location.
[0004] Currently, sonar technologies used to acquire target status information each have their advantages and disadvantages. While active sonar can provide both range and azimuth information simultaneously, its active signal transmission nature easily exposes the position and operational intentions of AUV platforms, which is extremely disadvantageous for missions requiring covert operation. In contrast, passive sonar offers the advantage of concealment; however, it typically only obtains pure azimuth observation data containing noise and lacking range information. This means that in a single observation, the target range cannot be directly obtained, resulting in the target's status exhibiting typical partially observable characteristics. This characteristic makes accurately determining the target's location very difficult, thus affecting the effectiveness of coordinated encirclement and suppression.
[0005] Furthermore, underwater acoustic communication also presents challenges in multi-AUV collaborative operations. Due to its low propagation speed and limited bandwidth, underwater acoustic communication suffers from significant latency, which varies with the distance between AUVs. This results in asynchronous and outdated information not arriving simultaneously and promptly. Such delays and inconsistencies severely impact real-time collaboration and information consistency among multiple AUVs, hindering their efficient coordination during collaborative capture missions.
[0006] To address the aforementioned issues, existing solutions mostly employ a phased approach of positioning and control, or rely on filtering algorithms and motion models to estimate the target state before implementing encirclement and control. However, when faced with complex situations such as high noise in pure azimuth observations, unknown target maneuvers, and significant communication delays, traditional methods often fail to simultaneously achieve positioning accuracy, control stability, and collaborative efficiency. On one hand, high observation noise can lead to significant deviations in positioning results; on the other hand, unknown target maneuvers increase the difficulty of accurately estimating the target state; and communication delays can cause lags in control decisions, thereby affecting collaborative efficiency and control stability.
[0007] In recent years, multi-agent deep reinforcement learning has provided a new technical path for achieving distributed collaborative decision-making in unknown environments. However, existing multi-agent reinforcement learning methods are often based on some idealized assumptions, such as the assumption that the observation information is relatively complete or the communication latency is low. Moreover, these methods often decouple the effects of perception uncertainty and latency from the control strategy, without fully considering their interrelationship. However, in the actual underwater environment, pure orientation perception uncertainty and information latency are coupled, and this decoupling approach is insufficient to effectively solve the problem of insufficient robustness in collaborative trapping caused by this.
[0008] In summary, existing multi-AUV cooperative encirclement control methods have limitations in addressing the challenges posed by pure azimuth observation and time-delay communication, making them unsuitable for practical applications. Therefore, there is an urgent need to propose a method that can perform end-to-end joint optimization of target localization and cooperative encirclement decision-making under conditions of pure azimuth observation and time-delay communication. Summary of the Invention
[0009] The technical problem to be solved by this invention is to provide a multi-AUV pure azimuth perception cooperative encirclement control method based on reinforcement learning, which is used to solve the problem that it is difficult to accurately achieve end-to-end joint optimization of target position estimation and cooperative encirclement decision under the conditions of passive sonar pure azimuth observation and underwater acoustic communication with distance-related time delay, thereby improving the success rate and robustness of encirclement and taking into account obstacle avoidance safety in complex environments.
[0010] To achieve the above-mentioned objectives, this invention provides a multi-AUV pure orientation-aware cooperative encirclement and control method based on reinforcement learning, comprising the following steps:
[0011] S1. Training Interaction and Sample Collection Phase:
[0012] Multiple autonomous underwater vehicles (AUVs) interact in a containment environment. Each AUV acquires its own state and passive sonar azimuth observations of the target at the current moment, and receives delayed communication information from neighboring AUVs. They construct observation and state inputs for policy learning and store the samples obtained from the interaction in an experience replay pool.
[0013] S2. Joint Training Phase:
[0014] Under the centralized training and distributed execution CTDE framework, a parameter-sharing multi-agent deep deterministic policy gradient network is constructed. The multi-agent deep deterministic policy gradient network integrates a time-aware deep geometric inversion module, an explicit belief guidance module, and an adaptive context-gated fusion module.
[0015] By sampling sequence samples through the experience replay pool, the global value network and the policy network with target belief estimation are jointly optimized to obtain the trained network parameters.
[0016] S3. Online Execution Phase:
[0017] By loading the network parameters obtained from training, each autonomous underwater vehicle generates inputs for the policy network based on its own state, passive sonar azimuth observation of the target, and delayed communication information with neighboring autonomous underwater vehicles. Based on the actions output by the policy network, control commands are generated to achieve coordinated encirclement and configuration maintenance by multiple autonomous underwater vehicles.
[0018] According to one aspect of the invention, in step S1, the samples stored in the experience replay pool are constructed in the form of six-tuples, wherein the sample is represented as:
[0019] ;
[0020] ;
[0021] ;
[0022] ;
[0023] in, This represents the global state of an autonomous underwater vehicle at adjacent moments. Indicates the first Local observations of an autonomous underwater vehicle at time t. This represents a local set of observations from all autonomous underwater vehicles at adjacent moments. This indicates coordinated actions between autonomous underwater vehicles. Indicates the first The actions output by the policy network of an autonomous underwater vehicle Indicates the first Speed control command for an autonomous underwater vehicle at time t Indicates the first Angular velocity control command of an autonomous underwater vehicle at time t Indicates the number of autonomous underwater vehicles. This indicates that the autonomous underwater vehicle is performing joint operations. They then received a reward.
[0024] According to one aspect of the invention, the local observations of each independent underwater vehicle at time t are obtained by concatenating the passive sonar pure azimuth observations of the target and the delayed communication vector containing an explicit time delay term, based on the vehicle's own state. The local observations of each independent underwater vehicle at time t are expressed as follows:
[0025] ;
[0026] ;
[0027] in, Indicates the first Local observations of an autonomous underwater vehicle at time t. Indicates the first The kinematic information of an autonomous underwater vehicle, including position, velocity, and heading angle. Indicates the first The perception information of an autonomous underwater vehicle includes orientation observation of targets and local map observation of obstacles. Indicates the first Delayed communication information vectors received by an autonomous underwater vehicle Represents a set of neighbors. Indicates the heading angle. Indicates azimuth observation. Indicates the first The location of the autonomous underwater vehicle Indicates the first An autonomous underwater vehicle in Location at any given moment Indicates the first An autonomous underwater vehicle in The speed of time Indicates the first The and the first Communication latency of an autonomous underwater vehicle Indicates the first The time it takes for an autonomous underwater vehicle to transmit information.
[0028] According to one aspect of the invention, in step S2, the time-aware depth geometry inversion module achieves target belief estimation by generating a relative position estimate of the target, comprising:
[0029] The delay-aware feature embedding submodule is used to jointly encode the delayed self-kinematic information and target perception information of neighboring autonomous underwater vehicles and their corresponding communication delays into a delay-aware embedding vector.
[0030] The cyclic memory submodule is used to perform cyclic reasoning on the observation history and delay-aware embedding vector within the time window and output the target belief features;
[0031] The target location estimation module for auxiliary supervision generates a relative location estimate of the target based on the input target belief features and using auxiliary regression.
[0032] According to one aspect of the present invention, in step S2, the step of sampling sequence samples through an experience replay pool, jointly optimizing the global value network and the policy network with target belief estimation to obtain the trained network parameters includes:
[0033] S21. Extract sequence samples of a preset length from the experience replay pool for training the global value network and the policy network with target belief estimation, wherein the policy network is used for target belief estimation based on the time-aware deep geometric inversion module.
[0034] S22. Use extracted sequence samples to minimize temporal difference error to update the network parameters of the global value network;
[0035] S23. The time-aware deep geometry inversion module is trained separately using the extracted sequence samples, wherein the auxiliary regression loss is minimized to update the network parameters of the time-aware deep geometry inversion module.
[0036] S24. The relative position estimate of the target output by the time-aware depth geometry inversion module is used as the input of the tracking belief feature into the policy network, and the gradient blocking operator is used to avoid the policy gradient from destroying the localization representation, thereby realizing explicit belief guidance and gradient decoupling.
[0037] S25. Adaptive context-gated fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target is performed to generate a fusion vector for input to the policy network;
[0038] S26. Under the centralized training and distributed execution CTDE framework, the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target position estimation module under training assistance supervision is introduced as a weighting term to jointly optimize the policy network. Under the condition of blocking the backpropagation of target belief gradient, the tracking belief features after stopping the gradient and the fusion vector are used together to form the fusion state as the input of the policy network.
[0039] S27. Perform soft updates on the target network of the policy network and the global value network, and repeat steps S21 to S26 until the preset training rounds are reached, obtain the trained network parameters, and save them.
[0040] According to one aspect of the invention, in step S22, the step of updating the network parameters of the global value network by minimizing the temporal difference error using the extracted sequence samples, wherein the global value network evaluates the global state and joint actions during the training phase and updates the network parameters of the global value network by minimizing the Bellman error, wherein the loss function for updating the network parameters of the global value network is expressed as:
[0041] ;
[0042] in, Indicates the expectation of a reward. Indicates the length of the training trajectory. Indicates the discount factor. Represents the target value network. Represents a value network. , These are global state features available at adjacent time steps during intensive training. For generated by the target policy network Always in coordination The next-time joint action is generated by the target policy network.
[0043] According to one aspect of the present invention, step S23, which involves training the time-aware depth geometry inversion module separately using extracted sequence samples, includes:
[0044] A delay-aware embedding vector for latency compensation is constructed based on the delay-aware feature embedding submodule, where the delay-aware embedding vector is represented as:
[0045] ;
[0046] ;
[0047] in, Represents a delay-aware embedding vector. Indicates embedded network, This represents a segment of delayed information from a neighboring autonomous underwater vehicle. Indicates the time delay position. Indicates the delay speed. Indicates the target orientation angle during time delay;
[0048] Based on the recurrent memory submodule, the observation history and delay-aware embedding vector within the time window are used for recurrent reasoning to output target belief features;
[0049] The supervised target location estimation module obtains the relative location estimate of the target from the target belief features and completes the update training of the network parameters of the time-aware deep geometric inversion module through supervised training using auxiliary regression loss; wherein, the auxiliary regression loss function used to represent the auxiliary regression loss is:
[0050] ;
[0051] in, This represents the relative position estimate of the target. This indicates the true relative position of the target that can be obtained during the training phase.
[0052] According to one aspect of the present invention, in step S24, in the step of using a gradient blocking operator to avoid gradient violation of the localization representation, the tracking belief feature after stopping the gradient is represented as follows:
[0053] ;
[0054] in, This represents the tracking belief feature after gradient cessation, and is the input state of the policy network. Indicates vector concatenation symbol. This represents the gradient blocking operator.
[0055] According to one aspect of the present invention, in step S25, the step of adaptively context-gated fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target to generate a fusion vector for input to the policy network is represented as follows:
[0056] ;
[0057] ;
[0058] in, Represents the fusion vector. Indicates the gating weight, and , This represents the spatial features of surrounding obstacles extracted from a locally occupied raster map via a convolutional network, and , Indicates the characteristics of tracking beliefs, and , Dimensions representing spatial features The dimension representing the tracking feature. This represents the weight matrix of the gated network. This represents the bias vector of the gated network. Indicates the concatenation operator;
[0059] In step S26, under the centralized training and distributed execution CTDE framework, the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target location estimation module under training supervision is introduced as a weighting term to jointly optimize the policy network. The joint optimization objective is expressed as:
[0060] ;
[0061] in, Represents the policy network, This represents the auxiliary loss weighting coefficient. The parameters represent the policy network. Represents the global state. This indicates a fusion state.
[0062] According to one aspect of the invention, the recurrent memory submodule is constructed based on a recurrent network, wherein the recurrent network is an LSTM network, and the LSTM gated computation is represented as:
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] ;
[0068] in, The output of the input gate determines how much new information should be updated in long-term memory. This represents the weight matrix of the input gate. This represents the bias vector of the input gate. Indicates candidate memory states, This indicates the current state of the cell. The weight matrix representing the candidate states. The bias vector represents the candidate state. This represents the output of the forget gate. This represents the weight matrix of the output gate. This represents the bias vector of the output gate. This indicates a timing input that incorporates both local and delay information. This indicates the hidden state at the current moment. For the Sigmoid function, This is element-wise multiplication.
[0069] The beneficial effects of this plan are as follows:
[0070] According to one aspect of the present invention, under the condition that passive sonar only provides pure orientation information, this scheme achieves time-delay-compensated reconstruction of the target position through time-aware depth geometric inversion, thereby improving the availability and accuracy of positioning. Through explicit belief guidance and gradient decoupling mechanisms, the positioning representation is driven by geometric consistency loss, reducing the interference of policy gradient noise on the belief representation and improving training stability and generalization ability. Through adaptive context-gated fusion, obstacle avoidance safety and capture efficiency are dynamically balanced in complex obstacle environments, improving the robustness of cooperative capture. The use of CTDE and parameter-sharing training methods improves sample efficiency and cooperative consistency, making it suitable for high-latency, partially observable underwater cooperative tasks.
[0071] According to one aspect of the present invention, under the condition of only passive sonar pure azimuth observation and underwater acoustic communication with distance-dependent time delay, this approach can fully realize end-to-end joint optimization of target relative position estimation and collaborative encirclement decision-making, improve the success rate of encirclement, robustness and generalization ability, and take into account obstacle avoidance safety in complex environments.
[0072] According to one aspect of the present invention, this approach can improve the accuracy of pure azimuth positioning and the robustness of multi-AUV cooperative capture under conditions of high latency and partial observability, and has good engineering application value. Attached Figure Description
[0073] Figure 1 This is a flowchart illustrating the steps of the reinforcement learning-based multi-AUV pure orientation perception cooperative encirclement and control method of the present invention.
[0074] Figure 2 This is a flowchart illustrating the principle of the multi-AUV pure orientation perception collaborative encirclement control method based on reinforcement learning according to the present invention. Detailed Implementation
[0075] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present solution will be described in detail below.
[0076] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The embodiments cannot be described in detail here, but the embodiments of the present invention are not limited to the following embodiments.
[0077] Combination Figure 1 and Figure 2 As shown, according to one embodiment of the present invention, a multi-AUV pure orientation perception cooperative encirclement control method based on reinforcement learning includes the following steps:
[0078] S1. Training Interaction and Sample Collection Phase:
[0079] Multiple autonomous underwater vehicles (AUVs) interact in a containment environment. Each AUV acquires its own state and passive sonar azimuth observations of the target at the current moment, and receives delayed communication information from neighboring AUVs. They construct observation and state inputs for policy learning and store the samples obtained from the interaction in an experience replay pool.
[0080] S2. Joint Training Phase:
[0081] Under the centralized training and distributed execution CTDE framework, a parameter-sharing multi-agent deep deterministic policy gradient network is constructed. The multi-agent deep deterministic policy gradient network integrates a time-aware deep geometric inversion module, an explicit belief guidance module, and an adaptive context-gated fusion module.
[0082] By sampling sequence samples through the experience replay pool, the global value network and the policy network with target belief estimation are jointly optimized to obtain the trained network parameters.
[0083] S3. Online Execution Phase:
[0084] By loading the network parameters obtained from training, each autonomous underwater vehicle generates inputs for the policy network based on its own state, passive sonar azimuth observation of the target, and delayed communication information with neighboring autonomous underwater vehicles. Based on the actions output by the policy network, control commands are generated to achieve coordinated encirclement and configuration maintenance by multiple autonomous underwater vehicles.
[0085] Combination Figure 1 and Figure 2 As shown, according to one embodiment of the present invention, in step S1, for a scenario of coordinated encirclement and capture of a target by multiple autonomous underwater vehicles (AUVs), the encirclement and capture formation is assumed to consist of... Composed of multiple autonomous underwater vehicles, discrete time At that time, the first The position, speed, and heading angle of the autonomous underwater vehicle are as follows: , and ; target location is , It represents the set of real numbers.
[0086] In this embodiment, passive sonar pure azimuth observation: under passive sonar conditions, the first An autonomous underwater vehicle that only acquires noisy azimuth observations of a target (without range information) can be represented as:
[0087] ;
[0088] in, Indicates the first An autonomous underwater vehicle in The velocity at any given moment is the observation noise; under two-dimensional horizontal conditions, This is the azimuth observation function.
[0089] Furthermore, receiving delayed communication information from neighboring autonomous underwater vehicles (AUVs): Since AUVs are submerged in water, their communication is underwater acoustic, which suffers from low propagation speed and distance-dependent propagation delay. Therefore, any two AUVs (such as the first...) can receive delayed communication information from neighboring AUVs. The autonomous underwater vehicle and the first Communication latency between autonomous underwater vehicles It can be given by the following formula:
[0090] ;
[0091] in, This is the speed of sound in water.
[0092] Therefore, the first An autonomous underwater vehicle at any time Received from the The messages from an autonomous underwater vehicle satisfy a delayed reception model:
[0093] ;
[0094] in, Indicates the first An autonomous underwater vehicle at any time Received from the The messages from an autonomous underwater vehicle satisfy a delayed reception model. Indicates the first An autonomous underwater vehicle at any time Add communication delay to messages The subsequent delay information.
[0095] In this embodiment, to facilitate the description of the autonomous underwater vehicle's current state and the definition of its actions, specifically, to meet the continuous control requirements of the encirclement maneuver, a continuous action space is adopted, namely the first... An autonomous underwater vehicle at any time Output action:
[0096] ;
[0097] in, For the first Speed control command for an autonomous underwater vehicle at time t For the first Angular velocity control command of an autonomous underwater vehicle at time t.
[0098] In this embodiment, each independent underwater vehicle performs perception and decision-making in a distributed manner, and its available information consists of two parts: local real-time observation and global latency information; wherein, the local real-time observation is expressed by constructing a local observation vector; specifically, the first An autonomous underwater vehicle at any time Constructing local observation vectors It includes at least its own state, passive sonar pure azimuth observations, and a delayed communication vector representing delayed communication information. Therefore, the local observation vector... Represented as:
[0099] ;
[0100] ;
[0101] ;
[0102] in, Indicates the first Local observations of an autonomous underwater vehicle at time t. Indicates the first The kinematic information of an autonomous underwater vehicle, including position, velocity, and heading angle. Indicates the first The sensing information of an autonomous underwater vehicle includes location observations of the target (such as azimuth angle observations). ) and local map observations of obstacles (such as local map / occupied grid) ), Indicates the first Delayed communication information vectors received by an autonomous underwater vehicle Represents a set of neighbors. Indicates the heading angle. Indicates azimuth observation. Indicates the first The location of the autonomous underwater vehicle Indicates the first An autonomous underwater vehicle in Location at any given moment Indicates the first An autonomous underwater vehicle in The speed of time Indicates the first The and the first Communication latency of an autonomous underwater vehicle Indicates the first The time it takes for an autonomous underwater vehicle to transmit information.
[0103] Combination Figure 1 and Figure 2 As shown, according to one embodiment of the present invention, during the training phase, each independent underwater vehicle uses... Or its historical sequence as input based on the output of the network according to the current policy And execute it; receive a reward after execution. With the next observation Therefore, in step S1, the samples stored in the experience replay pool are constructed in the form of six-tuples, where each sample is represented as:
[0104] ;
[0105] ;
[0106] ;
[0107] ;
[0108] in, This represents the global state of an autonomous underwater vehicle at adjacent moments. Indicates the first Local observations of an autonomous underwater vehicle at time t. This represents a local set of observations from all autonomous underwater vehicles at adjacent moments. This indicates coordinated actions between autonomous underwater vehicles. Indicates the first The actions output by the policy network of an autonomous underwater vehicle Indicates the first Speed control command for an autonomous underwater vehicle at time t Indicates the first Angular velocity control command of an autonomous underwater vehicle at time t Indicates the number of autonomous underwater vehicles. This indicates that the autonomous underwater vehicle is performing joint operations. They then received a reward.
[0109] According to one embodiment of the present invention, the local observation of each independent underwater vehicle at time t is obtained by concatenating the passive sonar pure azimuth observation of the target and the delayed communication vector containing an explicit time delay term based on its own state. The local observation of each independent underwater vehicle at time t is expressed as follows:
[0110] ;
[0111] ;
[0112] in, Indicates the first Local observations of an autonomous underwater vehicle at time t. Indicates the first The kinematic information of an autonomous underwater vehicle, including position, velocity, and heading angle. Indicates the first The perception information of an autonomous underwater vehicle includes orientation observation of targets and local map observation of obstacles. Indicates the first Delayed communication information vectors received by an autonomous underwater vehicle Represents a set of neighbors. Indicates the heading angle. Indicates azimuth observation. Indicates the first The location of the autonomous underwater vehicle Indicates the first An autonomous underwater vehicle in Location at any given moment Indicates the first An autonomous underwater vehicle in The speed of time Indicates the first The and the first Communication latency of an autonomous underwater vehicle Indicates the first The time it takes for an autonomous underwater vehicle to transmit information.
[0113] According to one embodiment of the present invention, in step S2, the time-aware depth geometry inversion module achieves target belief estimation by generating a relative position estimate of the target, which includes:
[0114] The delay-aware feature embedding submodule is used to jointly encode the delayed kinematic information of the neighboring autonomous underwater vehicle and the target perception information and their corresponding communication delay into a delay-aware embedding vector; in this embodiment, the delay-aware feature embedding submodule is constructed based on a multilayer perceptron (MLP);
[0115] The cyclic memory submodule is used to perform cyclic reasoning on the observation history and delay-aware embedding vector within the time window and output the target belief features;
[0116] The target position estimation module with auxiliary supervision generates a relative position estimate of the target based on the input target belief features and using auxiliary regression. In this embodiment, the target belief features are used to estimate the relative position of the target based on the geometric regression output layer, and an auxiliary supervision structure with the same topology as the geometric regression output layer is used to supervise the estimation process in an auxiliary regression manner.
[0117] In this embodiment, the time-aware depth geometry inversion module fuses passive sonar pure azimuth observations with time delay information to obtain the target's relative position estimate or tracking belief features, which can then be used for subsequent strategy decisions.
[0118] According to one embodiment of the present invention, the recurrent memory submodule is constructed based on a recurrent network, wherein the recurrent network is an LSTM network, and the LSTM gated computation is represented as follows:
[0119] ;
[0120] ;
[0121] ;
[0122] ;
[0123] ;
[0124] in, The output of the input gate determines how much new information should be updated in long-term memory. This represents the weight matrix of the input gate. This represents the bias vector of the input gate. Indicates candidate memory states, This indicates the current state of the cell. The weight matrix representing the candidate states. The bias vector represents the candidate state. This represents the output of the forget gate. This represents the weight matrix of the output gate. This represents the bias vector of the output gate. This indicates a timing input that incorporates both local and delay information. This indicates the hidden state at the current moment. For the Sigmoid function, This is element-wise multiplication.
[0125] According to one embodiment of the present invention, in step S2, the step of sampling sequence samples through an experience replay pool, jointly optimizing the global value network and the policy network with target belief estimation to obtain the trained network parameters includes:
[0126] S21. Extract sequence samples of a preset length from the experience replay pool for training the global value network and the policy network with target belief estimation, wherein the policy network is based on the time-aware deep geometric inversion module for target belief estimation.
[0127] S22. Use extracted sequence samples to minimize temporal difference error to update the network parameters of the global value network;
[0128] S23. The time-aware deep geometry inversion module is trained separately using the extracted sequence samples, wherein the auxiliary regression loss is minimized to update the network parameters of the time-aware deep geometry inversion module;
[0129] S24. The relative position estimate of the target output by the Time-Aware Deep Geometric Inversion (TDGI) module is used as the input of the tracking belief feature into the policy network, and the gradient blocking operator is used to avoid the policy gradient from destroying the localization representation, thereby realizing explicit belief guidance and gradient decoupling.
[0130] S25. Adaptive Gated Context Fusion (AGCF) is performed on the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target to generate a fusion vector for input to the policy network;
[0131] S26. Under the framework of Centralized Training for Decentralized Execution (CTDE), the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target position estimation module under training assistance supervision is introduced as a weighting term to jointly optimize the policy network. Under the condition of blocking the backpropagation of target belief gradient, the tracking belief features after the gradient is stopped and the fusion vector are used together to form the fusion state as the input of the policy network.
[0132] S27. Perform soft updates on the target network of the policy network and the global value network, and repeat steps S21 to S26 until the preset training rounds are reached, obtain the trained network parameters, and save them.
[0133] According to one embodiment of the present invention, in step S22, the step of updating the network parameters of the global value network by minimizing the temporal difference error using extracted sequence samples, wherein the global value network evaluates the global state and joint actions during the training phase, and updates the network parameters of the global value network by minimizing the Bellman error, wherein the loss function for updating the network parameters of the global value network is expressed as:
[0134] ;
[0135] in, Indicates the expectation of a reward. Indicates the length of the training trajectory. Indicates the discount factor. Represents the target value network. Represents a value network. These are global state features available during intensive training. The next-time joint action is generated by the target policy network.
[0136] According to one embodiment of the present invention, step S23, which involves training the time-aware depth geometry inversion module separately using extracted sequence samples, includes:
[0137] A delay-aware embedding vector for latency compensation is constructed based on the delay-aware feature embedding submodule, where the delay-aware embedding vector is represented as:
[0138] ;
[0139] ;
[0140] in, Represents a delay-aware embedding vector. Indicates embedded network, This represents a segment of delayed information from a neighboring autonomous underwater vehicle. Indicates the time delay position. Indicates the delay speed. Indicates the target orientation angle during time delay;
[0141] Based on the recurrent memory submodule, the observation history and delay-aware embedding vector within the time window are used for recurrent reasoning to output target belief features;
[0142] The supervised target location estimation module obtains the relative location estimate of the target from the target belief features and completes the update training of the network parameters of the time-aware deep geometric inversion module through supervised training using auxiliary regression loss; wherein, the auxiliary regression loss function used to represent the auxiliary regression loss is:
[0143] ;
[0144] in, This represents an estimate of the target's relative position. This indicates the true relative position of the target that can be obtained during the training phase.
[0145] According to one embodiment of the present invention, in step S24, the tracking belief feature after stopping the gradient is represented as follows: (The step of using the gradient blocking operator to avoid gradient violation of the localization representation is described in step S24.)
[0146] ;
[0147] in, This represents the tracking belief feature after gradient cessation, and is the input state of the policy network. This represents the vector concatenation symbol. This represents the gradient blocking operator.
[0148] According to one embodiment of the present invention, in step S25, the step of adaptively context-gated fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target to generate a fusion vector for input to the policy network is represented as follows:
[0149] ;
[0150] ;
[0151] in, Represents the fusion vector. Indicates the gating weight, and , This represents the spatial features of surrounding obstacles extracted from a locally occupied raster map via a convolutional network, and , Indicates the characteristics of tracking beliefs, and , Dimensions representing spatial features The dimension representing the tracking feature. This represents the weight matrix of the gated network. This represents the bias vector of the gated network. This represents the concatenation operator.
[0152] According to one embodiment of the present invention, in step S26, under the centralized training and distributed execution CTDE framework, the step of using the global value network to perform deterministic policy gradient updates on the policy network, and introducing the auxiliary loss of the target position estimation module under training-assisted supervision as a weighting term to jointly optimize the policy network, the joint optimization objective is expressed as:
[0153] ;
[0154] in, Represents the policy network, This represents the auxiliary loss weighting coefficient. The parameters represent the policy network. This represents a centralized value network. Represents the global state. This represents the fusion state formed by the tracking belief features and the fusion vector after the gradient stops.
[0155] Based on the above setup, the entire training phase can be summarized as the following iterative process: initializing the policy network, value network, target network, and experience replay pool; multiple autonomous underwater vehicles interact with the encirclement environment to collect and store samples; when the replay pool samples meet the update conditions, the extraction length is... The sequence samples are used to update the value network sequentially to minimize the loss function. Update the network parameters of the time-aware deep geometric inversion module to minimize the auxiliary regression loss function. Then update the policy network to minimize the joint optimization objective. Finally, perform a soft update on the target network and repeat the above process until the preset number of training rounds is reached, and save the network parameters.
[0156] According to one embodiment of the present invention, in step S3, loading the network parameters obtained from training, each autonomous underwater vehicle generates input for the policy network based on its own state, passive sonar pure azimuth observation of the target, and delayed communication information with neighboring autonomous underwater vehicles, and generates control commands based on the actions output by the policy network, thereby realizing the coordinated encirclement and configuration maintenance of multiple autonomous underwater vehicles, includes:
[0157] S31. Obtain the current state of the autonomous underwater vehicle (AUV), its passive sonar azimuth observation of the target, and delayed communication information with neighboring AUVs to construct a local observation at the current moment. ;
[0158] S32. The relative position estimate of the target is output by the time-aware deep geometric inversion module in the target belief estimation network. Or track belief characteristics ;
[0159] S33. Explicit belief guidance is performed through the explicit belief guidance module in the target belief estimation network, and a fusion vector is generated based on the adaptive context-gated fusion module. The policy network is based on the tracking belief features and fusion vector after stopping gradients. The combined fusion state output action ;
[0160] S34. Repeat steps S31 to S33 until the mission ends, to achieve multi-AUV encirclement configuration maintenance and target encirclement.
[0161] According to the present invention, for underwater collaborative capture scenarios where passive sonar only provides noisy azimuth information and underwater acoustic communication suffers from distance-dependent propagation delays, the solution, within a multi-agent deep reinforcement learning framework, jointly optimizes target localization and collaborative decision-making end-to-end. It introduces relative delay information through a time-aware deep geometric inversion module, performing motion compensation and fusion on delayed azimuth observations from the AUV and teammates to reconstruct and estimate the target's position, using this as an auxiliary learning task to improve feature representation and training stability. Furthermore, an explicit belief guidance mechanism is designed, embedding target tracking beliefs into the policy network, and employing an adaptive context-gated fusion structure to dynamically balance obstacle spatial features and temporal tracking beliefs based on environmental complexity, thereby balancing obstacle avoidance and capture efficiency. During online execution, each AUV generates speed and heading control commands based on local observations and received delay information, forming a stable collaborative capture configuration and improving the capture success rate.
[0162] The above description is merely an example of a specific solution of the present invention. For any devices and structures not described in detail herein, it should be understood that they are implemented using common devices and methods already available in the art.
[0163] The above description is merely one embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning, characterized in that, Includes the following steps: S1. Training Interaction and Sample Collection Phase: Multiple autonomous underwater vehicles (AUVs) interact in a containment environment. Each AUV acquires its own state and passive sonar azimuth observations of the target at the current moment, and receives delayed communication information from neighboring AUVs. They construct observation and state inputs for policy learning and store the samples obtained from the interaction in an experience replay pool. S2. Joint Training Phase: Under the centralized training and distributed execution CTDE framework, a parameter-sharing multi-agent deep deterministic policy gradient network is constructed. The multi-agent deep deterministic policy gradient network integrates a time-aware deep geometric inversion module, an explicit belief guidance module, and an adaptive context-gated fusion module. By sampling sequence samples through the experience replay pool, and jointly optimizing the global value network and the policy network with target belief estimation, the trained network parameters are obtained. S3. Online Execution Phase: The network parameters obtained from training are loaded, and each autonomous underwater vehicle generates inputs for the policy network based on its own state, passive sonar pure azimuth observation of the target, and delayed communication information with neighboring autonomous underwater vehicles. Based on the actions output by the policy network, control commands are generated to achieve cooperative encirclement and configuration maintenance of multiple autonomous underwater vehicles. In step S2, the time-aware depth geometry inversion module generates a relative position estimate of the target to achieve target belief estimation, which includes: The delay-aware feature embedding submodule is used to jointly encode the delayed self-kinematic information and target perception information of neighboring autonomous underwater vehicles and their corresponding communication delays into a delay-aware embedding vector. The cyclic memory submodule is used to perform cyclic reasoning on the observation history and delay-aware embedding vector within the time window and output the target belief features; The target location estimation module for auxiliary supervision generates a relative location estimate of the target based on the input target belief features and using auxiliary regression.
2. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 1, characterized in that, In step S1, the samples stored in the experience replay pool are constructed in the form of six-tuples, where each sample is represented as: in, This represents the global state of an autonomous underwater vehicle at adjacent moments. Indicates the first Local observations of an autonomous underwater vehicle at time t. This represents a local set of observations from all autonomous underwater vehicles at adjacent moments. This indicates coordinated actions between autonomous underwater vehicles. Indicates the first The actions output by the policy network of an autonomous underwater vehicle Indicates the first Speed control command for an autonomous underwater vehicle at time t Indicates the first Angular velocity control command of an autonomous underwater vehicle at time t Indicates the number of autonomous underwater vehicles. This indicates that the autonomous underwater vehicle is performing joint operations. They then received a reward.
3. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 2, characterized in that, The local observations of each independent underwater vehicle at time t are obtained by concatenating the target's passive sonar pure azimuth observations and the delayed communication vector containing explicit time delay terms, based on the vehicle's own state. The local observations of each independent underwater vehicle at time t are represented as follows: in, Indicates the first Local observations of an autonomous underwater vehicle at time t. Indicates the first The kinematic information of an autonomous underwater vehicle, including position, velocity, and heading angle. Indicates the first The perception information of an autonomous underwater vehicle includes orientation observation of targets and local map observation of obstacles. Indicates the first Delayed communication information vectors received by an autonomous underwater vehicle Represents a set of neighbors. Indicates the heading angle. Indicates azimuth observation. Indicates the first The location of the autonomous underwater vehicle Indicates the first An autonomous underwater vehicle in Location at any given moment Indicates the first An autonomous underwater vehicle in The speed of time Indicates the first The and the first Communication latency of an autonomous underwater vehicle Indicates the first The time it takes for an autonomous underwater vehicle to transmit information.
4. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 3, characterized in that, In step S2, the step of sampling sequence samples through the experience replay pool and jointly optimizing the global value network and the policy network with target belief estimation to obtain the trained network parameters includes: S21. Extract sequence samples of a preset length from the experience replay pool for training the global value network and the policy network with target belief estimation, wherein the policy network is used for target belief estimation based on the time-aware deep geometric inversion module. S22. Use extracted sequence samples to minimize temporal difference error to update the network parameters of the global value network; S23. The time-aware deep geometry inversion module is trained separately using the extracted sequence samples, wherein the auxiliary regression loss is minimized to update the network parameters of the time-aware deep geometry inversion module. S24. The relative position estimate of the target output by the time-aware depth geometry inversion module is used as the input of the tracking belief feature into the policy network, and the gradient blocking operator is used to avoid the policy gradient from destroying the localization representation, thereby realizing explicit belief guidance and gradient decoupling. S25. Adaptive context-gated fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target is performed to generate a fusion vector for input to the policy network; S26. Under the centralized training and distributed execution CTDE framework, the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target position estimation module under training assistance supervision is introduced as a weighting term to jointly optimize the policy network. Under the condition of blocking the backpropagation of target belief gradient, the tracking belief features after stopping the gradient and the fusion vector are used together to form the fusion state as the input of the policy network. S27. Perform soft updates on the target network of the policy network and the global value network, and repeat steps S21 to S26 until the preset training rounds are reached, obtain the trained network parameters, and save them.
5. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 4, characterized in that, In step S22, the step of updating the network parameters of the global value network by minimizing the temporal difference error using the extracted sequence samples, the global value network evaluates the global state and joint actions during the training phase and updates the network parameters of the global value network by minimizing the Bellman error. The loss function for updating the network parameters of the global value network is expressed as: in, Indicates the expectation of a reward. Indicates the length of the training trajectory. Indicates the discount factor. Represents the target value network. Represents a value network. , These are global state features available at adjacent time steps during intensive training. For generated by the target policy network Always in coordination The next-time joint action is generated by the target policy network.
6. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 5, characterized in that, Step S23, which involves training the time-aware depth geometric inversion module separately using the extracted sequence samples, includes: A delay-aware embedding vector for latency compensation is constructed based on the delay-aware feature embedding submodule, where the delay-aware embedding vector is represented as: in, Represents a delay-aware embedding vector. Indicates embedded network, This represents a segment of delayed information from a neighboring autonomous underwater vehicle. Indicates the time delay position. Indicates the delay speed. Indicates the target orientation angle during time delay; Based on the recurrent memory submodule, the observation history and delay-aware embedding vector within the time window are used for recurrent reasoning to output target belief features; The supervised target location estimation module obtains the relative location estimate of the target from the target belief features and completes the update training of the network parameters of the time-aware deep geometric inversion module through supervised training using auxiliary regression loss; wherein, the auxiliary regression loss function used to represent the auxiliary regression loss is: in, This represents the relative position estimate of the target. This indicates the true relative position of the target that can be obtained during the training phase.
7. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 6, characterized in that, In step S24, the step of using the gradient blocking operator to avoid gradient destruction of the localization representation is described as follows: The tracking belief feature after stopping the gradient is represented as follows: in, This represents the tracking belief feature after gradient cessation, and is the input state of the policy network. Indicates vector concatenation symbol. This represents the gradient blocking operator.
8. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 7, characterized in that, In step S25, the step of adaptively gating the fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target to generate a fusion vector for input to the policy network is represented as follows: in, Represents the fusion vector. Indicates the gating weight, and , This represents the spatial features of surrounding obstacles extracted from a locally occupied raster map via a convolutional network, and , Indicates the characteristics of tracking beliefs, and , Dimensions representing spatial features The dimension representing the tracking feature. This represents the weight matrix of the gated network. This represents the bias vector of the gated network. Indicates the concatenation operator; In step S26, under the centralized training and distributed execution CTDE framework, the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target location estimation module under training supervision is introduced as a weighting term to jointly optimize the policy network. The joint optimization objective is expressed as: in, Represents the policy network, This represents the auxiliary loss weighting coefficient. The parameters represent the policy network. Represents the global state. This indicates a fusion state.
9. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 1, characterized in that, The recurrent memory submodule is constructed based on a recurrent network, wherein the recurrent network is an LSTM network, and the LSTM gated computation is represented as follows: in, The output of the input gate determines how much new information should be updated in long-term memory. This represents the weight matrix of the input gate. This represents the bias vector of the input gate. Indicates candidate memory states, This indicates the current state of the cell. The weight matrix representing the candidate states. The bias vector represents the candidate state. This represents the output of the forget gate. This represents the weight matrix of the output gate. This represents the bias vector of the output gate. This indicates a timing input that incorporates both local and delay information. This indicates the hidden state at the current moment. For the Sigmoid function, This is element-wise multiplication.