Multi-AUV (Autonomous Underwater Vehicle) pure orientation perception cooperative hunting control method based on reinforcement learning
By employing a reinforcement learning-based multi-AUV pure orientation perception cooperative encirclement control method, and utilizing time-aware depth geometric inversion and explicit belief guidance modules, end-to-end optimization of target position estimation and cooperative decision-making under pure orientation observation and time-delay communication conditions is achieved, thereby improving the robustness and obstacle avoidance safety of multi-AUV cooperative encirclement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-03-16
- Publication Date
- 2026-04-10
AI Technical Summary
Existing multi-AUV cooperative encirclement control methods are difficult to achieve accurate target position estimation and cooperative encirclement decision-making under pure azimuth observation and time-delay communication conditions, resulting in insufficient robustness of cooperative encirclement.
A reinforcement learning-based multi-AUV pure orientation perception collaborative encirclement control method is adopted. Through training interaction and sample collection stages, a multi-agent deep deterministic policy gradient network with shared parameters is constructed. The time-aware deep geometric inversion module, explicit belief guidance module and adaptive context gating fusion module are integrated to achieve end-to-end optimization of target belief estimation and collaborative decision-making.
It improves target positioning accuracy and capture success rate, enhances obstacle avoidance safety and collaborative capture robustness in complex environments, and is suitable for high-latency and partially observable underwater collaborative missions.
Smart Images

Figure CN121832631A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of underwater unmanned systems and intelligent control technology, and particularly relates to a multi-AUV bearing-only perception cooperative hunting control method based on reinforcement learning. BACKGROUND
[0002] With the continuous progress of marine intelligent equipment, underwater sensing and underwater acoustic communication technology, multi-autonomous underwater vehicle (AUV) cooperative operation has been widely applied in many marine-related fields. In the marine security scene, multi-AUV cooperative operation can build a comprehensive monitoring network to effectively prevent illegal intrusion; in the target interception task, they can cooperate with each other to quickly intercept the threat target; in the marine inspection aspect, it can realize efficient patrol of a large area of sea area and timely discovery of potential problems; in the resource exploration field, multi-AUV cooperation can more comprehensively and accurately detect the distribution of marine resources.
[0003] Cooperative hunting control, as one of the key technologies of multi-AUV cluster tasks, faces many challenges. In the dynamically changing marine environment, each AUV needs to achieve sustained tracking and cooperative encirclement of non-cooperative targets, and under the constraints of complex marine topography and various obstacles, ensure its own safety and maintain a stable formation. In the actual underwater environment, the cooperative hunting performance is largely dependent on the accurate acquisition of the target state, especially the target position.
[0004] Currently, sonar technology for acquiring target state information has its advantages and disadvantages. Although the active sonar can provide the distance and bearing information of the target at the same time, due to the characteristics of its active emission signal, it is easy to expose the position and action intention of the AUV platform, which is extremely disadvantageous for some tasks that need to be executed in stealth. In contrast, passive sonar has the advantage of stealth, but it can usually only obtain bearing-only observation data containing noise and without distance information. This results in that the target distance cannot be directly obtained in a single observation, making the target state exhibit typical partially observable characteristics. This characteristic brings great difficulty to accurately grasp the target position, and further affects the effect of cooperative hunting.
[0005] In addition, underwater acoustic communication also has problems in multi-AUV cooperative operation. Due to the low propagation speed of underwater acoustic communication and the limited bandwidth, the communication delay is significant, and this delay will change with the distance between AUVs. This results in that the teammate information cannot be reached at the same time and in a timely manner, and the asynchronous arrival and obsolescence situation occurs. This untimely and inconsistent information seriously affects the real-time cooperation and information consistency among multi-AUVs, making them difficult to efficiently cooperate when performing the cooperative hunting task.
[0006] To solve the above problems, the existing solutions mostly adopt a phased scheme of positioning and control, or rely on filtering algorithms and motion models to estimate the target state, and then perform the pursuit control. However, when facing complex situations such as large bearing-only observation noise, unknown target maneuvering conditions, and significant communication time delay, the traditional methods often cannot simultaneously consider positioning accuracy, control stability, and coordination efficiency. On the one hand, large observation noise can cause large deviation in the positioning result; on the other hand, unknown target maneuvering conditions increase the difficulty of accurately estimating the target state; and the communication time delay can cause the control decision to lag, thereby affecting the coordination efficiency and the stability of the control.
[0007] In recent years, multi-agent deep reinforcement learning has provided a new technical path for realizing distributed collaborative decision-making in unknown environments. However, existing multi-agent reinforcement learning methods are usually based on some idealized assumptions, such as assuming that the observation information is relatively complete or the communication time delay is low. Moreover, these methods often decouple the perception uncertainty and the time delay impact from the control strategy, without fully considering their mutual relationship. However, in actual underwater environments, bearing-only perception uncertainty and information time delay are coupled with each other, and this decoupling approach cannot effectively solve the problem of insufficient robustness of collaborative pursuit caused by this coupling.
[0008] In summary, the existing multi-AUV collaborative pursuit control methods have limitations in dealing with the challenges brought by bearing-only observation and time-delay communication, and are difficult to meet the needs of practical applications. Therefore, it is urgent to propose a method that can jointly optimize the target positioning and collaborative pursuit decision in an end-to-end manner under the conditions of bearing-only observation and time-delay communication. SUMMARY
[0009] The technical problem to be solved by the present application is to provide a multi-AUV bearing-only perception collaborative pursuit control method based on reinforcement learning, which solves the problem of difficult end-to-end joint optimization of target position estimation and collaborative pursuit decision under the conditions of passive sonar bearing-only observation and distance-dependent time delay in underwater acoustic communication, improves the success rate and robustness of the pursuit, and also considers the obstacle avoidance safety in complex environments.
[0010] To achieve the above application purposes, the present application provides a multi-AUV bearing-only perception collaborative pursuit control method based on reinforcement learning, comprising the following steps:
[0011] S1. Training interaction and sample collection stage: The multiple autonomous underwater vehicles interact in the pursuit environment, and each autonomous underwater vehicle obtains its own state and passive sonar bearing-only observation of the target at the current time, and receives delayed communication information from the neighbor autonomous underwater vehicles, constructs the observation and state input for strategy learning, and stores the samples obtained through interaction into the experience replay pool; S2. Joint training phase: In the centralized training distributed execution (CTDE) framework, a parameter-shared multi-agent deep deterministic policy gradient network is constructed, wherein a time-aware deep geometric inversion module, an explicit belief guiding module, and an adaptive context gating fusion module are integrated in the multi-agent deep deterministic policy gradient network. By sampling sequence samples from the experience replay pool, the global value network and the policy network with target belief estimation are jointly optimized to obtain the trained network parameters. S3. Online execution phase: The trained network parameters are loaded, and each autonomous underwater vehicle generates an input for the policy network based on its own state, passive sonar bearing-only observation of the target, and delayed communication information with neighbor autonomous underwater vehicles, and generates a control instruction based on the action output by the policy network, thereby realizing cooperative hunting and configuration maintenance of multiple autonomous underwater vehicles.
[0012] According to one aspect of the present application, in step S1, the samples stored in the experience replay pool are in the form of six-tuples, wherein the sample is represented as: ; ; ; ; wherein, represents the global state of the autonomous underwater vehicle at the adjacent time, represents the local observation of the i-th autonomous underwater vehicle at time t, represents the set of local observations of all autonomous underwater vehicles at the adjacent time, represents the joint action between autonomous underwater vehicles, represents the action output by the policy network of the i-th autonomous underwater vehicle, represents the speed control instruction of the i-th autonomous underwater vehicle at time t, represents the angular velocity control instruction of the i-th autonomous underwater vehicle at time t, represents the number of autonomous underwater vehicles, represents the reward obtained by the autonomous underwater vehicle after executing the joint action .
[0013] According to one aspect of the invention, the local observations of each independent underwater vehicle at time t are obtained by concatenating the passive sonar pure azimuth observations of the target and the delayed communication vector containing an explicit time delay term, based on the vehicle's own state. The local observations of each independent underwater vehicle at time t are expressed as follows: ; ; in, Indicates the first Local observations of an autonomous underwater vehicle at time t. Indicates the first Kinematic information of an autonomous underwater vehicle, including position, velocity, and heading angle. Indicates the first The perception information of an autonomous underwater vehicle includes orientation observation of targets and local map observation of obstacles. Indicates the first Delayed communication information vectors received by an autonomous underwater vehicle Represents a set of neighbors. Indicates the heading angle. Indicates azimuth observation. Indicates the first The location of the autonomous underwater vehicle Indicates the first An autonomous underwater vehicle in Location at any given moment Indicates the first An autonomous underwater vehicle in The speed of time, Indicates the first The and the first Communication latency of an autonomous underwater vehicle Indicates the first The time it takes for an autonomous underwater vehicle to send information.
[0014] According to one aspect of the invention, in step S2, the time-aware depth geometry inversion module achieves target belief estimation by generating a relative position estimate of the target, comprising: The delay-aware feature embedding submodule is used to jointly encode the delayed self-kinematic information and target perception information of neighboring autonomous underwater vehicles and their corresponding communication delays into a delay-aware embedding vector. The cyclic memory submodule is used to perform cyclic reasoning on the observation history and delay-aware embedding vector within the time window and output the target belief features; The target location estimation module for auxiliary supervision generates a relative location estimate of the target based on the input target belief features and using auxiliary regression.
[0015] According to an aspect of the present application, in step S2, the step of jointly optimizing the global value network, the policy network with target belief estimation by sampling sequence samples from the experience replay pool to obtain the trained network parameters, comprises: S21. extracting a sequence sample of a preset length from the experience replay pool for training of the global value network and the policy network with target belief estimation, wherein the policy network is based on the time-aware deep geometric inversion module for target belief estimation; S22. using the extracted sequence sample to minimize the time difference error to update the network parameters of the global value network; S23. training the time-aware deep geometric inversion module separately using the extracted sequence sample, wherein the network parameters of the time-aware deep geometric inversion module are updated by minimizing the auxiliary regression loss; S24. inputting the relative position estimation of the target output by the time-aware deep geometric inversion module as the tracking belief feature into the policy network, and using a gradient blocking operator to avoid the policy gradient from destroying the positioning representation, to realize explicit belief guidance and gradient decoupling; S25. adaptively context-gated fusing the spatial features of the autonomous underwater vehicle itself and the tracking belief features of the target to generate a fusion vector for input of the policy network; S26. under the centralized training and decentralized execution (CTDE) framework, using the global value network to perform deterministic policy gradient update on the policy network, and introducing an auxiliary loss of a target position estimation module serving as a training aid as a weighted item to jointly optimize the policy network, wherein, under the condition of blocking the gradient return of the target belief, the tracking belief feature after gradient stop and the fusion vector jointly constitute a fusion state as input of the policy network; S27. soft updating the target network of the policy network and the global value network and repeating steps S21 to S26 until a preset training round is reached to obtain the trained network parameters and save them.
[0016] According to an aspect of the present application, in step S22, the step of using the extracted sequence sample to minimize the time difference error to update the network parameters of the global value network, wherein the global value network evaluates the global state and the joint action in the training stage, and updates the network parameters of the global value network by minimizing the Bellman error, and the loss function of the process of updating the network parameters of the global value network is represented as: ; wherein, represents the expected reward, represents the length of the training trajectory, represents the discount factor, represents the target value network, representing a value network, , is a global state feature available at the adjacent time for centralized training, is a joint action at the time generated by the target policy network, is a joint action at the next time generated by the target policy network. is a joint action at the next time generated by the target policy network.
[0017] According to one aspect of the present application, in the step of training the time-aware deep geometric inversion module separately using the extracted sequence samples in step S23, the step includes: delay-aware embedding vectors for delay compensation are constructed based on a delay-aware feature embedding sub-module, wherein the delay-aware embedding vectors are represented as: ; ; wherein, represents the delay-aware embedding vectors, represents the embedding network, represents a delay information segment from a neighbor autonomous underwater vehicle, represents a time delay position, represents a time delay velocity, represents a time delay target orientation angle; the observation history and the delay-aware embedding vectors within the time window are circularly inferred based on a circular memory sub-module, and a target belief feature is output; a target position estimation module based on auxiliary supervision obtains a relative position estimate of the target from the target belief feature, and updates and trains the network parameters of the time-aware deep geometric inversion module through auxiliary regression loss supervision; wherein an auxiliary regression loss function for representing the auxiliary regression loss is: ; wherein, represents the relative position estimate of the target, represents the true relative position of the target available in the training phase.
[0018] According to one aspect of the present application, in the step of avoiding policy gradient from destroying the positioning representation by using a gradient blocking operator in step S24, the tracking belief feature after stopping the gradient is represented as: ; wherein, represents the tracking belief feature after stopping the gradient, and the input state of the policy network is represents a vector concatenation symbol, represents the gradient blocking operator.
[0019] According to one aspect of the present application, in the step S25, the adaptive context gating fusion of the spatial features of the space where the autonomous underwater vehicle is located and the tracking belief features of the target is performed to generate a fusion vector for the input of the policy network, and the generated fusion vector is represented as: ; ; wherein, represents the fusion vector, represents the gating weight, and , represents the spatial features of the surrounding obstacles extracted by the local occupancy grid map convolutional network, and , represents the tracking belief features, and , represents the dimension of the spatial features, represents the dimension of the tracking features, represents the weight matrix of the gating network, represents the bias vector of the gating network, represents the splicing operator; In the step S26, under the centralized training distributed execution CTDE framework, the global value network is used to perform the deterministic policy gradient update of the policy network, and the auxiliary loss of the target position estimation module introduced for training is used as a weighted item to perform the joint optimization of the policy network, and the joint optimization target is represented as: ; wherein, represents the policy network, represents the auxiliary loss weight coefficient, represents the parameters of the policy network, represents the global state, represents the fusion state.
[0020] According to one aspect of the present application, the recurrent memory sub-module is constructed based on a recurrent network, wherein the recurrent network is an LSTM network, and the LSTM gating calculation is represented as: ; ; ; ; ; wherein, represents the output of the input gate, which determines how much new information should be updated to the long-time memory, represents the weight matrix of the input gate, a bias vector representing the input gate, a candidate memory state, a cell state at the current time, a weight matrix representing the candidate state, a bias vector representing the candidate state, an output of the forget gate, a weight matrix representing the output gate, a bias vector representing the output gate, a time series input fused with local and delayed information, a hidden state at the current time, is a Sigmoid function, is an element-wise multiplication.
[0021] The beneficial effects of the present scheme are as follows: According to one scheme of the present application, the scheme realizes time delay compensation type reconstruction of target position through time perception depth geometric inversion under the condition that the passive sonar only has pure bearing information, improves the positioning availability and accuracy. Through explicit belief guidance and gradient decoupling mechanism, the positioning representation is driven by geometric consistency loss, reduces the interference of policy gradient noise on belief representation, improves the training stability and generalization ability. Through adaptive context gate fusion, the safety of obstacle avoidance and the efficiency of encirclement are dynamically balanced in complex obstacle environment, and the robustness of cooperative encirclement is improved. The CTDE and parameter sharing training method are adopted to improve the sample efficiency and cooperative consistency, which is suitable for high time delay, partially observable underwater cooperative task.
[0022] According to one scheme of the present application, the scheme can realize end-to-end joint optimization of relative position estimation and cooperative hunting decision of the target under the condition that only passive sonar pure bearing observation and distance related time delay exist in underwater acoustic communication, improve the hunting success rate, robustness and generalization ability, and also consider the obstacle avoidance safety in complex environment.
[0023] According to one scheme of the present application, the scheme can improve the pure bearing positioning accuracy and the robustness of multi-AUV cooperative hunting under the condition of high time delay and partial observability, and has good engineering application value. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a step diagram of the multi-AUV pure bearing perception cooperative hunting control method based on reinforcement learning of the present application; Figure 2 is a flow principle diagram of the multi-AUV pure bearing perception cooperative hunting control method based on reinforcement learning of the present application. DETAILED DESCRIPTION
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the embodiments of the present application will be described in detail below.
[0026] The present application will be described in detail below in combination with the drawings and specific embodiments. The embodiments cannot be described one by one here, but the embodiments of the present application are not limited to the following embodiments.
[0027] In combination with Figure 1 and Figure 2 , according to an embodiment of the present application, a multi-AUV bearing-only perception cooperative hunting control method based on reinforcement learning of the present application comprises the following steps: S1. Training interaction and sample collection stage: The multi-autonomous underwater vehicles interact in the hunting environment. Each autonomous underwater vehicle obtains its own state and passive sonar bearing-only observation of the target at the current time, and receives delayed communication information from the neighbor autonomous underwater vehicle, constructs the observation and state input for policy learning, and stores the samples obtained by interaction to the experience replay pool; S2. Joint training stage: Under the centralized training distributed execution CTDE framework, a parameter-shared multi-agent deep deterministic policy gradient network is constructed; wherein the multi-agent deep deterministic policy gradient network integrates a time-aware deep geometric inversion module, an explicit belief guidance module and an adaptive context gating fusion module; By sampling sequence samples from the experience replay pool, the global value network and the policy network with target belief estimation are jointly optimized, and the trained network parameters are obtained; S3. Online execution stage: Load the network parameters obtained by training. Each autonomous underwater vehicle generates input for the policy network based on its own state, passive sonar bearing-only observation of the target, and delayed communication information from the neighbor autonomous underwater vehicle, and generates control instructions based on the action output by the policy network, to realize multi-autonomous underwater vehicle cooperative hunting and configuration maintenance.
[0028] In combination with Figure 1 and Figure 2 , according to an embodiment of the present application, in step S1, facing the scene of multi-autonomous underwater vehicle (AUV) cooperative hunting of the target, the hunting formation is composed of multi-autonomous underwater vehicles. At discrete time , the position, velocity and heading angle of the th autonomous underwater vehicle are , and , respectively. The target position is , R denotes the set of real numbers.
[0029] In this embodiment, passive sonar bearing-only observation: in the passive sonar condition, the first autonomous underwater vehicle only obtains the bearing observation of the target containing noise (no distance information), which can be expressed as: Wherein, denotes the velocity of the first autonomous underwater vehicle at time t, and is the observation noise; under the condition of two-dimensional horizontal plane, is the bearing angle observation function. Further, receive the delayed communication information from the neighbor autonomous underwater vehicle: since the autonomous underwater vehicle is in the water, it belongs to underwater acoustic communication, and further, the underwater acoustic communication propagation speed is low and there is distance-related propagation delay. Then, the communication delay between any two autonomous underwater vehicles (such as the first autonomous underwater vehicle and the second autonomous underwater vehicle) is
[0030] Given by the following formula: Wherein, is the sound speed in water.
[0031] Therefore, the message received by the first autonomous underwater vehicle from the second autonomous underwater vehicle at time t satisfies the delayed reception model: Wherein, denotes that the message received by the first autonomous underwater vehicle from the second autonomous underwater vehicle at time t satisfies the delayed reception model, denotes the time delay information of the message received by the first autonomous underwater vehicle from the second autonomous underwater vehicle at time t after adding the communication delay of the first autonomous underwater vehicle at time t. In this embodiment, in order to facilitate the description of the autonomous underwater vehicle obtaining its own state at the current time, the action of the autonomous underwater vehicle is defined, specifically, in order to meet the continuous control requirement of the pursuit maneuver, the continuous action space is adopted, that is, the first autonomous underwater vehicle outputs the action at time t:
[0032] Wherein, the velocity control command of the i-th autonomous underwater vehicle at time t, the velocity control command of the i-th autonomous underwater vehicle at time t, the angular velocity control command of the i-th autonomous underwater vehicle at time t. the angular velocity control command of the i-th autonomous underwater vehicle at time t.
[0033] In the embodiment, each autonomous underwater vehicle performs perception and decision in a distributed manner, and its available information is composed of two parts: local real-time observation and global time delay information; wherein the local real-time observation is expressed by constructing a local observation vector; specifically, the i-th autonomous underwater vehicle constructs a local observation vector at time t which at least includes its own state, passive sonar bearing-only observation, and delay communication vector representing delay communication information, so that the local observation vector is expressed as: ; ; ; wherein, represents the local observation of the i-th autonomous underwater vehicle at time t, represents the kinematics information of the i-th autonomous underwater vehicle, including position, velocity and heading angle, represents the perception information of the i-th autonomous underwater vehicle, including bearing observation (such as bearing angle observation ) of the target and local map observation (such as local map / occupancy grid ) of the obstacle, represents the delay communication information vector received by the i-th autonomous underwater vehicle, represents the neighbor set, represents the heading angle, represents the bearing angle observation, represents the position of the i-th autonomous underwater vehicle, represents the position of the i-th autonomous underwater vehicle at time t, represents the velocity of the i-th autonomous underwater vehicle at time t, represents the communication delay between the i-th and j-th autonomous underwater vehicles, represents the position of the i-th autonomous underwater vehicle, represents the position of the i-th autonomous underwater vehicle at time t, represents the velocity of the i-th autonomous underwater vehicle at time t, represents the communication delay between the i-th and j-th autonomous underwater vehicles, represents the position of the i-th autonomous underwater vehicle, represents the velocity of the i-th autonomous underwater vehicle at time t, represents the communication delay between the i-th and j-th autonomous underwater vehicles, represents the position of the i-th autonomous underwater vehicle, represents the velocity of the i-th autonomous underwater vehicle at time t, represents the communication delay between the i-th and j-th autonomous underwater vehicles, represents the position of the i-th autonomous underwater vehicle, represents the velocity of the i-th autonomous underwater vehicle at time t, The time at which the individual autonomous underwater vehicle sends information.
[0034] In combination Figure 1 And Figure 2 As shown in FIG. 1, according to an embodiment of the present application, each autonomous underwater vehicle is trained in a training phase to learn a policy network based on the output of the current policy network Or its historical sequence as input And executes, and obtains a reward after execution And the next observation Thus, in step S1, the sample stored in the experience replay pool is in the form of a six-tuple, where the sample is represented as: ; ; ; ; Wherein, Represents the global state of the autonomous underwater vehicle at the adjacent time, Represents the local observation of the t-th autonomous underwater vehicle at time t, Represents the set of local observations of all autonomous underwater vehicles at the adjacent time, Represents the joint action between autonomous underwater vehicles, Represents the action output by the policy network of the t-th autonomous underwater vehicle, Represents the action output by the policy network of the t-th autonomous underwater vehicle, Represents the velocity control command of the t-th autonomous underwater vehicle at time t, Represents the angular velocity control command of the t-th autonomous underwater vehicle at time t, Represents the number of autonomous underwater vehicles, Represents the reward obtained by the autonomous underwater vehicle after executing the joint action . According to an embodiment of the present application, the local observation of each autonomous underwater vehicle at time t is obtained by splicing the state of itself, the passive sonar bearing-only observation of the target, and the delayed communication vector containing an explicit time delay term, where the local observation of each autonomous underwater vehicle at time t is represented as:
[0035] ; ; ; Wherein, Represents the local observation of the t-th autonomous underwater vehicle at time t, Represents the local observation of the t-th autonomous underwater vehicle at time t, Represents the local observation of the t-th autonomous underwater vehicle at time t, kinematic information of the i-th autonomous underwater vehicle, including position, velocity and heading angle, perception information of the i-th autonomous underwater vehicle, including bearing observation of targets and local map observation of obstacles, perception information of the i-th autonomous underwater vehicle, including bearing observation of targets and local map observation of obstacles, delayed communication information vector received by the i-th autonomous underwater vehicle, neighbor set, heading angle, bearing observation, position of the i-th autonomous underwater vehicle, position of the i-th autonomous underwater vehicle, position of the i-th autonomous underwater vehicle at time t, velocity of the i-th autonomous underwater vehicle at time t, velocity of the i-th autonomous underwater vehicle at time t, communication delay between the i-th and j-th autonomous underwater vehicle, time at which the i-th autonomous underwater vehicle sends information. According to an embodiment of the present application, in step S2, the time-aware deep geometry inversion module realizes target belief estimation by generating relative position estimation of the target, which includes: a delay-aware feature embedding submodule, configured to jointly encode the delayed self-kinematic information and target perception information of the neighbor autonomous underwater vehicle with the corresponding communication delay as a delay-aware embedding vector; in this embodiment, the delay-aware feature embedding submodule is constructed based on a multi-layer perception (MLP); a recurrent memory submodule, configured to perform recurrent inference on the observation history and the delay-aware embedding vector within a time window, and output target belief features; an auxiliary-supervised target position estimation module, configured to generate relative position estimation of the target based on the input target belief features and in an auxiliary regression manner; in this embodiment, the target belief features are used to estimate the relative position of the target based on a geometric regression output layer, and an auxiliary supervision structure with the same topology as the geometric regression output layer is used to supervise the estimation process in an auxiliary regression manner. In this embodiment, the time-aware deep geometry inversion module obtains the relative position estimation or tracking belief features of the target by fusing the passive sonar bearing-only observation and the time delay information, which can be used for subsequent strategy decision-making.
[0036] According to an embodiment of the present application, in step S2, the time-aware deep geometry inversion module realizes target belief estimation by generating relative position estimation of the target, which includes: a delay-aware feature embedding submodule, configured to jointly encode the delayed self-kinematic information and target perception information of the neighbor autonomous underwater vehicle with the corresponding communication delay as a delay-aware embedding vector; in this embodiment, the delay-aware feature embedding submodule is constructed based on a multi-layer perception (MLP); a recurrent memory submodule, configured to perform recurrent inference on the observation history and the delay-aware embedding vector within a time window, and output target belief features; an auxiliary-supervised target position estimation module, configured to generate relative position estimation of the target based on the input target belief features and in an auxiliary regression manner; in this embodiment, the target belief features are used to estimate the relative position of the target based on a geometric regression output layer, and an auxiliary supervision structure with the same topology as the geometric regression output layer is used to supervise the estimation process in an auxiliary regression manner.
[0037] In this embodiment, the time-aware deep geometry inversion module obtains the relative position estimation or tracking belief features of the target by fusing the passive sonar bearing-only observation and the time delay information, which can be used for subsequent strategy decision-making.
[0038] According to an embodiment of the present application, the recurrent memory sub-module is constructed based on a recurrent network, wherein the recurrent network is an LSTM network, and wherein the LSTM gating computation is represented as: ; ; ; ; ; wherein, represents the output of the input gate, which determines how much new information should be updated into the long-term memory, represents the weight matrix of the input gate, represents the bias vector of the input gate, represents the candidate memory state, represents the cell state at the current time, represents the weight matrix of the candidate state, represents the bias vector of the candidate state, represents the output of the forget gate, represents the weight matrix of the output gate, represents the bias vector of the output gate, represents the time-series input that fuses local and delayed information, represents the hidden state at the current time, is a Sigmoid function, is an element-wise multiplication.
[0039] According to an embodiment of the present application, in step S2, the step of jointly optimizing the global value network, the policy network with target belief estimation, and obtaining the trained network parameters, comprises: S21. extracting a sequence sample of a preset length from the experience replay pool for training of the global value network and the policy network with target belief estimation, wherein the policy network is based on the time-aware deep geometric inversion module for target belief estimation; S22. using the extracted sequence sample to minimize the time difference error to update the network parameters of the global value network; S23. training the time-aware deep geometric inversion module separately using the extracted sequence sample, wherein the time-aware deep geometric inversion module network parameters are updated by minimizing the auxiliary regression loss; S24. The relative position estimation of the target output by the Time-Aware Deep Geometric Inversion (TDGI) module is input into the policy network as a tracking belief feature, and a gradient blocking operator is used to avoid the policy gradient from destroying the positioning representation, so as to realize explicit belief guidance and gradient decoupling; S25. Adaptive Gated Context Fusion (AGCF) is performed on the spatial features in which the autonomous underwater vehicle is located and the tracking belief features of the target, to generate a fusion vector for input into the policy network; S26. Under the centralized training for decentralized execution (CTDE) framework, the policy network is updated by the global value network using a deterministic policy gradient, and a loss of a target position estimation module used for training auxiliary supervision is introduced as a weighted item for joint optimization of the policy network, wherein, under the condition of blocking the return of the target belief gradient, the tracking belief feature after gradient stop and the fusion vector are combined to form a fusion state, which is input into the policy network. S27. The target network of the policy network and the global value network is soft-updated, and steps S21 to S26 are repeated until a preset training round is reached, to obtain trained network parameters and save them.
[0040] According to an embodiment of the present application, in step S22, in the step of updating the network parameters of the global value network by minimizing the time difference error using the extracted sequence samples, the global value network evaluates the global state and the joint action in the training phase, and updates the network parameters of the global value network by minimizing the Bellman error, wherein the loss function of the process of updating the network parameters of the global value network is represented as: ; wherein, represents the expected reward, represents the length of the training trajectory, represents the discount factor, represents the target value network, represents the value network, is the global state feature available during centralized training, is the joint action generated by the target policy network at the next time.
[0041] According to an embodiment of the present application, in step S23, in the step of training the Time-Aware Deep Geometric Inversion module separately using the extracted sequence samples, the step includes: The delay-aware embedding sub-module based on the delay perception feature constructs a delay-aware embedding vector for delay compensation, wherein the delay-aware embedding vector is represented as: ; ; wherein, represents the delay-aware embedding vector, represents the embedding network, represents the delay information segment from the neighbor autonomous underwater vehicle, represents the time delay position, represents the time delay velocity, represents the time delay target orientation angle. The recurrent memory sub-module based on the observation history and the delay-aware embedding vector within the time window performs recurrent inference, and outputs a target belief feature; The target position estimation module based on auxiliary supervision obtains a relative position estimate of the target from the target belief feature, and completes the update training of the network parameters of the time-aware deep geometric inversion module through auxiliary regression loss supervision training. The auxiliary regression loss function for representing the auxiliary regression loss is: ; wherein, represents the relative position estimate of the target, represents the real relative position of the target available in the training stage.
[0042] According to an embodiment of the present application, in the step of stopping the gradient in the step of adopting the gradient blocking operator to avoid the strategy gradient from destroying the positioning representation in step S24, the tracking belief feature after stopping the gradient is represented as: ; wherein, represents the tracking belief feature after stopping the gradient, and is the input state of the strategy network, represents a vector concatenation symbol, represents the gradient blocking operator.
[0043] According to an embodiment of the present application, in the step of generating a fusion vector for the input of the strategy network by adaptively fusing the space features in which the autonomous underwater vehicle is located and the tracking belief feature of the target in step S25, the generated fusion vector is represented as: ; ; wherein, represents the fusion vector, represents the gating weight, and , denotes the spatial feature of the surrounding obstacles extracted by the local occupancy grid map convolutional network, and , denotes the tracking belief feature, and , denotes the dimension of the spatial feature, denotes the dimension of the tracking feature, denotes the weight matrix of the gating network, denotes the bias vector of the gating network, denotes the concatenation operator.
[0044] According to an embodiment of the present application, in step S26, in the centralized training distributed execution CTDE framework, the global value network is used to perform a deterministic policy gradient update on the policy network, and an auxiliary loss of a target position estimation module for training auxiliary supervision is introduced as a weighted term to jointly optimize the policy network. The joint optimization target is expressed as: ; wherein, denotes the policy network, denotes the auxiliary loss weight coefficient, denotes the parameters of the policy network, denotes the centralized value network, denotes the global state, denotes a fusion state composed of the tracking belief feature after gradient stop and the fusion vector.
[0045] Based on the above settings, the entire training phase can be described as the following loop: initialize the policy network, the value network, the target network thereof, and the experience replay pool; interact with the surrounding environment to collect samples and store them; when the replay pool samples meet the update condition, extract sequence samples with a length of , update the value network in turn to minimize the loss function , update the network parameters of the time-aware deep geometric inversion module to minimize the auxiliary regression loss function , and then update the policy network to minimize the joint optimization target ; finally, soft update the target network and repeat the above process until the preset training round is reached, and save the network parameters.
[0046] According to an embodiment of the present application, in step S3, the network parameters obtained by training are loaded, each autonomous underwater vehicle generates an input for the policy network based on its own state, passive sonar bearing-only observation of the target, and delayed communication information with neighbor autonomous underwater vehicles, and generates a control instruction based on the action output by the policy network, to realize the steps of cooperative hunting and configuration keeping of the multiple autonomous underwater vehicles, comprising: S31. Obtain the current time's self-state of the autonomous underwater vehicle, the passive sonar bearing-only observation of the target, and the delayed communication information with the neighbor autonomous underwater vehicle, and construct the local observation at the current time ; S32. Output the relative position estimation of the target through the time-aware deep geometric inversion module in the target belief estimation network or track belief feature ; S33. Perform explicit belief guidance through the explicit belief guidance module in the target belief estimation network, and generate a fusion vector based on the adaptive context gating fusion module The policy network outputs actions based on the fusion state composed of the gradient-stopped tracking belief feature and the fusion vector ; ; S34. Repeat steps S31 to S33 until the task is completed, to realize the multi-AUV surrounding configuration and the target surrounding.
[0047] According to the scheme of the application, in the underwater cooperative hunting scene facing the passive sonar which only provides noisy bearing information and the underwater acoustic communication which has distance-related propagation delay, the target positioning and cooperative decision are optimized end-to-end under the multi-agent deep reinforcement learning framework: the relative time delay information is introduced through the time-aware deep geometric inversion module, the delayed bearing observations from the current ship and the teammates are motion compensated and fused, the reconstruction estimation of the target position is realized, and the auxiliary learning task is used to improve the feature expression and training stability; further, an explicit belief guidance mechanism is designed, the target tracking belief is embedded into the policy network, and an adaptive context gating fusion structure is adopted, the obstacle space features and the time sequence tracking belief are dynamically weighted according to the environmental complexity, so as to balance the safety obstacle avoidance and the hunting efficiency. When executed online, each AUV generates speed and heading control instructions according to the local observation and the received delayed information, forms a stable cooperative hunting configuration, and improves the capture success rate.
[0048] The above is only an example of the specific scheme of the application, and for the devices and structures not described in detail, it should be understood that the general devices and general methods in the art are used to implement them.
[0049] The above only describes one scheme of the application and is not used to limit the application. For those skilled in the art, the application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.
Claims
1. A multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning, characterized in that, Includes the following steps: S1. Training Interaction and Sample Collection Phase: Multiple autonomous underwater vehicles (AUVs) interact in a containment environment. Each AUV acquires its own state and passive sonar azimuth observations of the target at the current moment, and receives delayed communication information from neighboring AUVs. They construct observation and state inputs for policy learning and store the samples obtained from the interaction in an experience replay pool. S2. Joint Training Phase: Under the centralized training and distributed execution CTDE framework, a parameter-sharing multi-agent deep deterministic policy gradient network is constructed. The multi-agent deep deterministic policy gradient network integrates a time-aware deep geometric inversion module, an explicit belief guidance module, and an adaptive context-gated fusion module. By sampling sequence samples through the experience replay pool, the global value network and the policy network with target belief estimation are jointly optimized to obtain the trained network parameters. S3. Online Execution Phase: By loading the network parameters obtained from training, each autonomous underwater vehicle generates inputs for the policy network based on its own state, passive sonar azimuth observation of the target, and delayed communication information with neighboring autonomous underwater vehicles. Based on the actions output by the policy network, control commands are generated to achieve coordinated encirclement and configuration maintenance by multiple autonomous underwater vehicles.
2. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 1, characterized in that, In step S1, the samples stored in the experience replay pool are constructed in the form of six-tuples, where each sample is represented as: in, This represents the global state of an autonomous underwater vehicle at adjacent moments. Indicates the first Local observations of an autonomous underwater vehicle at time t. This represents a local set of observations from all autonomous underwater vehicles at adjacent moments. This indicates coordinated actions between autonomous underwater vehicles. Indicates the first The actions output by the policy network of an autonomous underwater vehicle Indicates the first Speed control command for an autonomous underwater vehicle at time t Indicates the first Angular velocity control command of an autonomous underwater vehicle at time t Indicates the number of autonomous underwater vehicles. This indicates that the autonomous underwater vehicle is performing joint operations. They then received a reward.
3. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 2, characterized in that, The local observations of each independent underwater vehicle at time t are obtained by concatenating the target's passive sonar pure azimuth observations and the delayed communication vector containing explicit time delay terms, based on the vehicle's own state. The local observations of each independent underwater vehicle at time t are represented as follows: in, Indicates the first Local observations of an autonomous underwater vehicle at time t. Indicates the first The kinematic information of an autonomous underwater vehicle, including position, velocity, and heading angle. Indicates the first The perception information of an autonomous underwater vehicle includes orientation observation of targets and local map observation of obstacles. Indicates the first Delayed communication information vectors received by an autonomous underwater vehicle Represents a set of neighbors. Indicates the heading angle. Indicates azimuth observation. Indicates the first The location of the autonomous underwater vehicle Indicates the first An autonomous underwater vehicle in Location at any given moment Indicates the first An autonomous underwater vehicle in The speed of time Indicates the first The and the first Communication latency of an autonomous underwater vehicle Indicates the first The time it takes for an autonomous underwater vehicle to transmit information.
4. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to any one of claims 1 to 3, characterized in that, In step S2, the time-aware depth geometry inversion module generates a relative position estimate of the target to achieve target belief estimation, which includes: The delay-aware feature embedding submodule is used to jointly encode the delayed self-kinematic information and target perception information of neighboring autonomous underwater vehicles and their corresponding communication delays into a delay-aware embedding vector. The cyclic memory submodule is used to perform cyclic reasoning on the observation history and delay-aware embedding vector within the time window and output the target belief features; The target location estimation module for auxiliary supervision generates a relative location estimate of the target based on the input target belief features and using auxiliary regression.
5. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 4, characterized in that, In step S2, the step of sampling sequence samples through the experience replay pool and jointly optimizing the global value network and the policy network with target belief estimation to obtain the trained network parameters includes: S21. Extract sequence samples of a preset length from the experience replay pool for training the global value network and the policy network with target belief estimation, wherein the policy network is used for target belief estimation based on the time-aware deep geometric inversion module. S22. Use extracted sequence samples to minimize temporal difference error to update the network parameters of the global value network; S23. The time-aware deep geometry inversion module is trained separately using the extracted sequence samples, wherein the auxiliary regression loss is minimized to update the network parameters of the time-aware deep geometry inversion module. S24. The relative position estimate of the target output by the time-aware depth geometry inversion module is used as the input of the tracking belief feature into the policy network, and the gradient blocking operator is used to avoid the policy gradient from destroying the localization representation, thereby realizing explicit belief guidance and gradient decoupling. S25. Adaptive context-gated fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target is performed to generate a fusion vector for input to the policy network; S26. Under the centralized training and distributed execution CTDE framework, the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target position estimation module under training assistance supervision is introduced as a weighting term to jointly optimize the policy network. Under the condition of blocking the backpropagation of target belief gradient, the tracking belief features after stopping the gradient and the fusion vector are used together to form the fusion state as the input of the policy network. S27. Perform soft updates on the target network of the policy network and the global value network, and repeat steps S21 to S26 until the preset training rounds are reached, obtain the trained network parameters, and save them.
6. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 5, characterized in that, In step S22, the step of updating the network parameters of the global value network by minimizing the temporal difference error using the extracted sequence samples, the global value network evaluates the global state and joint actions during the training phase and updates the network parameters of the global value network by minimizing the Bellman error. The loss function for updating the network parameters of the global value network is expressed as: in, Indicates the expectation of a reward. Indicates the length of the training trajectory. Indicates the discount factor. Represents the target value network. Represents a value network. , These are global state features available at adjacent time steps during intensive training. For generated by the target policy network Always in coordination The next-time joint action is generated by the target policy network.
7. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 6, characterized in that, Step S23, which involves training the time-aware depth geometric inversion module separately using the extracted sequence samples, includes: A delay-aware embedding vector for latency compensation is constructed based on the delay-aware feature embedding submodule, where the delay-aware embedding vector is represented as: in, Represents a delay-aware embedding vector. Indicates embedded network, This represents a segment of delayed information from a neighboring autonomous underwater vehicle. Indicates the time delay position. Indicates the delay speed. Indicates the target orientation angle during time delay; Based on the recurrent memory submodule, the observation history and delay-aware embedding vector within the time window are used for recurrent reasoning to output target belief features; The supervised target location estimation module obtains the relative location estimate of the target from the target belief features and completes the update training of the network parameters of the time-aware deep geometric inversion module through supervised training using auxiliary regression loss; wherein, the auxiliary regression loss function used to represent the auxiliary regression loss is: in, This represents an estimate of the target's relative position. This indicates the true relative position of the target that can be obtained during the training phase.
8. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 7, characterized in that, In step S24, the tracking belief feature after stopping the gradient is represented as follows: (The step involves using the gradient blocking operator to avoid gradient destruction in the localization representation.) in, This represents the tracking belief feature after gradient cessation, and is the input state of the policy network. Indicates vector concatenation symbol. This represents the gradient blocking operator.
9. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 8, characterized in that, In step S25, the step of adaptively gating the fusion of the spatial characteristics of the autonomous underwater vehicle and the tracking belief characteristics of the target to generate a fusion vector for input to the policy network is represented as follows: in, Represents the fusion vector. Indicates the gating weight, and , This represents the spatial features of surrounding obstacles extracted from a locally occupied raster map via a convolutional network, and , Indicates the characteristics of tracking beliefs, and , Dimensions representing spatial features The dimension representing the tracking feature. This represents the weight matrix of the gated network. This represents the bias vector of the gated network. Indicates the concatenation operator; In step S26, under the centralized training and distributed execution CTDE framework, the global value network is used to perform deterministic policy gradient updates on the policy network, and the auxiliary loss of the target location estimation module under training supervision is introduced as a weighting term to jointly optimize the policy network. The joint optimization objective is expressed as: in, Represents the policy network, This represents the auxiliary loss weighting coefficient. The parameters represent the policy network. Represents the global state. This indicates a fusion state.
10. The multi-AUV pure orientation perception cooperative encirclement and control method based on reinforcement learning according to claim 4, characterized in that, The recurrent memory submodule is constructed based on a recurrent network, wherein the recurrent network is an LSTM network, and the LSTM gated computation is represented as follows: in, The output of the input gate determines how much new information should be updated in long-term memory. This represents the weight matrix of the input gate. This represents the bias vector of the input gate. Indicates candidate memory states, This indicates the current state of the cell. The weight matrix representing the candidate states. The bias vector represents the candidate state. This represents the output of the forget gate. This represents the weight matrix of the output gate. This represents the bias vector of the output gate. This indicates a timing input that incorporates both local and delay information. This indicates the hidden state at the current moment. For the Sigmoid function, This is element-wise multiplication.
Citation Information
Patent Citations
Underwater multi-agent cooperative hunting method and device based on deep reinforcement learning
CN118153431A
Multi-unmanned aerial vehicle formation control method based on offline sample correction reinforcement learning
CN119620782A
UUV autonomous collision avoidance and navigation method and system based on improved ICM-DDQN
CN120143804A
Underwater multi-AUV attack and defense game method based on deep reinforcement learning under imperfect information
CN120525370A
MADDPG-based multi-AUV (Autonomous Underwater Vehicle) cooperative hunting method
CN120562748A