Multi-AUV (Autonomous Underwater Vehicle) cooperative sensing method based on multi-agent group relative strategy optimization

By constructing a multi-AUV collaborative perception system model and a multi-agent group relative strategy optimization algorithm, the problem of low data acquisition efficiency in multi-AUV collaborative perception tasks in underwater Internet of Things is solved, and the age optimization of error information under energy consumption and delay constraints is achieved, and the system performance is improved.

CN120491675APending Publication Date: 2025-08-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510490697.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the underwater Internet of Things, the synergistic effects between AUVs are not effectively utilized in the multi-AUV synergistic perception task, resulting in limited data acquisition efficiency. The traditional methods fail to effectively solve the problem of error information age (AOII), which has high calculation cost and is difficult to optimize under the constraints of energy consumption, delay and data coverage.

Method used

A multi-AUV collaborative perception system model is built, combining underwater communication, noise model and AUV motion constraints, and a multi-agent group relative strategy optimization algorithm is adopted. Through reinforcement learning, AUV agents are trained, and the perception strategy is optimized in a coordinated manner, and the punishment mechanism and KL divergence loss function are introduced to achieve a global optimal strategy.

Benefits of technology

It effectively reduces the age of error information collected by data, improves system performance, optimizes energy and time consumption, improves data coverage, and realizes stable and efficient collaborative perception of multi-AUV systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491675A_ABST
    Figure CN120491675A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of strategy optimization. The invention provides a multi-AUV (Autonomous Underwater Vehicle) cooperative sensing method based on multi-agent group relative strategy optimization. According to the embodiment of the invention, factors such as underwater communication, noise and the like and an environmental constraint and punishment mechanism are integrated to construct a multi-AUV collaborative sensing system model; a target function is constructed by taking reduction of data acquisition error information age as a core and combining with related factors; secondly, considering underwater environment and AUV information acquisition limitation, and converting an optimization problem into a partially observable Markov decision process; and finally, reinforcement learning training is performed on the AUV intelligent agent by using a multi-agent group relative strategy optimization algorithm, actions are coordinated, perception strategy optimization is realized, and system performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the technical field of strategy optimization, and in particular to a multi-AUV collaborative perception method based on relative strategy optimization of a multi-agent group. Background Art

[0002] In recent years, with the advent of the Smart Ocean concept, the Internet of Underwater Things (IoUT) has been widely applied in fields such as ocean exploration, environmental monitoring, and resource exploration. IoUT devices can be deployed on the seabed for long-term data collection. However, due to the complexity of the underwater environment, traditional electromagnetic wave communications have an extremely short propagation distance in water and are easily affected by factors such as tides and noise, resulting in communication delays or even interruptions. In contrast, while underwater acoustic communications can achieve long-distance data transmission, they still face technical challenges such as limited bandwidth, multipath interference, and the Doppler effect, which seriously affect the reliability and real-time performance of data transmission.

[0003] To improve the efficiency and reliability of underwater data collection, autonomous underwater vehicles (AUVs) are widely used to assist in data transmission between IoUT devices. AUV path planning is a core issue in optimizing data collection. Traditional approaches typically employ fixed paths or path optimization strategies based on environmental awareness to reduce energy consumption and improve data coverage. However, existing research primarily focuses on the age of information (AoI) to measure data freshness, but ignores the impact of high bit error rates and data loss in underwater communications. This can lead to the accumulation of large amounts of erroneous data and reduce the reliability of system decisions.

[0004] To address this issue, researchers have proposed the concept of Age of Incorrect Information (AOII), a metric that not only considers the timeliness of data but also comprehensively evaluates its correctness. However, in underwater multi-AUV collaborative perception tasks, minimizing AOII is a typical NP-hard problem. Its computational cost rises sharply with the increase in the number of AUVs and the complexity of environmental changes. In addition, most existing methods rely on single-agent optimization strategies and fail to fully utilize the synergy between AUVs, resulting in limited data collection efficiency. Therefore, how to minimize AOII using multi-agent optimization methods under constraints such as energy consumption, latency, and data coverage remains a major challenge facing the current underwater Internet of Things field.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the Invention

[0007] The purpose of the embodiments of the present disclosure is to provide a multi-AUV collaborative perception method based on multi-agent group relative strategy optimization, thereby overcoming one or more problems caused by the limitations and defects of related technologies to at least a certain extent.

[0008] According to an embodiment of the present disclosure, a multi-AUV collaborative perception method based on multi-agent group relative strategy optimization is provided, the method comprising: Construct a multi-AUV collaborative perception system model, which includes the deployment relationship between underwater IoT devices, multiple AUVs, and ground stations, and defines the underwater communication model, AUV data acquisition model, and AUV motion constraints. Based on the multi-AUV collaborative perception system model, under the constraints of energy consumption and communication bandwidth, the objective function is constructed with the goal of minimizing the age of erroneous information in data collection. The objective function is converted into a partially observable Markov decision process, defining a state space, an action space, and an individual AUV reward; wherein the state space includes the AUV's position, velocity, system data coverage, energy consumption, error information age, and penalty terms, and the action space includes the AUV's three-dimensional movement direction and velocity adjustment; Based on the state space, action space and individual AUV rewards, the multi-agent group relative strategy optimization algorithm is used to collaboratively optimize the perception strategies of multiple AUVs to obtain the optimal perception strategy.

[0009] Furthermore, the underwater communication model includes: Underwater noise models consisting of fluid noise, ship noise, wave noise, and thermal noise, as well as line-of-sight and non-line-of-sight communications between underwater IoT devices and multiple AUVs; The shortest path distance for non-line-of-sight communication is calculated based on the water depth and the height of the equipment and AUV above the seabed.

[0010] Furthermore, the AUV data acquisition model is determined based on the perception radius of the AUV and the spatial distribution of underwater IoT devices. The AUV data acquisition model includes an error information age model, an energy consumption model, and a total penalty model.

[0011] Furthermore, AUV motion constraints include AUV movement distance limits, obstacle avoidance rules, and boundary constraints.

[0012] Furthermore, the objective function is expressed as:

[0013] in, The age of misinformation for all underwater IoT devices, is the data collection coverage, For the minimum data collection coverage allowed, is the average energy consumption, is the maximum energy consumption allowed, Indicates time consumption, is the maximum time allowed, For AUV exist The location at the moment, For AUV exist The location at the moment, It is the boundary range of AUV movement.

[0014] Furthermore, the state space includes the state of the AUV and the state of the multi-AUV collaborative perception system model. The state of the AUV includes the position and speed of the AUV, and the state of the multi-AUV collaborative perception system model includes data collection coverage, energy consumption, error information age, time consumption and penalty. The action space includes the three-dimensional movement direction and speed adjustment of the AUV.

[0015] Furthermore, based on the state space, action space and individual AUV rewards, the multi-agent group relative strategy optimization algorithm is used to collaboratively optimize the perception strategies of multiple AUVs to obtain the optimal perception strategy, including the following steps: Calculate the relative advantage of each agent in the group based on the difference between the individual AUV reward and the group average reward; Based on the relative strength of each agent in the population, construct the policy loss: Based on the relative strength of each agent in the population, construct the policy loss:

[0016] in, For the current strategy Next, in state Take action probability; For reference strategy Next, in state Take action probability; It is the cutoff threshold, which is used to limit the amplitude of strategy update during strategy optimization; is the relative advantage of each agent in the group, is the expected value of the multiple sampling process, which represents the expected reward under a certain strategy; Introduce KL divergence loss to obtain KL divergence loss:

[0017] in, For reference strategy With the current strategy The KL divergence between is the regularization hyperparameter; According to the policy loss and KL divergence loss, the total loss is obtained:

[0018] The total loss is used to optimize the perception strategies of multiple AUVs to achieve global optimal behavior and obtain the optimal perception strategy.

[0019] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: In the embodiments disclosed herein, the multi-AUV collaborative perception method based on multi-agent group relative strategy optimization is implemented. First, a multi-AUV collaborative perception system model is constructed by integrating factors such as underwater communication and noise, as well as environmental constraints and penalty mechanisms. The objective function is constructed by combining relevant factors with the core goal of reducing the age of erroneous information in collected data. The optimization problem is then transformed into a partially observable Markov decision process, taking into account the underwater environment and the limitations of AUV information acquisition. Finally, the multi-agent group relative strategy optimization algorithm is used to conduct reinforcement learning training on the AUV agents, coordinate actions, optimize perception strategies, and improve system performance. Furthermore, the present application constructs an underwater communication model, an AUV data acquisition model, and AUV motion constraints, comprehensively considering practical factors such as line-of-sight and non-line-of-sight communication, obstacles, and boundary restrictions. A penalty mechanism is also introduced to ensure AUV operation, providing a solid foundation for optimized decision-making. The present application adopts a multi-agent group relative strategy optimization method combined with reinforcement learning to train AUV agents. The agents make adaptive decisions based on state information, and the group computing module coordinates actions, solving the problems of traditional methods and achieving a global optimal strategy. This application introduces relative advantage to incentivize agents to compete and cooperate, and designs a combined loss function that includes policy loss and KL divergence loss. This ensures stable policy updates, prevents overfitting, and provides a stable and efficient training method for multi-agent reinforcement learning, continuously improving system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0021] Figure 1A diagram showing the steps of a multi-AUV collaborative perception method based on relative strategy optimization of a multi-agent group in an exemplary embodiment of the present disclosure is shown; Figure 2 A framework diagram of a multi-AUV collaborative perception method based on relative strategy optimization of a multi-agent group in an exemplary embodiment of the present disclosure is shown; Figure 3 FIG. 4 shows the reward convergence during the training phase in an exemplary embodiment of the present disclosure; Figure 4 The AOII experimental results of the error information age in the exemplary embodiment of the present disclosure are shown; Figure 5 Showing the energy consumption in an exemplary embodiment of the present disclosure; Figure 6 Showing the time consumption in the exemplary embodiment of the present disclosure; Figure 7 Showing the performance of the present application and other methods in the exemplary embodiments of the present disclosure in terms of the age index of the error information; Figure 8 Showing the performance of the present application and other methods in data collection coverage in exemplary embodiments of the present disclosure; Figure 9 Showing the performance of the present application and other methods in terms of energy consumption in exemplary embodiments of the present disclosure; Figure 10 The time-consuming performance of the present application and other methods in the exemplary embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0023] In addition, the accompanying drawings are merely schematic illustrations of embodiments of the present disclosure and are not necessarily drawn to scale. Like reference numerals in the figures represent like or similar parts, and thus repeated descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically separate entities.

[0024] This example embodiment provides a multi-AUV collaborative perception method based on multi-agent group relative strategy optimization. Figure 1 As shown in , the multi-AUV collaborative perception method based on multi-agent group relative strategy optimization may include: Step S101: Construct a multi-AUV collaborative perception system model, which includes the deployment relationship between underwater IoT devices, multiple AUVs and ground stations, and defines an underwater communication model, an AUV data acquisition model and AUV motion constraints; Step S102: Based on the multi-AUV collaborative perception system model, under the constraints of energy consumption and communication bandwidth, an objective function is constructed with the goal of minimizing the age of erroneous information in data collection; Step S103: Convert the objective function into a partially observable Markov decision process, and define the state space, action space, and individual AUV reward; wherein the state space includes the AUV's position, speed, system data coverage, energy consumption, error information age, and penalty terms, and the action space includes the AUV's three-dimensional movement direction and speed adjustment; Step S104: Based on the state space, action space and individual AUV rewards, the multi-agent group relative strategy optimization algorithm is used to collaboratively optimize the perception strategies of multiple AUVs to obtain the optimal perception strategy.

[0025] Through the above-mentioned multi-AUV collaborative perception method based on multi-agent group relative strategy optimization, on the one hand, firstly, factors such as underwater communication, noise, as well as environmental constraints and penalty mechanisms are integrated to construct a multi-AUV collaborative perception system model; with the core of reducing the age of erroneous information in collected data, the objective function is constructed in combination with relevant factors; then, considering the underwater environment and the limitations of AUV information acquisition, the optimization problem is transformed into a partially observable Markov decision process; finally, the multi-agent group relative strategy optimization algorithm is used to conduct reinforcement learning training on the AUV agents, coordinate actions, achieve perception strategy optimization, and improve system performance. On the other hand, this application constructs an underwater communication model, an AUV data acquisition model, and an AUV motion constraint, comprehensively considering actual factors such as line-of-sight and non-line-of-sight communication, obstacles, boundary restrictions, etc., and also introduces a penalty mechanism to ensure AUV operation, providing a solid foundation for optimized decision-making. This application adopts a multi-agent group relative strategy optimization method, combined with reinforcement learning to train AUV agents. The agents make adaptive decisions based on state information, and the group computing module coordinates actions to solve the problems of traditional methods and achieve a global optimal strategy. This application introduces relative advantage to incentivize agents to compete and cooperate, and designs a combined loss function that includes policy loss and KL divergence loss. This ensures stable policy updates, prevents overfitting, and provides a stable and efficient training method for multi-agent reinforcement learning, continuously improving system performance.

[0026] Below, we will refer to Figures 1 to 10 Each step of the multi-AUV collaborative perception method based on multi-agent group relative strategy optimization in this example embodiment is described in more detail.

[0027] In step S101, a multi-AUV collaborative perception system model is constructed. The multi-AUV collaborative perception system model includes the deployment relationship of underwater Internet of Things devices, multiple AUVs and ground stations, and defines an underwater communication model, an AUV data acquisition model and AUV motion constraints.

[0028] Specifically, the target area of the device is a three-dimensional area, which includes a ground station, AUVs and Underwater Internet of Things (IoUT) devices. AUV needs to avoid obstacles such as reefs and sediments. Each AUV The three-dimensional coordinate position of is expressed as: Similarly, the three-dimensional coordinate position of the IoUT device is expressed as Each AUV is powered by a fully charged battery and equipped with sonar, underwater acoustic communication equipment, and a horizontal acoustic Doppler current profiler (HADCP) to measure ocean currents and optimize data collection. The AUV collects data from predefined IoUT devices, and the initial data volume of each device is recorded as At each time step ,AUV performs motion, data collection, and transmission sequentially, which consumes energy and affects the overall perception latency.,The data collected from underwater IoT devices are aggregated on the autonomous underwater vehicle,,which improves data accuracy and provides support for subsequent mission,execution.

[0029] Underwater Communication Model: This model considers the complex acoustic communication characteristics between underwater IoUT devices and AUVs, encompassing key factors such as ambient noise, channel attenuation, and communication rate. It also accounts for the impact of both line-of-sight (LOS) and non-line-of-sight (NLOS) communications. Due to the high attenuation and bandwidth limitations of underwater acoustic communication, noise interference, channel propagation characteristics, and communication rate must be considered to optimize data acquisition efficiency and ensure reliable information transmission.

[0030] Underwater noise model: The main noise sources of underwater acoustic communications include: Fluid noise ( ): caused by ocean current velocity, its model is:

[0031] Ship noise ( ): generated by ship activities, its model is:

[0032] in, Indicates the intensity of ship activity.

[0033] Wave noise ( ): caused by the motion of ocean waves, its model is:

[0034] in, Indicates the wave speed.

[0035] Thermal noise ( ): caused by thermal effect, its model is:

[0036] Total noise model:

[0037] Communication between IoUT devices and AUV: The communication between underwater IoUT devices and AUVs includes line-of-sight (LOS) communication and non-line-of-sight (NLOS) communication, which is affected by factors such as underwater turbulence, obstacles, and water depth changes.

[0038] Line-of-sight (LOS) communication: In LOS communication, sound waves travel directly between the IoUT device and the AUV, unobstructed by obstacles. This mode offers low path loss, high signal strength, and excellent communication quality, but is susceptible to underwater terrain limitations. The propagation distance of line-of-sight communication is:

[0039] Non-line-of-sight (NLOS) communications: In NLOS communications, sound waves reflect off the water surface or seabed, increasing propagation delay and signal attenuation. Water reflections introduce phase shifts and energy loss, particularly in high seas. Seabed reflections depend on the composition of the seabed (e.g., sand, rock, or mud). Harder surfaces reflect more strongly, while softer seabeds cause signal dispersion. The formula for calculating the shortest NLOS transmission distance is as follows:

[0040]

[0041] in, For water depth, is the height of the IoUT device from the seabed, is the height of AUV from the seabed.

[0042] Communication rate: Underwater acoustic communication is affected by path loss and noise interference. It is given by:

[0043] in, is the propagation loss factor, is the absorption coefficient, which is expressed as:

[0044] Therefore, the signal-to-noise ratio can be calculated as:

[0045] in, is the frequency-dependent noise factor.

[0046] Taking into account the reflection from the water surface and the seabed, the minimum signal-to-noise ratio for NLOS communication is:

[0047] in, and Represents the signal enhancement factor of the shortest non-line-of-sight path reflected from the water surface and the seabed, respectively.

[0048] Therefore, the communication rate It can be calculated as follows:

[0049] in, represents the total bandwidth, and the transmission power of the underwater IoUT device is , represents the electronic efficiency, the average communication rate is calculated as follows:

[0050] AUV motion and data acquisition model: AUV operates in a restricted 3D underwater environment and needs to avoid obstacles such as currents, turbulence, and reefs. ,AUV performs movement and data collection in sequence, and consumes energy.,To ensure that the AUV operates efficiently and safely, a penalty mechanism is introduced to,avoid bad behaviors such as boundary exceeding, energy exhaustion,,collision, and incomplete data collection.

[0051] AUV movement model: Each AUV moves in a specific direction Upward movement distance ,in, Represents the maximum moving distance in a single time slot. It is known that the speed of the AUV is , and its moving time is calculated as follows:

[0052] The AUV must avoid underwater obstacles and move within a predefined boundary. Therefore, a penalty term is introduced Constraints mainly include the following penalty constraints: Boundary Exceeding Penalty: If the AUV exceeds the operating area, it will be penalized:

[0053] Collision Penalty: If an AUV collides with another AUV or an obstacle, it will be penalized:

[0054] Energy depletion penalty: If the remaining energy of the AUV falls below the threshold and cannot continue the mission, it will receive a penalty:

[0055] Data collection model: AUV uses underwater acoustic communication to The underwater IoUT devices in the water collect sensory data.

[0056] in, Represents all IoUT devices within the UAV’s sensing range. Indicates that at the IoUT device time The amount of data collected.

[0057] The data acquisition time calculation is determined by the amount of data and the data transmission rate:

[0058] Age of Misinformation Model (AoII): Assume that an underwater autonomous underwater vehicle (AUV) decides to collect sensory data from an underwater Internet of Things (IoUT) device, and the data is transmitted through an unreliable channel that may have transmission errors. Assume that the channel transmission follows an independent and identically distributed (iid) Bernoulli distribution, that is, the data packets are transmitted with probability Successful collection, with probability Acquisition failed. The age of error information is a key metric for evaluating the freshness and accuracy of data in IoT systems, particularly in collaborative systems such as autonomous underwater vehicles (AUVs) in multi-agent environments. It considers both the mismatch between transmitted and received information and the duration of this mismatch. The age of error information for all devices in the underwater collaborative perception model is calculated as follows:

[0059] in, Indicates the number of devices that have collected data. and They represent the status of the collected dataset and the status of the real dataset respectively.

[0060] Energy consumption model: AUV consumes energy during movement and data collection. The average energy consumption of all autonomous underwater vehicles is given by the following formula:

[0061] in, and They represent the AUV's moving power and hovering power (data collection process) respectively.

[0062] Total penalty model: According to the AUV in the collection process, at time Total punishment It can be expressed as follows:

[0063] In step S102, based on the multi-AUV collaborative perception system model, under the constraints of energy consumption and communication bandwidth, an objective function is constructed with the goal of minimizing the age of erroneous information in data collection.

[0064] Specifically, we construct an objective function and formalize the problem: the multi-AUV collaborative perception problem is expressed as an optimization task. Under the constraints of communication rate, energy consumption, etc., the age of error information is minimized. The optimization problem can be expressed as follows:

[0065] in, Refers to the data collection coverage, which ensures that the autonomous underwater vehicle can cover the required area to achieve effective perception. represents the average energy consumption, is the maximum energy consumption allowed. This constraint ensures that the AUV can operate efficiently and does not deplete its energy reserves. Indicates time consumption, Ensure that the time consumption is within the acceptable range for real-time applications. Ensure that no two autonomous underwater vehicles are in the same location at any given moment, thus avoiding collisions and achieving safe navigation. Limiting the activities of autonomous underwater vehicles to predefined boundaries ensures that perception operations are carried out within designated operating areas for safe and efficient data collection.

[0066] In step S103, the objective function is converted into a partially observable Markov decision process to define the state space, action space and individual AUV reward.

[0067] Specifically, the optimization problem is transformed into a partially observable Markov decision problem. A multi-agent group relative policy optimization (GRPO) method is proposed. This method involves multiple agents (agent 1, …, agent i), each equipped with its own policy model and reference model. These models receive state information from the environment and generate corresponding policy outputs. Each agent's policy, together with its reward model, interacts with the environment to make decisions and with a group computation module, which ultimately optimizes the relative policies among the agents. The group computation module integrates the policies of all agents to achieve optimal collective behavior. This method helps coordinate the actions of individual agents in a multi-agent system to obtain a globally optimal behavioral strategy.

[0068] The state space includes the parameters of the autonomous underwater vehicle and the state of the entire system. The state of the autonomous underwater vehicle includes its position and velocity. The system state includes data collection coverage, energy consumption, error information age, time consumption and penalty. This optimization problem is a Markov decision process (MDP), where the reward Depends only on the current state , the state includes data coverage, age of error information (AOII), energy consumption, delay and penalty terms. This satisfies the Markov property because the future state and rewards Determined by the current state, i.e. . It conforms to the Markov property and can be solved by reinforcement learning methods.

[0069] During collaborative perception and data collection by multiple autonomous underwater vehicles (AUVs), each AUV must make decisions based on dynamic environmental changes, aiming to minimize the age of error information (AOII) while ensuring energy consumption and latency constraints and maximizing data collection coverage. Traditional methods struggle to find optimal solutions. Using reinforcement learning to train an AUV agent can help find the optimal path during data collection. The design of the agent primarily involves input states and an action space.

[0070] The state space includes the parameters of the AUV and the state of the entire system. The state of the AUV includes its position and velocity. The system state includes data collection coverage, energy consumption, error information age, time consumption, and penalty.

[0071] The action space consists of four dimensions: directional control for moving in three-dimensional space, and and The agent selects actions to adjust the trajectory of the AUV to optimize objectives such as coverage, communication, or energy consumption.

[0072] In step S104, based on the state space, action space and individual AUV rewards, the multi-agent group relative strategy optimization algorithm is used to collaboratively optimize the perception strategies of multiple AUVs to obtain the optimal perception strategy.

[0073] Specifically, such as Figure 2 The block diagram of the proposed algorithm is shown in Figure 2. The perception policy is optimized using a multi-agent group-based relative policy optimization algorithm. An AUV observes its current position and the position of a target underwater Internet of Things (IoUT) device and then selects an action using a policy network. Trajectories of multiple AUVs are collected in each episode, and rewards are calculated based on a reward function. After collecting a batch of trajectories, the relative advantage of each trajectory is calculated by comparing the reward with the group average reward. This relative advantage helps guide policy updates. During each trajectory update, the policy is adjusted through multiple GRPO iterations, using a clipped surrogate loss and KL divergence penalty similar to proximal policy optimization to ensure stability. The reward model is continuously updated via a replay mechanism, using experience stored in an experience buffer and a mean squared error loss (or contrastive loss) to improve reward predictions over time. The proposed method balances exploration and exploitation by simultaneously considering the individual performance of the agent and its performance relative to the group. The combination of clipped objectives and KL divergence regularization provides a stable and efficient training method for multi-agent reinforcement learning scenarios.

[0074] Group relative reward: each agent in the group The relative advantage of the individual agent is calculated as the reward Average reward for the group The difference between:

[0075]

[0076] This group relative advantage reward calculation encourages agents that perform better than the group average, while penalizing agents that perform poorly.

[0077] Strategy loss: It is to improve the strategy while maintaining stability, which is achieved through a clipping objective function. The loss is calculated using the following formula:

[0078] in, Under the current strategy, in the state Take action The probability of is the probability of taking the same action under the old policy. The parameterized clipping function keeps the ratio of the new probability to the old probability within a safe range to prevent drastic policy changes, thus ensuring the stability of the learning process.

[0079] KL Divergence Loss: To further standardize the training process and prevent the model from overfitting to the current policy, we introduce the KL Divergence Loss between the reference policy and the current policy. This ensures that the updated policy does not deviate significantly from the reference policy. The KL Divergence Loss is calculated as follows:

[0080] in, Reference strategy With the current strategy The KL divergence between is a regularization hyperparameter that controls the weight of KL divergence in the total loss.

[0081] Total loss It is the sum of the policy loss and the KL divergence loss. The policy is optimized by minimizing this combined loss function. The total loss is calculated as:

[0082] Through the above method, this application realizes the minimization of the error information age of data collection during the collaborative perception process of multiple AUVs under the constraints of energy consumption, communication bandwidth, etc.

[0083] In a specific embodiment, the key parameters are configured as follows: the learning rate of the policy network is 1×10 -4 , the learning rate of the reward network is 2×10 -5 , discount factor , clipping threshold , KL divergence coefficient The experimental environment simulates a 1km×1km×1km three-dimensional detection area, in which underwater Internet of Things (IoUT) devices are randomly distributed. In this environment, multiple autonomous underwater vehicles collaborate (under energy constraints (each AUV has an initial energy of 20 kWh) to dynamically calculate state and action dimensions. The communication frequency range is 2.4-2.483 GHz, and the wavelength is , with a noise power of -90dBm. Energy consumption follows a 10W model during propulsion and a 0.398W model during hovering. Optimization objectives include coverage, age of error information (AOII), time delay, and energy efficiency. A two-hidden-layer actor-critic architecture (256 nodes per layer) with ReLU activation function is used for optimization, enhanced by L2 regularization (weight decay coefficient of 1×10-2), gradient clipping, and layer normalization. This configuration achieves joint optimization of energy efficiency and information freshness in a multi-AUV system through an adaptive exploration-exploitation balance. Specifically, group updates (group size 64) and generalized proximal policy optimization (GRPO) iterations (3 cycles) are performed over 100,000 training rounds.

[0084] During the training phase, the reward convergence is as follows Figure 3 As shown in , the reward increases steadily during training and eventually stabilizes, which shows that the policy optimization is effective. Figure 4 As shown in Figure 2, the AOII initially fluctuates, and then continues to decrease, which indicates that as the training progresses, the proposed method can effectively reduce the overall system error information age. Figure 5 , showing that the cumulative energy consumption gradually decreases, which reflects the optimization of energy use by the model. Figure 6 As shown, the overall time expenditure is reduced, which indicates that the efficiency is improved in the perception task execution and the real-time performance is optimized.

[0085] To evaluate the performance of the proposed method, we compared it with several representative research results in the field of multi-agent reinforcement learning. The Multi-Agent Dual Deep Q-Network (MADDQN) employs a dual Q-learning approach to handle multi-agent environments, improving stability and reducing overestimation bias, thereby providing more accurate action value estimates. Cooperative Deep Q-Learning (CDQL) leverages deep Q-learning and communication mechanisms to promote collaboration between agents. The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm employs an actor-critic architecture to address behavioral coordination in mixed cooperative and competitive environments. Finally, the Distributed Proximal Policy Optimization (DPPO) algorithm improves learning efficiency and stability in multi-agent systems through a distributed policy optimization approach.

[0086] Figure 7-10The performance of different methods in terms of error information age index (AOII), data collection coverage, energy consumption, and time consumption as the number of underwater Internet of Things (IoUT) devices changes is presented. Overall, the proposed multi-agent group (MAGRPO) method significantly outperforms other compared methods in the multi-autonomous underwater vehicle (AUV) collaborative perception scenario. Specifically, Figure 7 In the data, as the number of IoUT devices increases, the growth rate of average AOII under the MAGRPO method is much lower than that of other methods, and always remains at a relatively low level. In contrast, the upward trend of the distributed proximal policy optimization (DPPO) method is the most obvious, increasing by more than four times, while MAGRPO only increases by about 1.5 times. Figure 8 Among them, the MAGRPO method can maintain a relatively stable and high data collection coverage. On the contrary, for other methods such as DPPO and multi-agent deep deterministic policy gradient algorithm (MADDPG), the coverage shows a more obvious downward trend as the number of devices increases. Figure 9 The energy consumption growth under the MAGRPO method is relatively slow. Compared with other methods, its advantages in energy consumption control are very significant. Figure 10 In terms of time consumption, MAGRPO also maintains good experimental results and does not show a significant increase.

[0087] Through the above-mentioned multi-AUV collaborative perception method based on multi-agent group relative strategy optimization, on the one hand, firstly, factors such as underwater communication, noise, as well as environmental constraints and penalty mechanisms are integrated to construct a multi-AUV collaborative perception system model; with the core of reducing the age of erroneous information in collected data, the objective function is constructed in combination with relevant factors; then, considering the underwater environment and the limitations of AUV information acquisition, the optimization problem is transformed into a partially observable Markov decision process; finally, the multi-agent group relative strategy optimization algorithm is used to conduct reinforcement learning training on the AUV agents, coordinate actions, achieve perception strategy optimization, and improve system performance. On the other hand, this application constructs an underwater communication model, an AUV data acquisition model, and an AUV motion constraint, comprehensively considering actual factors such as line-of-sight and non-line-of-sight communication, obstacles, boundary restrictions, etc., and also introduces a penalty mechanism to ensure AUV operation, providing a solid foundation for optimized decision-making. This application adopts a multi-agent group relative strategy optimization method, combined with reinforcement learning to train AUV agents. The agents make adaptive decisions based on state information, and the group computing module coordinates actions to solve the problems of traditional methods and achieve a global optimal strategy. This application introduces relative advantage to incentivize agents to compete and cooperate, and designs a combined loss function that includes policy loss and KL divergence loss. This ensures stable policy updates, prevents overfitting, and provides a stable and efficient training method for multi-agent reinforcement learning, continuously improving system performance.

[0088] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0089] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. A multi-AUV collaborative perception method based on multi-agent group relative strategy optimization, characterized in that: The method includes: Construct a multi-AUV collaborative perception system model, which includes the deployment relationship between underwater IoT devices, multiple AUVs, and ground stations, and defines the underwater communication model, AUV data acquisition model, and AUV motion constraints. Based on the multi-AUV collaborative perception system model, under the constraints of energy consumption and communication bandwidth, the objective function is constructed with the goal of minimizing the age of erroneous information in data collection. The objective function is converted into a partially observable Markov decision process, defining a state space, an action space, and an individual AUV reward; wherein the state space includes the AUV's position, velocity, system data coverage, energy consumption, error information age, and penalty terms, and the action space includes the AUV's three-dimensional movement direction and velocity adjustment; Based on the state space, action space and individual AUV rewards, the multi-agent group relative strategy optimization algorithm is used to collaboratively optimize the perception strategies of multiple AUVs to obtain the optimal perception strategy.

2. The multi-AUV collaborative perception method based on multi-agent group relative strategy optimization according to claim 1 is characterized in that: Underwater communication models include: Underwater noise models consisting of fluid noise, ship noise, wave noise, and thermal noise, as well as line-of-sight and non-line-of-sight communications between underwater IoT devices and multiple AUVs; The shortest path distance for non-line-of-sight communication is calculated based on the water depth and the height of the equipment and AUV above the seabed.

3. The multi-AUV collaborative perception method based on multi-agent group relative strategy optimization according to claim 2 is characterized in that: The AUV data collection model is determined based on the AUV's perception radius and the spatial distribution of underwater IoT devices. The AUV data collection model includes an error information age model, an energy consumption model, and a total penalty model.

4. The multi-AUV collaborative perception method based on multi-agent group relative strategy optimization according to claim 3 is characterized in that: AUV motion constraints include AUV movement distance limit, obstacle avoidance rules and boundary constraints.

5. The multi-AUV collaborative perception method based on multi-agent group relative strategy optimization according to claim 4 is characterized in that: The expression of the objective function is: in, The age of misinformation for all underwater IoT devices, is the data collection coverage, For the minimum data collection coverage allowed, is the average energy consumption, is the maximum energy consumption allowed, Indicates time consumption, is the maximum time allowed, For AUV exist The location at the moment, For AUV exist The location at the moment, It is the boundary range of AUV movement.

6. The multi-AUV collaborative perception method based on multi-agent group relative strategy optimization according to claim 5 is characterized in that: The state space includes the state of the AUV and the state of the multi-AUV collaborative perception system model. The state of the AUV includes the position and speed of the AUV. The state of the multi-AUV collaborative perception system model includes data collection coverage, energy consumption, error information age, time consumption and penalty. The action space includes the three-dimensional movement direction and speed adjustment of the AUV.

7. The multi-AUV collaborative perception method based on multi-agent group relative strategy optimization according to claim 6 is characterized in that: Based on the state space, action space, and individual AUV rewards, the multi-agent group relative strategy optimization algorithm is used to collaboratively optimize the perception strategies of multiple AUVs to obtain the optimal perception strategy. The steps include: Calculate the relative advantage of each agent in the group based on the difference between the individual AUV reward and the group average reward; Based on the relative strength of each agent in the population, construct the policy loss: in, For the current strategy Next, in state Take action probability; For reference strategy Next, in state Take action probability; It is the cutoff threshold, which is used to limit the amplitude of strategy update during strategy optimization; is the relative advantage of each agent in the group, is the expected value of the multiple sampling process, which represents the expected reward under a certain strategy; Introduce KL divergence loss to obtain KL divergence loss: in, For reference strategy With the current strategy The KL divergence between is the regularization hyperparameter; According to the policy loss and KL divergence loss, the total loss is obtained: The total loss is used to optimize the perception strategies of multiple AUVs to achieve global optimal behavior and obtain the optimal perception strategy.