Cluster path planning method and device, equipment, storage medium and program product

Through multi-objective optimization reinforcement learning algorithm training target graph attention network, combined with multiple sub-reward functions, the problem of solving quality and computing speed in underwater robot cluster path planning is solved, and efficient and robust path planning is achieved.

CN120008618AActive Publication Date: 2025-05-16TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202510488425.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the path planning algorithm of underwater robot clusters is difficult to take into account both the solution quality and the calculation speed, especially in a large-scale and dynamically changing environment.

Method used

The multi-objective optimization reinforcement learning algorithm is used to train the target graph attention network, and the target path of the underwater vehicle is determined through the graph structure data and the target graph attention network, combining the sub-reward functions of total energy consumption, energy distribution differences, task scheduling time constraints and information collection efficiency.

Benefits of technology

It realizes the global optimization capability and robustness of path generation in complex and non-stationary environments, and improves the quality of efficiency and reconciliation of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120008618A_ABST
    Figure CN120008618A_ABST
Patent Text Reader

Abstract

The invention relates to a cluster path planning method and device, equipment, a storage medium and a program product. The method comprises the steps that graph structure data of a cluster are acquired, the cluster comprises a plurality of underwater vehicles, nodes of the graph structure data are used for representing position information of the underwater vehicles, and edges connecting the nodes in the graph structure data are used for representing communication links between the underwater vehicles; a target path of each underwater vehicle is determined according to the graph structure data and a target graph attention network, the target graph attention network is obtained by training an initial graph attention network through a multi-target optimization reinforcement learning algorithm, and a reward function of the multi-target optimization reinforcement learning algorithm comprises a plurality of sub-reward functions; the plurality of sub-reward functions comprise a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function and an information acquisition efficiency sub-reward function. By adopting the method, the solution quality and the calculation speed can be considered during path planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of path planning, and in particular to a cluster path planning method, device, equipment, storage medium and program product. Background Art

[0002] Faced with the complex ocean hydrological environment and the dangers of underwater manual operations, underwater robots have gradually become a reliable way to explore the ocean. Under the task of information collection, research on path planning algorithms for underwater robot swarms has also gradually deepened.

[0003] In the prior art, the collaborative path planning problem of underwater robot clusters can be abstracted as a combinatorial optimization problem, and the path planning problem can be solved by a heuristic algorithm, such as a graph search-based heuristic algorithm, a sampling-based heuristic algorithm, or a biological heuristic algorithm.

[0004] However, heuristic algorithms have poor adaptability to large-scale, dynamically changing path planning problems, and it is difficult to strike a balance between solution quality and calculation speed. Summary of the invention

[0005] Based on this, it is necessary to provide a cluster path planning method, device, equipment, storage medium and program product that can take into account both solution quality and calculation speed in order to address the above technical problems.

[0006] In a first aspect, the present application provides a cluster path planning method, comprising:

[0007] Acquire graph structure data of a cluster, the cluster including a plurality of underwater vehicles, the nodes of the graph structure data being used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data being used to represent the communication links between the underwater vehicles;

[0008] The target path of each underwater vehicle is determined according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: total energy consumption sub-reward function, energy distribution difference sub-reward function, task scheduling time constraint sub-reward function and information collection efficiency sub-reward function.

[0009] In one embodiment, the target graph attention network includes a graph attention encoder and a graph decoder, and determines the target path of each underwater vehicle according to the graph structure data and the target graph attention network, including:

[0010] The graph structure data is subjected to graph pooling and graph normalization to determine the target information; the target information is encoded according to the graph attention encoder to determine the encoding result of each node; the target path of each underwater vehicle is determined based on the encoding result and the graph decoder.

[0011] In one embodiment, the initial graph attention network includes an initial graph attention encoder and an initial decoder, and the training process of the target graph attention network includes:

[0012] Generate multiple initial paths based on graph structure data and the initial graph attention network; determine the cumulative reward value of each initial path based on the initial path and the reward function; train the initial graph attention encoder and initial decoder based on each cumulative reward value and the policy gradient algorithm until the convergence conditions are met.

[0013] In one embodiment, graph structure data is subjected to graph pooling and graph normalization to determine target information, including:

[0014] The importance score of each node is calculated according to the feature representation, activation function and weight parameters of each node; the graph structure data is pooled according to each importance score; the graph is normalized according to the feature representation of each node, the node feature mean and node feature standard deviation of the graph structure data after graph pooling, and the target information of each node is determined.

[0015] In one embodiment, the graph decoder is a masked autoregressive decoder, and the target path of each underwater vehicle is determined according to the encoding result and the graph decoder, including:

[0016] The dynamic mask information is obtained, which includes the path nodes that have been visited, the task scheduling time constraints and the energy consumption constraints. For each underwater vehicle, the encoding result and the dynamic mask information are input into the masked autoregressive decoder, and the target path of the underwater vehicle is determined according to the multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

[0017] In one embodiment, training the initial graph attention network according to each accumulated reward value and the policy gradient algorithm includes:

[0018] For each initial path, the policy gradient of the initial path is calculated based on the cumulative reward value and the baseline function value of the initial path; the average gradient is determined based on the policy gradient of each initial path; and the parameters of the initial graph attention network are updated based on the gradient ascent method and the average gradient.

[0019] In a second aspect, the present application also provides a cluster path planning device, including:

[0020] An acquisition module, used to acquire graph structure data of a cluster, the cluster includes a plurality of underwater vehicles, the nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles;

[0021] The planning module is used to determine the target path of each underwater vehicle based on the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions: total energy consumption sub-reward function, energy distribution difference sub-reward function, task scheduling constraint sub-reward function and information collection efficiency sub-reward function.

[0022] In one of the embodiments, the target graph attention network includes a graph attention encoder and a graph decoder, and a planning module, which is specifically used to perform graph pooling and graph normalization processing on the graph structure data to determine the target information; encode the target information according to the graph attention encoder to determine the encoding result of each node; and determine the target path of each underwater vehicle based on the encoding result and the graph decoder.

[0023] In one of the embodiments, the initial graph attention network includes an initial graph attention encoder and an initial decoder, and the cluster path planning device also includes a training module, which is specifically used to generate multiple initial paths based on graph structure data and the initial graph attention network; determine the cumulative reward value of each initial path based on the initial path and the reward function; and train the initial graph attention encoder and the initial decoder based on each cumulative reward value and the policy gradient algorithm until the convergence condition is met.

[0024] In one of the embodiments, the planning module is specifically used to calculate the importance score of each node based on the feature representation, activation function and weight parameters of each node; perform graph pooling on the graph structure data according to each importance score; perform graph normalization on each node based on the feature representation of each node, the node feature mean and node feature standard deviation of the graph structure data after graph pooling processing, and determine the target information of each node.

[0025] In one embodiment, the graph decoder is a masked autoregressive decoder, and the planning module is specifically used to obtain dynamic mask information, which includes path nodes that have been visited, task scheduling time constraints, and energy consumption constraints; for each underwater vehicle, the encoding result and the dynamic mask information are input into the masked autoregressive decoder, and the target path of the underwater vehicle is determined according to multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

[0026] In one of the embodiments, the training module is specifically used to calculate the policy gradient of each initial path based on the cumulative reward value and the baseline function value of the initial path; determine the average gradient based on the policy gradient of each initial path; and update the parameters of the initial graph attention network based on the gradient ascent method and the average gradient.

[0027] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, any of the methods described in the first aspect is implemented.

[0028] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the methods described in the first aspect above.

[0029] In a fifth aspect, the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements any of the methods described in the first aspect above.

[0030] The above-mentioned cluster path planning method, device, equipment, storage medium and program product obtain the graph structure data of the cluster, the cluster includes multiple underwater vehicles, the nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting each node in the graph structure data are used to represent the communication link between each underwater vehicle; the target path of each underwater vehicle is determined according to the graph structure data and the target graph attention network, the target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function and an information collection efficiency sub-reward function. The graph attention network is trained by a multi-objective optimization reinforcement learning algorithm, so that by designing a weighted sub-reward function, multiple optimization objectives such as minimizing total energy consumption, energy balance constraints, task scheduling constraints and maximizing information collection efficiency are incorporated into a unified framework, and a dynamic balance between different objectives is achieved. The graph attention network obtained by training in this way can have the global optimization ability and robustness of path generation in a complex and non-stationary environment, thereby improving the efficiency of path planning and the quality of the solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 A schematic diagram of a flow chart of a cluster path planning method in one embodiment;

[0033] Figure 2 A schematic diagram of the correspondence between the local coordinate system and the global coordinate system in one embodiment;

[0034] Figure 3 A schematic diagram of a flow chart of steps for determining a target path of each underwater vehicle according to graph structure data and a target graph attention network in one embodiment;

[0035] Figure 4 A schematic diagram of a flow chart of steps of performing graph pooling and graph normalization processing on graph structure data in one embodiment;

[0036] Figure 5 A schematic diagram of a flow chart of a step of determining a target path of each underwater vehicle according to an encoding result and a graph decoder in one embodiment;

[0037] Figure 6 Schematic diagram of a process of training a target graph attention network in one embodiment;

[0038] Figure 7 A schematic diagram of a flow chart of target graph attention network training steps in another embodiment;

[0039] Figure 8 A schematic diagram of a training process in an embodiment;

[0040] Fig. 9 A schematic diagram of a flow chart of a cluster path planning method in another embodiment;

[0041] Fig.10 A schematic diagram of the relationship between average path planning time and the number of fixed nodes in one embodiment;

[0042] Fig.11 A schematic diagram of the relationship between multi-objective optimization values ​​and the number of nodes in an embodiment;

[0043] Fig.12 A schematic diagram of a routing result of path planning in one embodiment;

[0044] Fig.13 A schematic diagram of a routing result of path planning in another embodiment;

[0045] Fig.14 It is a structural block diagram of a cluster path planning device in one embodiment;

[0046] Fig.15 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] Faced with the complex ocean hydrological environment and the dangers of underwater manual operations, underwater robots have gradually become a reliable way to explore the ocean. Under the task of information collection, research on path planning algorithms for underwater robot swarms has also gradually deepened.

[0049] In the prior art, the collaborative path planning problem of underwater robot clusters can be abstracted as a combinatorial optimization problem, and the path planning problem can be solved by a heuristic algorithm, such as a graph search-based heuristic algorithm, a sampling-based heuristic algorithm, or a biological heuristic algorithm.

[0050] However, heuristic algorithms have poor adaptability to large-scale, dynamically changing path planning problems, and it is difficult to strike a balance between solution quality and calculation speed.

[0051] In view of this, an embodiment of the present application provides a cluster path planning method that can take into account both solution quality and calculation speed. The cluster path planning method provided in the embodiment of the present application, its execution subject can be a cluster path planning device, and the cluster path planning device can be implemented by software, hardware, or a combination of software and hardware. It can be embedded in or independent of the processor in the computer device in the form of hardware, or it can be stored in the memory of the computer device in the form of software. In the following method embodiments, the execution subject is a computer device as an example for explanation, wherein the computer device can be a server or a desktop computer. The embodiment of the present application does not limit the specific type of the computer device.

[0052] In an exemplary embodiment, Figure 1 As shown, a cluster path planning method is provided, including the following steps 101 to 102. Among them:

[0053] Step 101, obtaining graph structure data of a cluster, the cluster including a plurality of underwater vehicles, the nodes of the graph structure data are used to represent the location information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles.

[0054] Optionally, in an underwater environment, an Autonomous Underwater Vehicle (AUV) can collect underwater information through sensors carried by itself and communicate with other underwater vehicles to complete the underwater information collection task.

[0055] Optionally, the graph structure data also includes multiple seabed fixed sensors, the nodes of the graph structure are also used to represent the location information of the multiple seabed fixed sensors, and the edges connecting the nodes are also used to represent the communication links between each underwater vehicle and each seabed fixed sensor.

[0056] Optionally, the sensor node network can be represented by a graph structure, and the topology of the cluster's graph structure is dynamic and uncertain, which is closely related to various factors of the underwater environment (such as water currents, ocean weather, equipment failures, etc.).

[0057] For example, assume that the graph structure data consists of N nodes, and each node i (i=1,2,…,N) has a basic attribute position p i =(x i ,y i ), the basic attribute position can represent the position of the underwater vehicle in the local coordinate system, and each node also includes the basic attribute capacity C i , the basic attribute capacity can represent the amount of information carried by the underwater vehicle’s sensors at the current moment. In graph structure data, the edges between nodes represent the connection links between them, and the weight of the edge w ij It is usually the communication cost considering interference such as ocean currents. To simplify the analysis, it is assumed that the connection link between each pair of nodes is stable within a specific time period. The final representation is a fully connected weighted undirected graph: G=(V,E), where V={v1,v2,…,v n} is the set of sensor nodes, is an edge set, representing the connection link. At time t, the state of each underwater vehicle can be ,in is the position of the AUV at time t, It consumes energy.

[0058] Optionally, the goal of the AUV swarm is to dynamically adjust the task allocation and path planning of swarm members according to mission requirements and energy constraints.

[0059] Step 102, determine the target path of each underwater vehicle based on the graph structure data and the target graph attention network, the target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function and an information collection efficiency sub-reward function.

[0060] Optionally, the initial graph attention network can be trained by a multi-objective optimization reinforcement learning algorithm to obtain a target graph attention network. In this way, after the acquired graph structure data is input into the target graph attention network, the target graph attention network can input the actions of each underwater vehicle at the next moment, for example, the underwater vehicle's forward direction, forward speed, or forward acceleration.

[0061] Optionally, in the deep-sea Internet, when the cluster collaborates to perform information collection tasks, each AUV in the cluster needs to collaborate to maximize the efficiency of information collection and overcome challenges including energy consumption, task allocation, and communication limitations. The main characteristics of AUV clusters include the sparsity of task scheduling, dynamic changes in the environment, energy limitations, and timeliness of information. At the same time, the high dynamics of the marine environment also poses a serious challenge to the Markov nature of the path planning method. Therefore, underwater cluster path planning can be regarded as a data-driven weakly supervised learning Markov decision problem, which can be solved using the characteristics of reinforcement learning algorithms.

[0062] Among them, the sparsity of the above-mentioned task scheduling may mean that different tasks may be more important to certain AUVs in the cluster, resulting in sparse distribution of information collection tasks in time and space. Therefore, AUVs must be dynamically adjusted according to the importance of the current task and the energy situation. Dynamic changes in the environment may be that factors such as currents, waves, and weather in the marine environment will continue to change, affecting the path planning and communication quality of AUVs, which requires the AUV cluster to have the ability to adapt to environmental changes in real time. Energy limitations and the timeliness of information may mean that each AUV has limited energy, and it is necessary to reasonably allocate energy when performing tasks to ensure that the tasks can be completed and returned. At the same time, the timeliness of information requires that AUVs must complete the collection tasks in the shortest possible time and return information in a timely manner.

[0063] Optionally, the Graph Attention Network (GAT) is a deep learning model based on graph structured data, focusing on processing graph data with non-Euclidean structure. GAT introduces an attention mechanism to dynamically assign weights between graph nodes, thereby more effectively highlighting the influence of key nodes when aggregating neighborhood information. Compared with the traditional Graph Convolutional Network (GCN), GAT has stronger flexibility and expressiveness, especially in dynamic graphs, sparse graphs and large-scale graph data processing tasks.

[0064] Optionally, a multi-objective optimization reinforcement learning algorithm guides the AUV to learn the optimal strategy through multiple sub-reward functions. In the multi-objective optimization reinforcement learning algorithm, the path planning problem corresponding to the cluster information collection task can be regarded as a Markov decision process (MDP). The unmanned underwater vehicle optimizes the long-term cumulative reward through interactive learning strategies. In this process, the sub-reward function is designed as multiple independent optimization indicators, each of which corresponds to an optimization goal, and is weighted and combined through adjustable weight coefficients to achieve a dynamic balance between different goals. Under the MDP framework, the underwater vehicle moves in the underwater environment in a state. Make decisions by taking actions Impact the environment, get rewarded , and enter the next state This process has obvious Markov properties, that is, the future state depends only on the current state and current action, and has nothing to do with the historical state.

[0065] Optionally, a Markov decision process can be defined using a five-tuple It indicates that, represents the state space, i.e., the set of states of the underwater vehicle and the environment; A represents the action space, i.e., the set of decisions that the underwater vehicle can execute; Represents the state transition probability, that is, the state s is transferred to a new state after taking action a The probability distribution of ; R{s,a} represents the reward function, which defines the immediate reward obtained under the current state and action; γ∈(0,1] Represents the discount factor, which controls the weight of future rewards in the overall reward. ,The information collection task of underwater vehicles involves multiple key state variables, the state space It can be expressed as:

[0066]

[0067] in, is the position information of the i-th AUV; is the current remaining energy of the i-th AUV; is the information collection progress of the i-th AUV target, that is, the information distribution of the fixed sensor nodes on the seabed; The task scheduling status of the i-th AUV, including task completion status and remaining time.

[0068] Optionally, when performing a mission, an underwater vehicle needs to consume energy for propulsion, sensor operation, and data communication. Therefore, the total energy consumption is a key factor affecting the mission duration and overall performance. The total energy consumption sub-reward function of the AUV is It can be expressed by the following formula:

[0069]

[0070] Where N is the total number of AUVs, M is the total number of other AUVs connected to the i-th AUV, and the total energy consumption of each AUV can include its anti-ocean current movement and communication processes. In the following text, the energy consumption of anti-ocean current movement is respectively and communication energy consumption Make an introduction.

[0071] Alternatively, in a swarm mission, the uneven consumption of individual energy may cause some underwater vehicles to fail prematurely, affecting the overall mission completion. To measure the distribution divergence in a data set, the difference or divergence between two probability distributions can be quantified by a metric.

[0072] For example, the energy consumption distribution of different underwater vehicles can be measured by Kullback-Leibler divergence (KL divergence) to make it as close as possible to an ideal equilibrium distribution. The remaining energy distribution of each AUV is obtained as P(x), which represents the probability distribution of the remaining energy of the AUV, and the distribution of the optimal energy distribution of all AUVs in the cluster Q(x). The energy distribution difference sub-reward function It can be expressed by the following formula:

[0073]

[0074] Optionally, to ensure that each task is completed within the specified time and the action time of each AUV does not exceed its energy limit, that is , The time required for the i-th underwater vehicle to collect information from the j-th fixed node on the seafloor, that is, the task completion time, and the task scheduling time constraint sub-reward function It can be expressed as:

[0075]

[0076]

[0077] in, is the preset coefficient, is the timeout penalty function, which depends on whether the time window constraint of the task is guaranteed and the degree of exceeding the time limit.

[0078] Optionally, in order to ensure that as many information collection tasks as possible are completed within a limited time and resources, the information collection efficiency, i.e., the sub-reward function corresponding to the information collection success rate, is The ratio of the amount of effective information collected to the total amount of mission target information can be calculated as follows:

[0079]

[0080] in, is the amount of effective information collected by the i-th AUV, is the total amount of mission target information.

[0081] Alternatively, in a multi-objective optimization reinforcement learning algorithm, maximizing the cumulative discounted reward can be expressed by the following formula:

[0082]

[0083] in, is the instantaneous reward in the Markov decision process at time t, that is, the reward function, which is determined by the weighted sum of each sub-reward function, that is, ,in, It is an adjustable weight parameter used to control the relative importance of different optimization objectives. It can be understood that in maximizing the cumulative discounted reward, it is necessary to minimize the total energy consumption, minimize the energy distribution difference, satisfy the time window constraint of the task scheduling time, and maximize the information collection efficiency.

[0084] Optionally, when planning the cluster path, it is necessary to adapt to environmental changes, perform dynamic path planning based on the AUV's real-time energy status, task completion status, and environmental changes (such as water flow speed, communication delay), and optimize multiple objectives at the same time.

[0085] Next, the energy consumption of moving against the ocean currents and communication energy consumption To introduce:

[0086] For example, under the assumption of graph structure modeling, that is, two-dimensional modeling, the six-degree-of-freedom AUV motion can be further simplified to the motion in the plane. Ignoring the motion in the vertical direction, the motion problem of the AUV can be simplified to three degrees of freedom, that is, only considering the translation (forward, lateral) and rotation (yaw) motion in the horizontal plane, such as Figure 2 As shown in the figure, it is a schematic diagram of the coordinate system, where u represents the forward speed of this type of underwater robot, v is the lateral movement speed, θ is the pitch angle, φ is the barrel roll angle, ψ is the yaw angle, and r represents the yaw angular velocity. The following relationship exists between the global coordinate system and the local coordinate system of the underwater vehicle:

[0087]

[0088] in, represents the velocity component in the X-axis direction in the global coordinate system, represents the velocity component in the Y-axis direction in the global coordinate system, Represents the rate of change of the yaw angle over time.

[0089] Optionally, the kinetic model is as follows:

[0090]

[0091] Where m is the mass of the AUV; T is the thrust of the propeller water jet; and are the longitudinal and lateral additional masses respectively; and are the moment of inertia and additional moment of inertia about the center of gravity respectively; , , , and is the linear hydrodynamic derivative; , , , , , and is the nonlinear hydrodynamic derivative. The hydrodynamic derivative is usually captured by the Planar Motion Technique (PMM) test. In particular, represents the length of the real AUV between the vertical lines, It's real With model The reciprocal of the scaling factor between . , and Represents the disturbance of the wave.

[0092] For example, the dynamic system can be simplified by distributed modeling, and the hydrodynamic model can be decoupled into two independent parts: the drag-free motion equation and the turbulence field influence model, where the influence of the turbulence field directly acts on the velocity component of the AUV. In this way, the algorithm complexity can be reduced by simplifying the model. Secondly, the turbulence field characteristics of the research environment can be analyzed and set according to needs. The turbulence field at the working plane can be modeled based on the world coordinate system through the classic two-dimensional Navier-Stokes equation to be close to the actual ocean situation:

[0093]

[0094]

[0095]

[0096] in, , and are the velocity, vorticity and viscosity of the fluid, respectively. and They are the gradient operator and the Laplace operator, which can approximate a point in a plane in the world coordinate system. At the moment The water flow velocity is:

[0097]

[0098]

[0099] in, is the vortex center, and are the strength of the vortex and the radius of influence. Since the horizontal turbulence field has a dominant influence on the AUV, the resistance of the AUV can be approximated by the method of computational fluid dynamics (CFD):

[0100]

[0101] in, is the mass density of the fluid, is the longitudinal section area of ​​the AUV, is the drag coefficient, is the relative speed of the ocean current and the AUV, from which we can deduce the speed of the AUV in the time slot Energy consumption of translational and rotational motion:

[0102]

[0103]

[0104] in, is the electrical conversion efficiency, is the length of the AUV, is the mass of the AUV.

[0105] For example, the communication consumption caused by exchanging information For , it can be determined with the help of the classic Thorp model, as shown below:

[0106] The attenuation coefficient of the underwater acoustic signal with a frequency of f in the underwater acoustic channel at a distance l can be expressed as:

[0107]

[0108] Where A0 is the unit normalization constant, k is the propagation coefficient, is the absorption coefficient.

[0109] The acoustic path loss (dB) is given by:

[0110]

[0111] Among them, the first term on the right side of the above equation represents the propagation loss, and the second term represents the absorption loss. The constant k is usually between 2 and 4. Some of its classic values ​​depend on the different propagation conditions of the underwater acoustic signal: when the underwater acoustic signal is spherical diffusion, k=2; when it is cylindrical diffusion, k=1. In the embodiment of the present application, considering the actual underwater acoustic propagation situation, k=1.5 is taken.

[0112] When the acoustic frequency is kHz, the above equation can be simplified to the following equation:

[0113]

[0114] For low frequency signals, it can also be approximated using the following simpler equation form:

[0115]

[0116] In addition to considering underwater path loss, underwater acoustic signal propagation often needs to consider three noise sources to simulate signal noise, namely turbulence, ships, waves and thermal noise. The sum of the noise power can be expressed as

[0117]

[0118] in represents the total power spectral density of the ambient noise, Represents the factor by which noise decreases with frequency.

[0119] According to Shannon's theorem, the data transmission rate of AUV information output is defined as R c , assuming that the maximum transmission power of the AUV is , then the maximum communication distance in the plane can be obtained At the maximum communication efficiency, the communication consumption of the i-th AUV can be obtained as:

[0120]

[0121] in, is the total number of AUVs in the working plane, B is the bandwidth, , H represents the total power spectrum density of the ambient noise, represents the coefficient by which noise decreases with frequency, It is an indicative function. When the judgment condition in the brackets is met, the function value is 1, otherwise the function value is 0.

[0122] The above cluster path planning method obtains the graph structure data of the cluster, the cluster includes multiple underwater vehicles, the nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting each node in the graph structure data are used to represent the communication link between each underwater vehicle; the target path of each underwater vehicle is determined according to the graph structure data and the target graph attention network, the target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: total energy consumption sub-reward function, energy distribution difference sub-reward function, task scheduling time constraint sub-reward function and information collection efficiency sub-reward function. The graph attention network is trained by the multi-objective optimization reinforcement learning algorithm, so that by designing the weighted sub-reward function, multiple optimization objectives such as minimizing total energy consumption, energy balance constraint, task scheduling constraint and maximizing information collection efficiency are incorporated into a unified framework, and a dynamic balance between different objectives is achieved. The graph attention network obtained by training in this way can have the global optimization ability and robustness of path generation in a complex and non-stationary environment, thereby improving the efficiency of path planning and the quality of the solution.

[0123] In an exemplary embodiment, Figure 3 As shown, optionally, the target graph attention network includes a graph attention encoder and a graph decoder, and the target path of each underwater vehicle is determined according to the graph structure data and the target graph attention network, including the following steps 301 to 303. Among them:

[0124] Step 301, perform graph pooling and graph normalization processing on the graph structure data to determine the target information.

[0125] Optionally, graph pooling can be a pooling technology for graph structured data, which can gradually compress graph structured data into a fixed size, thereby reducing computational complexity and extracting higher-level graph structure information. Graph pooling can compress graph structured data into a fixed-size representation and retain key global and local information without destroying the graph structure through a preset node selection strategy or feature aggregation strategy.

[0126] Exemplarily, graph pooling can be performed by node importance evaluation, and important nodes can be selected according to the importance scores of each node, and finally a new coarsened graph can be constructed using these important nodes; nodes can also be mapped to several clusters according to similarity or other criteria through node clustering. In the clustering process, the attention mechanism can also be combined to assign clustering weights to each node. Nodes with higher weights are more likely to become representative nodes of the cluster, thereby retaining important information of the node in the pooled graph structure.

[0127] Optionally, graph normalization can be a normalization technology for graph structured data, which can solve the problem of inconsistent node feature distribution in graph neural networks. Since graph structured data usually have different scales and structures, traditional normalization methods may not be able to effectively adapt to these differences when processing graph structured data. Therefore, through graph normalization, part of the node features can be stabilized, reducing the impact of the graph structure size on the features, and improving the stability and convergence speed of model training.

[0128] Optionally, the target information may be a graph structure after graph pooling and graph normalization and a feature representation of each node in the graph structure.

[0129] For example, Figure 4 As shown, optionally, graph pooling and graph normalization are performed on the graph structure data to determine the target information, including the following steps 401 to 403. Among them:

[0130] Step 401, calculating the importance score of each node according to the feature representation, activation function and weight parameter of each node.

[0131] Optionally, for each node in the graph structure data, its importance score can be expressed by the following formula: :

[0132]

[0133] in, represents the feature representation of node i, W and b are learnable parameters, is the activation function, is the global normalized task weight parameter, The weight ratio of learnable task importance and attention mechanism.

[0134] Step 402: Perform graph pooling processing on the graph structure data according to each importance score.

[0135] Optionally, important nodes in the graph structure data can be selected according to importance scores. For example, a preset importance score threshold is obtained, and nodes with importance scores greater than the preset importance score threshold are regarded as important nodes. Alternatively, the nodes are sorted according to importance scores, and nodes with top importance scores are regarded as important nodes.

[0136] Optionally, a new coarsened graph can be constructed based on important nodes.

[0137] Step 403 , performing graph normalization processing on each node according to the feature representation of each node, the node feature mean and the node feature standard deviation of the graph structure data after graph pooling processing, and determining the target information of each node.

[0138] Optionally, for each node i in the graph structure G, its normalized feature It can be expressed by the following formula:

[0139]

[0140] in, is the original feature of the node, is the mean of graph-level node features after graph pooling, is the standard deviation of the graph-level features after graph pooling, and is the preset coefficient.

[0141] Through the above-mentioned graph pooling and normalization processing, the graph structure can be adjusted in real time according to the changes in nodes in the task to adapt to the ever-changing environmental topology. It has better robustness when processing tasks of different scales and ensures that the target graph attention network can still work stably in a dynamic environment.

[0142] Step 302: Encode the target information according to the graph attention encoder to determine the encoding result of each node.

[0143] Optionally, the encoding result of each node may be a feature vector obtained by aggregating information of neighboring nodes through an attention mechanism.

[0144] Optionally, the graph attention encoder can be a target graph attention network, whose backbone network mainly includes attention weight calculation, feature update and multi-head attention mechanism.

[0145] For example, GAT can assign dynamic weights to the neighborhood of each node through a learnable attention mechanism. Specifically, GAT can calculate the attention weight through the following formula:

[0146]

[0147] in, is the attention weight between node i and its neighbor node j; and are the feature representations of nodes i and j respectively; W is the learnable feature transformation matrix; a is the weight vector of the attention mechanism; | represents the concatenation operation of the features; LeakyRELU is the activation function.

[0148] Calculated attention weights It can be used to weight the features of the neighborhood nodes to generate updated node representations:

[0149]

[0150] in, It is a nonlinear activation function, usually ReLU can be used.

[0151] Optionally, the attention weight calculation can be combined with the characteristics of dynamic changes of nodes during task scheduling. By increasing the temporal dynamic modeling capability of node importance weights, it can more flexibly cope with the dynamics and randomness of the deep-sea environment.

[0152] In order to improve the robustness and stability of the model, GAT introduces a multi-head attention mechanism, which concatenates or averages the outputs of multiple independent attention heads to obtain richer node representations:

[0153]

[0154] Among them, k represents the kth head, || represents the splicing operation, represents the attention weight between node i corresponding to the kth head and its neighbor node j, Represents the feature transformation matrix that can be learned by the k-th head.

[0155] Optionally, a multi-scale attention mechanism can be introduced to further capture the global characteristics and local correlations of deep-sea environment tasks by integrating node features in different ranges, thereby improving GAT's modeling capabilities for sparse targets.

[0156] The above-mentioned GAT has the advantages of strong adaptability, efficient aggregation, and multi-head mechanism to enhance representation. The attention mechanism enables GAT to dynamically adjust the influence between nodes according to the different characteristics of the graph data, which is particularly suitable for sparse graphs and irregular graphs. GAT automatically selects the importance of neighbor nodes through attention weights, thereby capturing global and local information more accurately. Through the multi-head attention mechanism, GAT can learn feature representations from multiple different perspectives, thereby improving the robustness of the model.

[0157] Step 303: determine the target path of each underwater vehicle according to the encoding result and the graph decoder.

[0158] Optionally, the graph decoder may be a neural network module, and the encoding result may be converted into a target path, that is, the action of each underwater vehicle at the next moment, through the graph decoder.

[0159] Exemplarily, the graph decoder can be a recurrent neural network decoder, a Transformer decoder, or an autoregressive decoder, etc., which is not limited in the embodiments of the present application.

[0160] The above-mentioned graph structure data is subjected to graph pooling and graph normalization processing to determine the target information, the target information is encoded according to the graph attention encoder, the encoding result of each node is determined, and the target path of each underwater vehicle is determined according to the encoding result and the graph decoder, which can improve adaptability and robustness when determining the target path.

[0161] In an exemplary embodiment, Figure 5 As shown, optionally, the graph decoder is a masked autoregressive decoder, and the target path of each underwater vehicle is determined according to the encoding result and the graph decoder, including the following steps 501 to 502. Among them:

[0162] Step 501 , obtaining dynamic mask information, where the dynamic mask information includes the path nodes that have been visited, the task scheduling time constraints and the energy consumption constraints.

[0163] Optionally, the path nodes that have been visited can be obtained from the motion trajectory of each underwater vehicle. It can be understood that the motion trajectory of each underwater vehicle is continuously updated along with the path planning.

[0164] Optionally, by obtaining the task scheduling time constraints and energy consumption constraints, invalid target paths can be filtered out in the subsequent decoding process. For example, if a path does not meet the task scheduling time constraints and energy consumption constraints, that is, the path is unavailable, it can be excluded through mask information, so that the decoder can focus on valid path information.

[0165] Step 502: for each underwater vehicle, the encoding result and the dynamic mask information are input into a masked autoregressive decoder, and the target path of the underwater vehicle is determined according to multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

[0166] Optionally, when determining the target path of each underwater vehicle, the encoding result and dynamic mask information of the node can be input into the mask autoregressive decoder to predict the probability distribution from the current node to other reachable nodes. This probability distribution represents the possibility of choosing a different path to continue moving forward under the current state.

[0167] Optionally, when determining the target path of the underwater vehicle based on multiple candidate paths and corresponding probability distributions, the candidate path with the highest probability can be selected as the target path, or one of the multiple candidate paths can be randomly selected as the target path according to a random algorithm. Alternatively, multiple paths with higher probabilities can be selected at each step, that is, a certain number of optimal paths are retained, and then these selected paths are continued to be expanded and evaluated in subsequent steps, and finally the path with the highest score is selected as the target path.

[0168] The above method obtains dynamic mask information, which includes path nodes that have been visited, task scheduling time constraints and energy consumption constraints. For each underwater vehicle, the encoding result and the dynamic mask information are input into the masked autoregressive decoder. The target path of the underwater vehicle is determined according to the multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution. The dynamics of the cluster can be queried in each decision time step, that is, in the process of generating the target path, the generated path nodes are masked to avoid path duplication and improve the efficiency of the algorithm. At the same time, the target path is dynamically output by autoregressive decoding, so that the path generation process can flexibly respond to the complex changes of environmental dynamics, task execution dynamics and underwater environmental dynamics.

[0169] In an exemplary embodiment, Figure 6 As shown, optionally, the initial graph attention network includes an initial graph attention encoder and an initial decoder, and the training process of the target graph attention network includes the following steps 601 to 603. Among them:

[0170] Step 601, generating multiple initial paths according to the graph structure data and the initial graph attention network.

[0171] Optionally, the acquired graph structure data can be first subjected to graph pooling and graph normalization processing, and the obtained target information is input into the initial graph attention encoder for encoding processing to obtain the encoding result, and then the encoding result is input into the initial decoder. It can be understood that the initial decoder is an initial masked autoregressive decoder, thereby obtaining multiple initial paths.

[0172] Optionally, the processes of graph pooling, graph normalization, encoding and decoding are consistent with those in the above-mentioned embodiments, except that the parameters of the network models corresponding to the encoder and decoder are different, which will not be elaborated in the embodiments of the present application.

[0173] Step 602, determining the cumulative reward value of each initial path according to the initial path and the reward function.

[0174] Optionally, the cumulative reward value of each initial path may be determined according to each initial path and a reward function including a plurality of sub-reward functions, which may be specifically expressed by the following formula:

[0175] J (θ)= E τ~ π θ [ ∑ t=0 T γ t R( s t , a t ) ]

[0176] in, is the cumulative reward value, is the instantaneous reward at time t determined by weighted summation of each sub-reward function, Represents the parameters of the initial graph attention encoder and initial decoder, that is, the parameters of the policy network; Represent action and state trajectories; is a discount factor used to balance the importance of current rewards and future rewards; T is the maximum length of the trajectory.

[0177] Step 603: Train the initial graph attention encoder and the initial decoder according to the accumulated reward values ​​and the policy gradient algorithm until the convergence condition is met.

[0178] Optionally, the parameters of the initial graph attention encoder and initial decoder can be determined based on the cumulative reward values The gradient of the graph is calculated, and then the network model parameters are updated according to the calculated gradient. The above steps 601 to 603 are repeated until the convergence condition is met, and the target graph attention network is determined, where the convergence condition can be that the number of iterations reaches an upper limit, or the performance of the graph attention network is no longer improved.

[0179] Optionally, multi-size graph joint training can be used during the training process. In different iterations, the number of fixed nodes is different, that is, the nodes in the graph structure data of the cluster are different. For example, taking the number of fixed nodes as 50, when the number of iterations is n, 30 nodes can be selected as nodes in the graph structure data, and when the number of iterations is n+1, 40 nodes can be selected as nodes in the graph structure data. In this way, by being exposed to graph structure data of different scales during the training process, the target graph attention network can learn more universal feature representations and have stronger generalization ability and adaptability. At the same time, graph structure data are easily affected by various factors, and graph structure data of different sizes may exhibit different characteristics when facing interference factors. Through multi-scale joint training, the target graph attention network can learn to deal with interference at different scales, thereby providing the target graph attention network with robustness and being able to work stably in various complex environments.

[0180] For example, Figure 7 As shown, optionally, the initial graph attention network is trained according to each cumulative reward value and the policy gradient algorithm, including the following steps 701 to 703. Among them:

[0181] Step 701, for each initial path, the policy gradient of the initial path is calculated according to the accumulated reward value and the baseline function value of the initial path.

[0182] Optionally, a low-variance baseline strategy can be used to reduce the instability of the training process.

[0183] For example, by subtracting the baseline value from the immediate reward, the contribution of actions whose immediate rewards are close to the baseline in the gradient calculation is reduced, thereby reducing the variance, which can avoid training failure or instability caused by excessive variance.

[0184] Optionally, the baseline function is a function that is independent of the action, and may be a state value function or other estimated value. For example, the baseline function in the embodiment of the present application may be expressed by the following formula:

[0185] b( s t )=E a t ~ π θ [ R s t , a t ]

[0186] Optionally, for each initial path, the policy gradient determined by the cumulative reward value and the baseline function value of the initial path can be expressed by the following formula:

[0187]

[0188] Step 702: Determine the average gradient according to the policy gradients of each initial path.

[0189] Optionally, after determining the policy gradient of each initial path according to the above policy gradient formula, multiple policy gradients are averaged to obtain an average gradient.

[0190] Step 703, update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

[0191] Optionally, the process of updating the parameters of the initial graph attention network according to the gradient ascent method and the average gradient can be expressed by the following formula:

[0192]

[0193] in, is the updated network parameter value, is the current network parameter value, is the average gradient, is the learning rate, which controls the magnitude of parameter updates in each iteration.

[0194] Optional, such as Figure 8 As shown in the figure, it is a training flowchart for training the initial graph attention network according to the multi-objective reinforcement learning algorithm. The training process is weakly supervised learning optimization.

[0195] The above training process based on the multi-objective reinforcement learning algorithm generates multiple optimal solutions for the same initial problem in each training round through a multi-strategy parallel mechanism, and uses these solutions as training signals to update the initial graph attention network. This can avoid the problem that traditional reinforcement learning methods are prone to falling into local optimality, and significantly improve the efficiency and effect of strategy optimization by utilizing the symmetry and diversity of the solution space in combinatorial optimization problems.

[0196] As an optional implementation, Fig. 9 As shown, the cluster path planning method provided in the embodiment of the present application may include the following specific steps:

[0197] Step 901, generating multiple initial paths according to the graph structure data and the initial graph attention network;

[0198] Step 902, determining the cumulative reward value of each initial path according to the initial path and the reward function;

[0199] The reward function includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function and an information collection efficiency sub-reward function;

[0200] Step 903, for each initial path, the policy gradient of the initial path is calculated according to the accumulated reward value and the baseline function value of the initial path;

[0201] Step 904, determining an average gradient according to the policy gradients of each initial path;

[0202] Step 905, updating the parameters of the initial graph attention network according to the gradient ascent method and the average gradient until the convergence condition is met, and obtaining the target graph attention network, which includes a graph attention encoder and a mask autoregressive decoder;

[0203] Step 906, obtaining graph structure data of a cluster, where the cluster includes a plurality of underwater vehicles, the nodes of the graph structure data are used to represent the location information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles;

[0204] Step 907, calculating the importance score of each node according to the feature representation, activation function and weight parameter of each node;

[0205] Step 908, performing graph pooling processing on the graph structure data according to each importance score;

[0206] Step 909 , performing graph normalization processing on each node according to the feature representation of each node, the node feature mean and the node feature standard deviation of the graph structure data after graph pooling processing, and determining the target information of each node.

[0207] Step 910, encoding the target information according to the graph attention encoder to determine the encoding result of each node;

[0208] Step 911, obtaining dynamic mask information, the dynamic mask information including the path nodes that have been visited, the task scheduling time constraint and the energy consumption constraint;

[0209] Step 912: for each underwater vehicle, the encoding result and the dynamic mask information are input into a masked autoregressive decoder, and the target path of the underwater vehicle is determined according to the multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

[0210] For example, by comparing the path planning results obtained by existing methods (such as heuristic algorithms) and the open source software OR-tools for solving optimization problems, when using 200 fixed submarine sensors, the implementation of this application reduces the scheduling route length and reasoning time by about 66.84% and 89.23% respectively compared with the heuristic algorithm. Fig.10is the relationship between the average path planning time and the number of fixed nodes, wherein 1001 is the average path planning time of the heuristic algorithm under different fixed nodes, 1002 is the average path planning time of OR-tools under different fixed nodes, and 1003 is the average path planning time of the method provided in the embodiment of the present application under different fixed nodes; Fig.11 1101 is the multi-objective optimization value of the heuristic algorithm at different fixed nodes, 1002 is the multi-objective optimization value of OR-tools at different fixed nodes, and 1003 is the multi-objective optimization value of the method provided in the embodiment of the present application at different fixed nodes.

[0211] For example, Fig.12 and Fig.13 As shown, it is the routing result of path planning according to the embodiment of the present application, the horizontal axis is the x direction, and the vertical axis is the y direction. Fig.12 This is the path planning result when the number of AUVs is 1 and the number of seafloor fixed sensors is 50. Fig.13 The path planning result is determined when the number of AUVs is 7 and the number of seabed fixed sensors is 50. It can be seen from the figure that the embodiment of the present application shows good adaptability in graph structures of different sizes.

[0212] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0213] Based on the same inventive concept, the embodiment of the present application also provides a cluster path planning device for implementing the cluster path planning method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more cluster path planning device embodiments provided below can refer to the limitations of the cluster path planning method above, and will not be repeated here.

[0214] In an exemplary embodiment, Fig.14As shown, a cluster path planning device 1400 is provided, including: an acquisition module 1401 and a planning module 1402, wherein:

[0215] An acquisition module 1401 is used to acquire graph structure data of a cluster, where the cluster includes a plurality of underwater vehicles, the nodes of the graph structure data are used to represent the location information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles;

[0216] The planning module 1402 is used to determine the target path of each underwater vehicle based on the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling constraint sub-reward function, and an information collection efficiency sub-reward function.

[0217] In one embodiment, the target graph attention network includes a graph attention encoder and a graph decoder, and a planning module 1402, which is specifically used to perform graph pooling and graph normalization processing on the graph structure data to determine the target information; encode the target information according to the graph attention encoder to determine the encoding result of each node; and determine the target path of each underwater vehicle according to the encoding result and the graph decoder.

[0218] In one embodiment, the initial graph attention network includes an initial graph attention encoder and an initial decoder, and the cluster path planning device 1400 also includes a training module, which is specifically used to generate multiple initial paths based on graph structure data and the initial graph attention network; determine the cumulative reward value of each initial path based on the initial path and the reward function; and train the initial graph attention encoder and the initial decoder based on each cumulative reward value and the policy gradient algorithm until the convergence condition is met.

[0219] In one embodiment, the planning module 1402 is specifically used to calculate the importance score of each node based on the feature representation, activation function and weight parameters of each node; perform graph pooling on the graph structure data according to each importance score; perform graph normalization on each node based on the feature representation of each node, the node feature mean and node feature standard deviation of the graph structure data after graph pooling, and determine the target information of each node.

[0220] In one embodiment, the graph decoder is a masked autoregressive decoder, and the planning module 1402 is specifically used to obtain dynamic mask information, which includes path nodes that have been visited, task scheduling time constraints, and energy consumption constraints; for each underwater vehicle, the encoding result and the dynamic mask information are input into the masked autoregressive decoder, and the target path of the underwater vehicle is determined according to multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

[0221] In one of the embodiments, the training module is specifically used to calculate the policy gradient of each initial path based on the cumulative reward value and the baseline function value of the initial path; determine the average gradient based on the policy gradient of each initial path; and update the parameters of the initial graph attention network based on the gradient ascent method and the average gradient.

[0222] Each module in the above cluster path planning device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module above.

[0223] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.15 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a cluster path planning method is implemented.

[0224] Those skilled in the art will understand that Fig.15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0225] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps described in any of the above method embodiments when executing the computer program.

[0226] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps described in any of the above method embodiments are implemented.

[0227] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the steps described in any of the above method embodiments when executed by a processor.

[0228] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0229] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0230] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A cluster path planning method, characterized in that: The method comprises: Acquire graph structure data of a cluster, the cluster comprising a plurality of underwater vehicles, the nodes of the graph structure data being used to represent position information of each of the underwater vehicles, and the edges connecting the nodes in the graph structure data being used to represent communication links between the underwater vehicles; The target path of each underwater vehicle is determined according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function and an information collection efficiency sub-reward function.

2. The method according to claim 1, characterized in that The target graph attention network includes a graph attention encoder and a graph decoder, and determines the target path of each underwater vehicle according to the graph structure data and the target graph attention network, including: Performing graph pooling and graph normalization processing on the graph structure data to determine target information; Encoding the target information according to the graph attention encoder to determine the encoding result of each node; The target path of each of the underwater vehicles is determined according to the encoding result and the graph decoder.

3. The method according to claim 1 or 2, characterized in that: The initial graph attention network includes an initial graph attention encoder and an initial decoder, and the training process of the target graph attention network includes: Generating a plurality of initial paths according to the graph structure data and the initial graph attention network; Determine the cumulative reward value of each of the initial paths according to the initial paths and the reward function; The initial graph attention encoder and the initial decoder are trained according to each of the cumulative reward values ​​and the policy gradient algorithm until a convergence condition is met.

4. The method according to claim 2, characterized in that: The performing graph pooling and graph normalization processing on the graph structure data to determine target information includes: Calculate the importance score of each node according to the feature representation, activation function and weight parameter of each node; Performing graph pooling processing on the graph structure data according to each of the importance scores; Graph normalization is performed on each of the nodes according to the feature representation of each of the nodes, the node feature mean value and the node feature standard deviation of the graph structure data after graph pooling processing, and the target information of each of the nodes is determined.

5. The method according to claim 2, characterized in that: The graph decoder is a masked autoregressive decoder, and determining the target path of each underwater vehicle according to the encoding result and the graph decoder includes: Acquire dynamic mask information, wherein the dynamic mask information includes path nodes that have been visited, task scheduling time constraints, and energy consumption constraints; For each of the underwater vehicles, the encoding result and the dynamic mask information are input into the masked autoregressive decoder, and the target path of the underwater vehicle is determined according to multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

6. The method according to claim 3, characterized in that The training of the initial graph attention network according to each of the cumulative reward values ​​and the policy gradient algorithm includes: For each initial path, calculating the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path; Determine the average gradient based on the policy gradient of each initial path; Update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

7. A cluster path planning device, characterized in that: The device comprises: an acquisition module, configured to acquire graph structure data of a cluster, wherein the cluster includes a plurality of underwater vehicles, the nodes of the graph structure data are used to represent the position information of each of the underwater vehicles, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles; A planning module is used to determine the target path of each underwater vehicle based on the graph structure data and the target graph attention network, wherein the target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling constraint sub-reward function, and an information collection efficiency sub-reward function.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Cooperative disinfection robot control method and system based on reinforcement learning

    CN115933639A

  • Water surface unmanned cluster route planning method based on multi-agent reinforcement learning

    CN116501069A

  • Missile time sequence planning method based on graph neural network

    CN116894392A

  • Multi-robot task planning method based on graph neural network and reinforcement learning

    CN116900539A

  • Unmanned ship track generation method based on graph neural network and deep reinforcement learning

    CN116952235A

Cited By

  • Cboth case generation method and device based on large model, equipment and medium

    CN120257948A

  • Airport intelligent robot navigation method and system

    CN120427011A

  • Robot, power management method and device thereof, and program product

    CN120941453A