Cluster Path Planning Method, Device, Equipment, Storage Medium and Program Product

Through multi-objective optimization reinforcement learning algorithm, the target graph attention network is trained, combined with the graph attention encoder and decoder, the problem of quality and speed is solved in dynamic environments of underwater robot cluster path planning, and efficient and robust path generation is achieved.

CN120008618BActive Publication Date: 2025-07-08TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510488425.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-08
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the path planning algorithm of underwater robot clusters is difficult to take into account both the solution quality and the calculation speed in a large-scale and dynamically changing environment, and the heuristic algorithm is poorly adaptable.

Method used

The multi-objective optimization reinforcement learning algorithm is used to train the target graph attention network. Through four sub-reward functions, total energy consumption, energy distribution differences, task scheduling time constraints and information acquisition efficiency, the weighted sub-reward function is designed to achieve dynamic balance of path planning, and the graph attention encoder and decoder are used to generate the target path.

Benefits of technology

It improves the quality of efficiency and reconciliation of path planning, can achieve global optimization in complex and non-stationary environments, and enhances the robustness and adaptability of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120008618B_ABST
    Figure CN120008618B_ABST
Patent Text Reader

Abstract

The present application relates to a cluster path planning method, apparatus, device, storage medium and program product. The method includes: obtaining graph structure data of a cluster, where the cluster includes multiple underwater vehicles, nodes of the graph structure data are used to represent the position information of each underwater vehicle, and edges connecting the nodes in the graph structure data are used to represent communication links between the underwater vehicles; determining the target paths of the underwater vehicles according to the graph structure data and a target graph attention network, the target graph attention network is obtained by training an initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function, and an information collection efficiency sub-reward function. Using this method can take into account both the solution quality and the calculation speed during path planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of path planning, and in particular, to a method, device, equipment, storage medium, and program product for cluster path planning. Background Art

[0002] In the face of the complex marine hydrographic environment and the dangerous situation of underwater manual operations, underwater robots have gradually become a reliable way to explore the ocean. In the information collection task, the research on the path planning algorithm for underwater robot swarms has also gradually deepened.

[0003] In the prior art, the cooperative path planning problem of underwater robot swarms can be abstracted as a combinatorial optimization problem, and the path planning problem can be solved by heuristic algorithms, such as heuristic algorithms based on graph search, sampling-based heuristic algorithms, or bio-inspired algorithms.

[0004] However, heuristic algorithms have poor adaptability to large-scale and dynamically changing path planning problems, and it is difficult to balance the solution quality and computational speed. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, equipment, storage medium, and program product for cluster path planning that can balance the solution quality and computational speed.

[0006] In a first aspect, the present application provides a method for cluster path planning, including:

[0007] Obtain the graph structure data of the cluster. The cluster includes multiple underwater vehicles. The nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between each underwater vehicle;

[0008] Determine the target path of each underwater vehicle according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function, and an information collection efficiency sub-reward function.

[0009] In one embodiment, the target graph attention network includes a graph attention encoder and a graph decoder. Determining the target path of each underwater vehicle according to the graph structure data and the target graph attention network includes:

[0010] Perform graph pooling processing and graph normalization processing on the graph structure data to determine the target information; encode the target information according to the graph attention encoder to determine the encoding results of each node; determine the target paths of each underwater vehicle according to the encoding results and the graph decoder.

[0011] In one embodiment, the initial graph attention network includes an initial graph attention encoder and an initial decoder, and the training process of the target graph attention network includes:

[0012] Generate multiple initial paths according to the graph structure data and the initial graph attention network; determine the cumulative reward values of each initial path according to the initial path and the reward function; train the initial graph attention encoder and the initial decoder according to each cumulative reward value and the policy gradient algorithm until the convergence condition is met.

[0013] In one embodiment, performing graph pooling processing and graph normalization processing on the graph structure data to determine the target information includes:

[0014] Calculate the importance scores of each node according to the feature representations of each node, the activation function, and the weight parameters; perform graph pooling processing on the graph structure data according to each importance score; perform graph normalization processing on each node according to the feature representations of each node, the node feature mean, and the node feature standard deviation of the graph structure data after graph pooling processing to determine the target information of each node.

[0015] In one embodiment, the graph decoder is a masked autoregressive decoder, and determining the target paths of each underwater vehicle according to the encoding results and the graph decoder includes:

[0016] Obtain dynamic mask information, where the dynamic mask information includes the path nodes that have been visited, the task scheduling time constraint, and the energy consumption constraint; for each underwater vehicle, input the encoding result and the dynamic mask information into the masked autoregressive decoder, and determine the target path of the underwater vehicle according to the multiple candidate paths output by the masked autoregressive decoder and the corresponding probability distribution.

[0017] In one embodiment, training the initial graph attention network according to each cumulative reward value and the policy gradient algorithm includes:

[0018] For each initial path, calculate the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path; determine the average gradient according to the policy gradients of each initial path; update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

[0019] In a second aspect, the present application also provides a cluster path planning device, including:

[0020] An acquisition module for acquiring the graph structure data of a cluster, where the cluster includes multiple underwater vehicles, the nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles;

[0021] A planning module for determining the target paths of each underwater vehicle according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling constraint sub-reward function, and an information acquisition efficiency sub-reward function.

[0022] In one embodiment, the target graph attention network includes a graph attention encoder and a graph decoder. The planning module is specifically configured to perform graph pooling processing and graph normalization processing on the graph structure data to determine target information; encode the target information according to the graph attention encoder to determine the encoding results of each node; and determine the target paths of each underwater vehicle according to the encoding results and the graph decoder.

[0023] In one embodiment, the initial graph attention network includes an initial graph attention encoder and an initial decoder. The cluster path planning device further includes a training module. The training module is specifically configured to generate multiple initial paths according to the graph structure data and the initial graph attention network; determine the cumulative reward values of each initial path according to the initial paths and the reward function; and train the initial graph attention encoder and the initial decoder according to each cumulative reward value and the policy gradient algorithm until the convergence condition is met.

[0024] In one embodiment, the planning module is specifically configured to calculate the importance scores of each node according to the feature representations of each node, the activation function, and the weight parameters; perform graph pooling processing on the graph structure data according to each importance score; and perform graph normalization processing on each node according to the feature representations of each node, the node feature mean, and the node feature standard deviation of the graph structure data after graph pooling processing to determine the target information of each node.

[0025] In one embodiment, the graph decoder is a masked autoregressive decoder. The planning module is specifically configured to obtain dynamic mask information, where the dynamic mask information includes the path nodes that have been visited, the task scheduling time constraint, and the energy consumption constraint; for each underwater vehicle, input the encoding result and the dynamic mask information into the masked autoregressive decoder, and determine the target path of the underwater vehicle according to the multiple candidate paths and the corresponding probability distributions output by the masked autoregressive decoder.

[0026] In one embodiment, the training module is specifically configured to, for each initial path, calculate the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path; determine the average gradient according to the policy gradients of the initial paths; and update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

[0027] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the method described in any one of the first aspects above is implemented.

[0028] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the first aspects above is implemented.

[0029] In a fifth aspect, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspects above is implemented.

[0030] For the above cluster path planning method, device, equipment, storage medium and program product, the graph structure data of the cluster is obtained. The cluster includes multiple underwater vehicles, the nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles; according to the graph structure data and the target graph attention network, the target paths of each underwater vehicle are determined. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function, and an information collection efficiency sub-reward function. By training the graph attention network through the multi-objective optimization reinforcement learning algorithm, in this way, by designing the weighted sub-reward function, multiple optimization objectives such as minimizing the total energy consumption, energy balance constraint, task scheduling constraint, and maximizing the information collection efficiency are incorporated into a unified framework, realizing the dynamic balance between different objectives. The graph attention network trained in this way can have the global optimization ability and robustness of path generation in a complex and non-stationary environment, thereby improving the efficiency and solution quality of path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description in the embodiments of the present application or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a schematic flowchart of the cluster path planning method in an embodiment;

[0033] Figure 2 It is a schematic diagram of the correspondence between the local coordinate system and the global coordinate system in an embodiment;

[0034] Figure 3 It is a schematic flowchart of the steps of determining the target paths of each underwater vehicle according to the graph structure data and the target graph attention network in an embodiment;

[0035] Figure 4 It is a schematic flowchart of the steps of performing graph pooling processing and graph normalization processing on the graph structure data in an embodiment;

[0036] Figure 5 It is a schematic flowchart of the steps of determining the target paths of each underwater vehicle according to the encoding result and the graph decoder in an embodiment;

[0037] Figure 6 It is a schematic flowchart of the training steps of the target graph attention network in an embodiment;

[0038] Figure 7 It is a schematic flowchart of the training steps of the target graph attention network in another embodiment;

[0039] Figure 8 It is a schematic flowchart of the training process in an embodiment;

[0040] Figure 9 It is a schematic flowchart of the cluster path planning method in another embodiment;

[0041] Figure 10 It is a schematic diagram of the relationship between the average path planning time and the number of fixed nodes in an embodiment;

[0042] Figure 11 It is a schematic diagram of the relationship between the multi-objective optimization value and the number of nodes in an embodiment;

[0043] Figure 12 It is a schematic diagram of the routing result of path planning in an embodiment;

[0044] Figure 13 It is a schematic diagram of the routing result of path planning in another embodiment;

[0045] Figure 14 It is a structural block diagram of a cluster path planning device in an embodiment;

[0046] Figure 15 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0047] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] Facing the complex marine hydrological environment and the dangerous situation of underwater manual operations, underwater robots have gradually become a reliable way to explore the ocean. Under the information collection task, the research on the path planning algorithm of underwater robot clusters has gradually deepened.

[0049] In the prior art, the cooperative path planning problem of underwater robot clusters can be abstracted as a combinatorial optimization problem, and the path planning problem can be solved by heuristic algorithms, such as heuristic algorithms based on graph search, sampling-based heuristic algorithms or bio-inspired algorithms.

[0050] However, heuristic algorithms have poor adaptability to large-scale and dynamically changing path planning problems and are difficult to balance the solution quality and calculation speed.

[0051] In view of this, the embodiments of the present application provide a cluster path planning method that can balance the solution quality and calculation speed. The cluster path planning method provided by the embodiments of the present application may be executed by a cluster path planning device. The cluster path planning device may be implemented in a software, hardware, or a combination of software and hardware manner, and may be embedded in or independent of a processor in a computer device in a hardware form, or may be stored in a memory of the computer device in a software form. In the following method embodiments, the computer device is taken as an example of the execution subject for description. The computer device may be a server or a desktop computer. The embodiments of the present application do not limit the specific type of the computer device.

[0052] In an exemplary embodiment, as Figure 1 shown, a cluster path planning method is provided, including the following steps 101 to 102. Among them:

[0053] Step 101, obtaining graph structure data of a cluster. The cluster includes multiple underwater vehicles. The nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles.

[0054] Optionally, in an underwater environment, an autonomous underwater vehicle (AUV) can collect underwater information through sensors carried by itself and communicate with other AUVs to complete the underwater information collection task.

[0055] Optionally, the graph-structured data further includes a plurality of seabed fixed sensors. The nodes of the graph structure are further used to represent the position information of the plurality of seabed fixed sensors, and the edges connecting the nodes are further used to represent the communication links between each AUV and each seabed fixed sensor.

[0056] Optionally, the sensor node network can be represented by a graph structure. The topological structure of the graph structure of the cluster has dynamics and uncertainty, which is closely related to various factors in the underwater environment (such as water flow, ocean weather, equipment failure, etc.).

[0057] Exemplarily, assume that the graph-structured data consists of N nodes. Each node i (i = 1, 2, …, N) has a basic attribute position p i =(x i , y i ). The basic attribute position can represent the position of the AUV in the local coordinate system. Each node further includes a basic attribute capacity C i . The basic attribute capacity can represent the amount of information carried by the sensors of the AUV at the current moment. In the graph-structured data, the edges between nodes represent the connection links between them. The weight w ij of the edge is usually the communication cost considering interference such as ocean currents. For simplicity of analysis, assume that the connection links between each pair of nodes are stable within a specific time period, and finally the fully connected weighted undirected graph can be represented as: G = (V, E), where V = {v1, v2, …, v n} is the set of sensor nodes, is the set of edges representing the connection links. At time t, the state of each AUV can be , where is the position of the AUV at time t, is its consumed energy.

[0058] Optionally, the goal of the AUV cluster is to dynamically adjust the task assignment and path planning of the cluster members according to the task requirements and energy constraints.

[0059] Step 102: Determine the target paths of each underwater vehicle according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: total energy consumption sub-reward function, energy distribution difference sub-reward function, task scheduling time constraint sub-reward function, and information acquisition efficiency sub-reward function.

[0060] Optionally, the initial graph attention network can be trained through a multi-objective optimization reinforcement learning algorithm to obtain the target graph attention network. In this way, after the obtained graph structure data is input into the target graph attention network, the target graph attention network can input the actions of each underwater vehicle at the next moment. For example, the forward direction and forward speed or forward acceleration of the underwater vehicle, etc.

[0061] Optionally, in the deep-sea Internet, when the cluster collaboratively executes information acquisition tasks, each AUV in the cluster needs to cooperate to maximize the information acquisition efficiency and overcome challenges including energy consumption, task allocation, communication limitations, etc. The main characteristics of the AUV cluster include the sparsity of task scheduling, the dynamic changes of the environment, energy limitations, and the timeliness of information, etc. At the same time, the high dynamicity of the marine environment also poses serious challenges to the Markov property of the path planning method. Therefore, underwater cluster path planning can be regarded as a data-driven weakly supervised learning Markov decision problem, and the characteristics of the reinforcement learning algorithm can be used to solve it.

[0062] Among them, the sparsity of the above task scheduling may refer to that different tasks may be more important for some AUVs in the cluster, resulting in a sparse distribution of information acquisition tasks in time and space. Therefore, the AUV must make dynamic adjustments according to the importance of the current task and the energy situation. The dynamic changes of the environment may mean that factors such as water flow, swell, and weather in the marine environment will constantly change, affecting the path planning and communication quality of the AUV. This requires the AUV cluster to have the ability to adapt to environmental changes in real time. The energy limitation and the timeliness of information may refer to that the energy of each AUV is limited, and it is necessary to reasonably allocate energy when performing tasks to ensure that the tasks can be completed and returned. At the same time, the timeliness of information requires the AUV to complete the acquisition task as quickly as possible and return the information in time.

[0063] Optionally, the Graph Attention Network (GAT) is a deep learning model based on graph-structured data, focusing on processing graph data with non-Euclidean structures. By introducing an attention mechanism, GAT dynamically assigns weights between graph nodes, enabling it to more effectively highlight the influence of key nodes when aggregating neighborhood information. Compared with traditional Graph Convolutional Networks (GCNs), GAT has stronger flexibility and expressive power, especially performing excellently in tasks involving dynamic graphs, sparse graphs, and large-scale graph data processing.

[0064] Optionally, the multi-objective optimization reinforcement learning algorithm guides the AUV to learn the optimal strategy through multiple sub-reward functions. In the multi-objective optimization reinforcement learning algorithm, the path planning problem corresponding to the cluster information collection task can be regarded as a Markov Decision Process (MDP). The unmanned submersible learns the strategy through interaction and optimizes the long-term cumulative reward. In this process, the sub-reward functions are designed as multiple independent optimization metrics, each corresponding to an optimization goal, and are weighted and combined through adjustable weight coefficients to achieve a dynamic balance between different goals. Under the MDP framework, the underwater vehicle makes decisions in the underwater environment with state and affects the environment by taking action to obtain reward and enters the next state . This process has obvious Markov properties, that is, the future state depends only on the current state and the current action, and is independent of the historical state.

[0065] Optionally, a Markov decision process can be defined by a five-tuple where, represents the state space, that is, the set of states of the underwater vehicle and the environment; A represents the action space, that is, the set of decisions that the underwater vehicle can execute; represents the state transition probability, that is, the probability distribution of transferring to the new state after taking action a in state s; R{s,a} represents the reward function, that is, the immediate reward obtained under the current state and action is defined; γ ∈ (0, 1] represents the discount factor, that is, it controls the weight of future rewards in the overall reward. For the state space , the information collection task of the underwater vehicle involves multiple key state variables, and the state space can be expressed as:

[0066]

[0067] wherein, is the position information of the i-th AUV; is the remaining energy of the i-th AUV at present; is the information acquisition progress of the i-th AUV, that is, the information quantity distribution of the undersea fixed sensor nodes; is the task scheduling status of the i-th AUV, including the task completion situation and the remaining time.

[0068] Optionally, when the underwater vehicle executes a task, it needs to consume energy for propulsion, sensor operation and data communication. Therefore, the total energy consumption is a key factor affecting the task duration and overall performance. The total energy consumption sub-reward function of the AUV can be expressed by the following formula:

[0069]

[0070] wherein, N is the total number of AUVs, M is the total number of other AUVs connected to the i-th AUV. The total energy consumption of each AUV can include the processes such as moving against ocean currents and communication. The energy consumption for moving against ocean currents and the communication energy consumption will be introduced hereinafter.

[0071] Optionally, in the cluster task, the unbalanced consumption of individual energy may cause some underwater vehicles to fail in advance, affecting the overall task completion situation. To measure the distribution divergence in the data group, the difference or divergence between two probability distributions can be quantified through a metric method.

[0072] Exemplarily, the energy consumption distribution of different underwater vehicles can be measured through the Kullback-Leibler divergence (KL divergence) to make it as close as possible to an ideal balanced distribution. Obtain the remaining energy distribution of each AUV as P(x), representing the probability distribution of the remaining energy of the AUV, and the distribution Q(x) of the best energy distribution of all AUVs in the cluster. The energy distribution difference sub-reward function can be expressed by the following formula:

[0073]

[0074] Optionally, to ensure that each task is completed within the specified time and the operation time of each AUV does not exceed its energy limit, that is , is the time required for the i-th underwater vehicle to collect the information of the j-th undersea fixed node, that is, the task completion time. The task scheduling time constraint sub-reward function can be expressed as:

[0075]

[0076]

[0077] Among them, is a preset coefficient, is a timeout penalty function, which depends on whether the time window constraint of the task is guaranteed and the degree of exceeding the time limit.

[0078] Optionally, in order to determine to complete as many information collection tasks as possible under limited time and resources, the information collection efficiency, that is, the sub-reward function corresponding to the information collection success rate can be represented by the ratio of the amount of valid information collected to the total amount of task target information:

[0079]

[0080] Among them, is the amount of valid information collected by the i-th AUV, is the total amount of task target information.

[0081] Optionally, in the multi-objective optimization reinforcement learning algorithm, maximizing the cumulative discounted reward can be represented by the following formula:

[0082]

[0083] Among them, is the immediate reward in the Markov decision process at time t, that is, the reward function, which is determined by weighted summation of each sub-reward function, that is , among which, is an adjustable weight parameter used to control the relative importance between different optimization objectives. It can be understood that in maximizing the cumulative discounted reward, it is necessary to satisfy the minimization of total energy consumption, the minimization of energy distribution difference, the satisfaction of the time window constraint by the task scheduling time, and the maximization of information collection efficiency.

[0084] Optionally, during cluster path planning, it is necessary to adapt to environmental changes, perform dynamic path planning according to the real-time energy state of the AUV, the task completion situation, and environmental changes (such as water flow speed, communication delay), etc., and optimize multiple objectives simultaneously.

[0085] Next, the energy consumption against ocean currents and the communication energy consumption will be introduced:

[0086] Exemplarily, under the assumption of graph structure modeling, that is, two-dimensional modeling, the six-degree-of-freedom motion of the AUV can be further simplified to the motion in the plane, ignoring the vertical motion. The motion problem of the AUV can be simplified to three degrees of freedom, that is, only considering the translational (forward, lateral) and rotational (yaw) motions in the horizontal plane, asFigure 2 As shown in the figure, it is a schematic diagram of a coordinate system. Among them, \(u\) represents the forward velocity of this type of underwater robot, \(v\) is the lateral movement velocity, \(\theta\) is the pitch angle, \(\varphi\) is the barrel roll angle, \(\psi\) is the yaw angle, \(r\) represents the yaw angular velocity. There is the following relationship between the global coordinate system and the local coordinate system of the underwater vehicle:

[0087]

[0088] Among them, represents the velocity component in the \(X\)-axis direction in the global coordinate system, represents the velocity component in the \(Y\)-axis direction in the global coordinate system, represents the rate of change of the yaw angle with time.

[0089] Optionally, the dynamic model is as follows:

[0090]

[0091] Among them, \(m\) is the mass of this AUV; \(T\) is the thrust of the propeller water jet; and are the longitudinal and lateral added masses respectively; and are the moment of inertia about the center of gravity and the added moment of inertia respectively; , , , and are the linear hydrodynamic derivatives; , , , , , and are the nonlinear hydrodynamic derivatives. The hydrodynamic derivatives are usually captured by planar motion technique (PMM) tests. In particular, represents the length of the real AUV between the vertical lines, is the real and the model is the reciprocal of the scale factor between them. , and represent the wave disturbances.

[0092] Exemplarily, the dynamic system can be simplified by means of distributed modeling, decoupling the hydrodynamic model into two independent parts: the drag-free motion equation and the turbulent field influence model. Among them, the influence of the turbulent field directly acts on the velocity components of the AUV. Thus, the algorithm complexity can be reduced through model simplification. Secondly, the characteristics of the turbulent field in the research environment can be analyzed and set according to requirements. The turbulent field at the working plane can be modeled based on the world coordinate system through the classical two-dimensional Navier-Stokes equation to approximate the actual ocean situation:

[0093]

[0094]

[0095]

[0096] Among them, , and are the velocity, vorticity and viscosity of the fluid respectively, and are the gradient operator and the Laplace operator respectively, which can approximately represent a point in a certain plane under the world coordinate system at time The water flow velocity is:

[0097]

[0098]

[0099] Among them, is the vortex center, and are the intensity and the influence range radius of the vortex respectively. Since the influence of the horizontal turbulent field on the AUV dominates, the resistance suffered by the AUV can be approximately deduced by means of Computational Fluid Dynamics (CFD) as:

[0100]

[0101] Among them, is the mass density of the fluid, is the longitudinal cross-sectional area of the AUV, is the drag coefficient, is the relative velocity between the ocean current and the AUV. From this, the translational and rotational motion energy consumption of the AUV within the time slot can be deduced:

[0102]

[0103]

[0104] Among them, is the electrotransformation efficiency, is the length of the AUV, is the mass of the AUV.

[0105] Exemplarily, for the communication consumption caused by exchanging information it can be determined with the aid of the classical Thorp model, specifically as follows:

[0106] The attenuation coefficient of an underwater acoustic signal with frequency f at a distance l in an underwater acoustic channel can be expressed as:

[0107]

[0108] Among them, A0 is the unit normalization constant, and k is the propagation coefficient, is the absorption coefficient.

[0109] The path loss (dB) is given by the following formula:

[0110]

[0111] Among them, for the right - hand side of the above equation, the first term represents the propagation loss, and the second term represents the absorption loss. The constant k is usually between 2 and 4. Some of its classical values depend on different situations of underwater acoustic signal propagation: when the underwater acoustic signal is spherically diffused, k = 2; when it is cylindrically diffused, k = 1. In the embodiments of the present application, considering the actual underwater acoustic propagation situation, k = 1.5 is taken.

[0112] When the acoustic frequency is kHz, the above equation can be simplified to the following equation:

[0113]

[0114] For low - frequency signals, the following simpler equation form can also be used for approximation:

[0115]

[0116] In addition to considering the underwater path loss, when propagating underwater acoustic signals, three noise sources often need to be considered to simulate signal noise, namely turbulence, ships, waves, and thermal noise. And the total noise power can be expressed as

[0117]

[0118] Among them represents the total power spectral density of the ambient noise, represents the coefficient by which the noise decreases with frequency.

[0119] According to Shannon's theorem, the data transmission rate of the AUV information output is defined as R c , assuming that the maximum transmission power of the AUV is , then the maximum communication distance in the plane can be obtained . At the maximum communication efficiency, the communication consumption of the i-th AUV can be obtained as follows:

[0120]

[0121] where is the total number of AUVs in the working plane, B is the bandwidth, , H represents the total power spectral density of the ambient noise, represents the coefficient by which the noise decreases with frequency, is an indicator function. When the decision condition in the parentheses is satisfied, the function value is 1; otherwise, the function value is 0.

[0122] The above-mentioned cluster path planning method obtains the graph structure data of the cluster. The cluster includes multiple underwater vehicles. The nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between each underwater vehicle; the target paths of each underwater vehicle are determined according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling time constraint sub-reward function, and an information acquisition efficiency sub-reward function. By training the graph attention network through a multi-objective optimization reinforcement learning algorithm, in this way, by designing weighted sub-reward functions, multiple optimization objectives such as minimizing the total energy consumption, energy balance constraint, task scheduling constraint, and maximizing the information acquisition efficiency are incorporated into a unified framework, realizing the dynamic balance between different objectives. The graph attention network trained in this way can have the global optimization ability and robustness of path generation in a complex and non-stationary environment, thereby improving the efficiency and quality of path planning.

[0123] In an exemplary embodiment, as Figure 3 shown, optionally, the target graph attention network includes a graph attention encoder and a graph decoder. Determining the target paths of each underwater vehicle according to the graph structure data and the target graph attention network includes the following steps 301 to step 303. Wherein:

[0124] Step 301, perform graph pooling processing and graph normalization processing on the graph structure data to determine the target information.

[0125] Optionally, graph pooling can be a pooling technique for graph-structured data, which can gradually compress the graph-structured data into a fixed size, thereby reducing the computational complexity and extracting higher-level graph-structured information. Through a preset node selection strategy or feature aggregation strategy, graph pooling can compress the graph-structured data into a fixed-size representation without destroying the graph structure and retain key global and local information.

[0126] Exemplarily, graph pooling processing can be performed by evaluating node importance. Important nodes are selected based on the importance scores of each node, and finally, a new coarsened graph is constructed using these important nodes; alternatively, nodes can be mapped into several clusters according to similarity or other criteria through node clustering. The attention mechanism can also be combined during the clustering process to assign clustering weights to each node. Nodes with higher weights are more likely to become the representative nodes of the clusters, thereby retaining the important information of the nodes in the pooled graph structure.

[0127] Optionally, graph normalization can be a normalization technique for graph-structured data, which can solve the problem of inconsistent node feature distributions in graph neural networks. Since graph-structured data usually has different scales and structures, traditional normalization methods may not be able to effectively adapt to these differences when dealing with graph-structured data. Therefore, through graph normalization, the part of node features can be stabilized, the influence of the graph structure size on features can be reduced, and the stability and convergence speed of model training can be improved.

[0128] Optionally, the target information can be the graph structure after graph pooling processing and graph normalization processing, as well as the feature representations of each node in the graph structure.

[0129] Exemplarily, as Figure 4 shown, optionally, performing graph pooling processing and graph normalization processing on the graph-structured data to determine the target information includes the following steps 401 to 403. Wherein:

[0130] Step 401, calculate the importance scores of each node according to the feature representations, activation functions, and weight parameters of each node.

[0131] Optionally, for each node in the graph-structured data, its importance score can be represented by the following formula :

[0132]

[0133] Wherein, represents the feature representation of node i, W and b are learnable parameters, is the activation function, is the global normalization task weight parameter, is the learnable task importance and attention mechanism weight ratio.

[0134] Step 402: Perform graph pooling on the graph structure data according to each importance score.

[0135] Optionally, important nodes in the graph structure data can be selected according to the importance scores. For example, obtain a preset importance score threshold, and use the nodes with importance scores greater than the preset importance score threshold as important nodes, or sort the nodes according to the importance scores and use the nodes with the top several importance scores as important nodes.

[0136] Optionally, a new coarsened graph can be formed according to the important nodes.

[0137] Step 403: Perform graph normalization on each node according to the feature representation of each node, the node feature mean value, and the node feature standard deviation of the graph structure data after graph pooling, and determine the target information of each node.

[0138] Optionally, for each node i in the graph structure G, its normalized feature can be represented by the following formula:

[0139]

[0140] where is the original feature of the node, is the mean value of the graph-level node features after graph pooling, is the standard deviation of the graph-level features after graph pooling, and are preset coefficients.

[0141] Through the above graph pooling and graph normalization, the graph structure can be adjusted in real time according to the changes of nodes in the task to adapt to the continuously changing environmental topology, and it has better robustness when dealing with tasks of different scales, ensuring that the target graph attention network can still work stably in a dynamic environment.

[0142] Step 302: Encode the target information according to the graph attention encoder to determine the encoding results of each node.

[0143] Optionally, the encoding results of each node can be feature vectors obtained after aggregating the information of neighbor nodes through the attention mechanism.

[0144] Optionally, the graph attention encoder can be a target graph attention network, and its backbone network mainly includes attention weight calculation, feature update, and multi-head attention mechanism.

[0145] Exemplarily, GAT can assign dynamic weights to the neighborhood of each node through a learnable attention mechanism. Specifically, GAT can calculate the attention weights through the following formula:

[0146]

[0147] Among them, is the attention weight between node i and its neighbor node j; and are the feature representations of nodes i and j respectively; W is a learnable feature transformation matrix; a is the weight vector of the attention mechanism; | represents the concatenation operation of features; LeakyRELU is the activation function.

[0148] The calculated attention weight can be used to weighted aggregate the features of neighboring nodes, thereby generating an updated node representation:

[0149]

[0150] Among them, is a non-linear activation function, and usually ReLU can be used.

[0151] Optionally, the characteristics of the dynamic changes of nodes in the task scheduling process can be combined to calculate the attention weight calculation. By increasing the temporal dynamic modeling ability of the node importance weight, it can more flexibly cope with the dynamics and randomness of the deep-sea environment.

[0152] To improve the robustness and stability of the model, GAT introduces a multi-head attention mechanism, which concatenates or averages the outputs of multiple independent attention heads to obtain a richer node representation:

[0153]

[0154] Among them, k represents the k-th head, || represents the concatenation operation, represents the attention weight between node i and its neighbor node j corresponding to the k-th head, represents the learnable feature transformation matrix corresponding to the k-th head.

[0155] Optionally, a multi-scale attention mechanism can also be introduced. By integrating the node features within different ranges, it can further capture the global features and local correlations of the deep-sea environment tasks, and improve the modeling ability of GAT for sparse targets.

[0156] The above GAT has the advantages of strong self-adaptability, efficient aggregation, and enhanced representational power by the multi-head mechanism. The attention mechanism enables GAT to dynamically adjust the influence between nodes according to the different characteristics of graph data, and is especially suitable for sparse graphs and irregular graphs. GAT automatically selects the importance of neighbor nodes through attention weights, thereby more accurately capturing global and local information. Through the multi-head attention mechanism, GAT can learn feature representations from multiple different perspectives, thereby improving the robustness of the model.

[0157] Step 303. Determine the target paths of each underwater vehicle according to the encoding result and the graph decoder.

[0158] Optionally, the graph decoder can be a neural network module, which can convert the encoding result into a target path through the graph decoder, that is, the actions of each underwater vehicle at the next moment.

[0159] Exemplarily, the graph decoder can be a recurrent neural network decoder, a Transformer decoder, an autoregressive decoder, etc., which are not limited in the embodiments of the present application.

[0160] Performing the above graph pooling processing and graph normalization processing on the graph structure data to determine the target information, encoding the target information according to the graph attention encoder to determine the encoding results of each node, and determining the target paths of each underwater vehicle according to the encoding results and the graph decoder can improve the adaptability and robustness when determining the target paths.

[0161] In an exemplary embodiment, as Figure 5 shown, optionally, the graph decoder is a masked autoregressive decoder. Determining the target paths of each underwater vehicle according to the encoding result and the graph decoder includes the following steps 501 to 502. Wherein:

[0162] Step 501. Obtain dynamic mask information, where the dynamic mask information includes the path nodes that have been visited, the task scheduling time constraint, and the energy consumption constraint.

[0163] Optionally, the path nodes that have been visited can be obtained from the movement trajectories of each underwater vehicle. It can be understood that the movement trajectories of each underwater vehicle are continuously updated as the path planning progresses.

[0164] Optionally, by obtaining the task scheduling time constraint and the energy consumption constraint, invalid target paths can be filtered out in the subsequent decoding process. For example, if a certain path does not meet the task scheduling time constraint and the energy consumption constraint, that is, the path is unavailable, it can be excluded through the mask information, so that the decoder focuses on the valid path information.

[0165] Step 502. For each underwater vehicle, input the encoding result and the dynamic mask information into the masked autoregressive decoder, and determine the target path of the underwater vehicle according to the multiple candidate paths and the corresponding probability distributions output by the masked autoregressive decoder.

[0166] Optionally, when determining the target path of each underwater vehicle, the encoded result of the node and the dynamic mask information can be input into the masked autoregressive decoder, so as to predict the probability distribution from the current node to other reachable nodes. This probability distribution represents the possibility of choosing different paths to continue moving forward in the current state.

[0167] Optionally, when determining the target path of the underwater vehicle according to multiple candidate paths and the corresponding probability distribution, the candidate path with the highest probability can be selected as the target path, or a candidate path can be randomly selected from multiple candidate paths according to a random algorithm. Alternatively, multiple paths with relatively high probabilities can be selected at each step, that is, a certain number of optimal paths are retained, and then these selected paths are continuously expanded and evaluated in subsequent steps, and finally the path with the highest score is selected as the target path.

[0168] By obtaining the dynamic mask information as described above, the dynamic mask information includes the path nodes that have been visited, the task scheduling time constraint, and the energy consumption constraint. For each underwater vehicle, the encoded result and the dynamic mask information are input into the masked autoregressive decoder, and the target path of the underwater vehicle is determined according to the multiple candidate paths and the corresponding probability distribution output by the masked autoregressive decoder. It is possible to query the dynamics of the cluster at each decision time step, that is, during the target path generation process, by masking the generated path nodes, path duplication is avoided and the algorithm efficiency is improved. At the same time, the autoregressive decoding method is used to dynamically output the target path, so that the path generation process can flexibly respond to the complex changes of the environmental dynamics, task execution dynamics, and underwater environment dynamics.

[0169] In an exemplary embodiment, as Figure 6 shown, optionally, the initial graph attention network includes an initial graph attention encoder and an initial decoder. The training process of the target graph attention network includes the following steps 601 to step 603. Among them:

[0170] Step 601, generate multiple initial paths according to the graph structure data and the initial graph attention network.

[0171] Optionally, the obtained graph structure data can be first subjected to graph pooling processing and graph normalization processing, and the obtained target information is input into the initial graph attention encoder for encoding processing to obtain an encoded result, and then the encoded result is input into the initial decoder. It can be understood that the initial decoder is an initial masked autoregressive decoder, so as to obtain multiple initial paths.

[0172] Optionally, the processes of graph pooling processing, graph normalization processing, encoding processing, and decoding processing are the same as those in the above embodiment, except that the parameter values of the network models corresponding to the encoder and the decoder are different, and the embodiments of the present application will not elaborate on this.

[0173] Step 602, determining the cumulative reward value of each initial path according to the initial path and the reward function.

[0174] Optionally, the cumulative reward value of each initial path may be determined according to each initial path and a reward function including a plurality of sub-reward functions, which may be specifically expressed by the following formula:

[0175] J (θ) = E τ ~ π θ [ ∑ t=0 T γ t R( s t , a t ) ]

[0176] in, is the cumulative reward value, is the instantaneous reward at time t determined by weighted summation of each sub-reward function, Represents the parameters of the initial graph attention encoder and initial decoder, that is, the parameters of the policy network; Represent action and state trajectories; is a discount factor used to balance the importance of current rewards and future rewards; T is the maximum length of the trajectory.

[0177] Step 603: Train the initial graph attention encoder and the initial decoder according to the accumulated reward values ​​and the policy gradient algorithm until the convergence condition is met.

[0178] Optionally, the parameters of the initial graph attention encoder and initial decoder can be determined based on the cumulative reward values The gradient of the graph is calculated, and then the network model parameters are updated according to the calculated gradient. The above steps 601 to 603 are repeated until the convergence condition is met, and the target graph attention network is determined, where the convergence condition can be that the number of iterations reaches an upper limit, or the performance of the graph attention network is no longer improved.

[0179] Optionally, multi-scale graph joint training can be adopted during the training process. During different iteration processes, the number of fixed nodes is different, that is, the nodes in the graph structure data of the cluster are different. For example, taking the fixed number of nodes as 50 as an example, when the number of iterations is n, 30 nodes can be selected as the nodes in the graph structure data, and when the number of iterations is n + 1, 40 nodes can be selected as the nodes in the graph structure data. In this way, during the training process, the target graph attention network can learn more general feature representations by being exposed to graph structure data of different scales, and has stronger generalization ability and adaptability. At the same time, graph structure data is easily affected by various factors, and graph structure data of different sizes may exhibit different characteristics when facing interference factors. Through multi-scale joint training, the target graph attention network can learn to handle interference in different sizes, improve the robustness of the target graph attention network, and enable it to work stably in various complex environments.

[0180] Exemplarily, such as Figure 7 As shown, optionally, the initial graph attention network is trained according to each cumulative reward value and the policy gradient algorithm, including the following steps 701 to step 703. Wherein:

[0181] Step 701, for each initial path, calculate the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path.

[0182] Optionally, the low-variance baseline strategy can be adopted to reduce the instability during the training process.

[0183] Exemplarily, by subtracting the baseline value from the immediate reward, the contribution of actions whose immediate rewards are close to the baseline in the gradient calculation is reduced, thereby reducing the variance and avoiding problems such as training failure or instability caused by excessive variance.

[0184] Optionally, the baseline function is a function independent of the action, and can be the state value function or other estimated values. For example, the baseline function in the embodiments of the present application can be represented by the following formula:

[0185] b( s t )=E a t ~ π θ [ R s t , a t ]

[0186] Optionally, for each initial path, the policy gradient determined by the cumulative reward value and the baseline function value of the initial path can be represented by the following formula:

[0187]

[0188] Step 702: Determine the average gradient according to the policy gradients of each initial path.

[0189] Optionally, after determining the policy gradients of each initial path according to the above policy gradient formula, taking the average of multiple policy gradients can obtain the average gradient.

[0190] Step 703: Update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

[0191] Optionally, the process of updating the parameters of the initial graph attention network according to the gradient ascent method and the average gradient can be represented by the following formula:

[0192]

[0193] where, is the updated network parameter value, is the current network parameter value, is the average gradient, is the learning rate, which controls the amplitude of parameter update in each iteration.

[0194] Optionally, as Figure 8 shown, it is the training flow chart for training the initial graph attention network according to the multi-objective reinforcement learning algorithm, and the training process is weak supervision learning optimization.

[0195] The above training process based on the multi-objective reinforcement learning algorithm, through the multi-policy parallel mechanism, generates multiple optimal solutions for the same initial problem in each training episode, and uses these solutions as training signals to update the initial graph attention network, which can avoid the problem that traditional reinforcement learning methods are prone to fall into local optima, and significantly improve the efficiency and effect of policy optimization by utilizing the symmetry and diversity of the solution space in the combinatorial optimization problem.

[0196] As an optional implementation manner, as Figure 9 shown, the cluster path planning method provided in the embodiments of the present application may include the following specific steps:

[0197] Step 901: Generate multiple initial paths according to the graph structure data and the initial graph attention network;

[0198] Step 902: Determine the cumulative reward value of each initial path according to the initial path and the reward function;

[0199] where, the reward function includes multiple sub-reward functions; the multiple sub-reward functions include: total energy consumption sub-reward function, energy distribution difference sub-reward function, task scheduling time constraint sub-reward function, and information collection efficiency sub-reward function;

[0200] Step 903: For each initial path, calculate the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path.

[0201] Step 904: Determine the average gradient according to the policy gradients of the initial paths.

[0202] Step 905: Update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient until the convergence condition is met, and obtain the target graph attention network. The target graph attention network includes a graph attention encoder and a masked autoregressive decoder.

[0203] Step 906: Obtain the graph structure data of the cluster. The cluster includes multiple underwater vehicles. The nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles.

[0204] Step 907: Calculate the importance scores of the nodes according to the feature representations of the nodes, the activation function, and the weight parameters.

[0205] Step 908: Perform graph pooling processing on the graph structure data according to the importance scores.

[0206] Step 909: Perform graph normalization processing on the nodes according to the feature representations of the nodes, the mean of the node features, and the standard deviation of the node features in the graph structure data after graph pooling processing, and determine the target information of the nodes.

[0207] Step 910: Encode the target information according to the graph attention encoder to determine the encoding results of the nodes.

[0208] Step 911: Obtain the dynamic mask information, which includes the path nodes that have been visited, the task scheduling time constraint, and the energy consumption constraint.

[0209] Step 912: For each underwater vehicle, input the encoding result and the dynamic mask information into the masked autoregressive decoder, and determine the target path of the underwater vehicle according to the multiple candidate paths and the corresponding probability distributions output by the masked autoregressive decoder.

[0210] Exemplarily, by comparing with the path planning results obtained by existing methods (such as heuristic algorithms) and the open-source software OR-tools for solving optimization problems, when using 200 fixed underwater sensors, the implementation of this application reduces the scheduling route length and the inference time by approximately 66.84% and 89.23% respectively compared with the heuristic algorithm. Figure 10It is the relationship between the average path planning time and the number of fixed nodes. Among them, 1001 is the average path planning time of the heuristic algorithm under different fixed nodes, 1002 is the average path planning time of OR-tools under different fixed nodes, and 1003 is the average path planning time of the method provided in the embodiment of the present application under different fixed nodes; Figure 11 It is the relationship between the multi-objective optimization value and the number of nodes. Among them, 1101 is the multi-objective optimization value of the heuristic algorithm under different fixed nodes, 1002 is the multi-objective optimization value of OR-tools under different fixed nodes, and 1003 is the multi-objective optimization value of the method provided in the embodiment of the present application under different fixed nodes.

[0211] Exemplarily, such as Figure 12 and Figure 13 As shown, it is the routing result of path planning according to the embodiment of the present application. The abscissa is the x direction and the ordinate is the y direction. Figure 12 It is the path planning result determined when the number of AUVs is 1 and the number of undersea fixed sensors is 50. Figure 13 It is the path planning result determined when the number of AUVs is 7 and the number of undersea fixed sensors is 50. It can be seen from the figure that the embodiment of the present application shows good adaptability in graph structures of different sizes.

[0212] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0213] Based on the same inventive concept, the embodiment of the present application also provides a cluster path planning device for implementing the above-mentioned cluster path planning method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the following cluster path planning devices can refer to the limitations on the cluster path planning method in the above text, and will not be repeated here.

[0214] In an exemplary embodiment, such as Figure 14As shown, a cluster path planning device 1400 is provided, including: an acquisition module 1401 and a planning module 1402, where:

[0215] The acquisition module 1401 is configured to acquire the graph structure data of the cluster. The cluster includes multiple underwater vehicles. The nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles.

[0216] The planning module 1402 is configured to determine the target paths of each underwater vehicle according to the graph structure data and the target graph attention network. The target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm. The reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling constraint sub-reward function, and an information acquisition efficiency sub-reward function.

[0217] In one embodiment, the target graph attention network includes a graph attention encoder and a graph decoder. The planning module 1402 is specifically configured to perform graph pooling processing and graph normalization processing on the graph structure data to determine the target information; perform encoding processing on the target information according to the graph attention encoder to determine the encoding results of each node; and determine the target paths of each underwater vehicle according to the encoding results and the graph decoder.

[0218] In one embodiment, the initial graph attention network includes an initial graph attention encoder and an initial decoder. The cluster path planning device 1400 further includes a training module. The training module is specifically configured to generate multiple initial paths according to the graph structure data and the initial graph attention network; determine the cumulative reward values of each initial path according to the initial paths and the reward function; and train the initial graph attention encoder and the initial decoder according to each cumulative reward value and the policy gradient algorithm until the convergence condition is met.

[0219] In one embodiment, the planning module 1402 is specifically configured to calculate the importance scores of each node according to the feature representations of each node, the activation function, and the weight parameters; perform graph pooling processing on the graph structure data according to each importance score; and perform graph normalization processing on each node according to the feature representations of each node, the mean of the node features, and the standard deviation of the node features of the graph structure data after graph pooling processing to determine the target information of each node.

[0220] In one embodiment, the graph decoder is a masked autoregressive decoder. The planning module 1402 is specifically configured to obtain dynamic masking information, where the dynamic masking information includes the path nodes that have been visited, task scheduling time constraints, and energy consumption constraints. For each underwater vehicle, the encoded result and the dynamic masking information are input into the masked autoregressive decoder, and the target path of the underwater vehicle is determined according to the multiple candidate paths and the corresponding probability distributions output by the masked autoregressive decoder.

[0221] In one embodiment, the training module is specifically configured to, for each initial path, calculate the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path; determine the average gradient according to the policy gradients of the initial paths; and update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

[0222] Each module in the above cluster path planning device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above modules.

[0223] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, a cluster path planning method is implemented.

[0224] Those skilled in the art can understand that Figure 15 the structure shown in

[0225] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps described in any one of the above method embodiments are implemented.

[0226] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps described in any one of the above method embodiments are implemented.

[0227] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps described in any one of the above method embodiments are implemented.

[0228] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0229] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.

[0230] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for cluster path planning, characterized in that, The method includes: Obtaining the graph structure data of the cluster, where the cluster includes multiple underwater vehicles, the nodes of the graph structure data are used to represent the position information of each underwater vehicle, and the edges connecting the nodes in the graph structure data are used to represent the communication links between the underwater vehicles; Determining the target paths of each underwater vehicle according to the graph structure data and the target graph attention network, where the target graph attention network is obtained by training the initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions; the multiple sub-reward functions include: total energy consumption sub-reward function, energy distribution difference sub-reward function, task scheduling time constraint sub-reward function, and information acquisition efficiency sub-reward function; The initial graph attention network includes an initial graph attention encoder and an initial decoder, and the training process of the target graph attention network includes: Generating multiple initial paths according to the graph structure data and the initial graph attention network; Determining the cumulative reward values of each initial path according to the initial path and the reward function; Training the initial graph attention encoder and the initial decoder according to each cumulative reward value and the policy gradient algorithm until the convergence condition is met.

2. The method according to claim 1, wherein The target graph attention network includes a graph attention encoder and a graph decoder. Determining the target paths of each underwater vehicle according to the graph structure data and the target graph attention network includes: Performing graph pooling processing and graph normalization processing on the graph structure data to determine target information; Encoding the target information according to the graph attention encoder to determine the encoding results of each node; Determining the target paths of each underwater vehicle according to the encoding results and the graph decoder.

3. The method according to claim 2, wherein The performing graph pooling processing and graph normalization processing on the graph structure data to determine target information includes: Calculating the importance scores of each node according to the feature representations of each node, activation function, and weight parameters; Performing graph pooling processing on the graph structure data according to each importance score; Performing graph normalization processing on each node according to the feature representations of each node, the node feature mean, and the node feature standard deviation of the graph structure data after graph pooling processing to determine the target information of each node.

4. The method according to claim 2, wherein The graph decoder is a masked autoregressive decoder. Determining the target paths of each underwater vehicle according to the encoding results and the graph decoder includes: Obtaining dynamic mask information, where the dynamic mask information includes the path nodes that have been visited, task scheduling time constraints, and energy consumption constraints; For each underwater vehicle, inputting the encoding results and the dynamic mask information into the masked autoregressive decoder, and determining the target path of the underwater vehicle according to the multiple candidate paths and the corresponding probability distributions output by the masked autoregressive decoder.

5. The method according to claim 1, characterized in that The training the initial graph attention network according to each cumulative reward value and the policy gradient algorithm includes: For each initial path, calculate the policy gradient of the initial path according to the cumulative reward value and the baseline function value of the initial path; Determine the average gradient according to the policy gradients of the initial paths; Update the parameters of the initial graph attention network according to the gradient ascent method and the average gradient.

6. A cluster path planning device, characterized in that, The device includes: An acquisition module, configured to acquire graph structure data of a cluster, where the cluster includes multiple underwater vehicles, nodes of the graph structure data are used to represent the position information of each underwater vehicle, and edges connecting the nodes in the graph structure data are used to represent communication links between the underwater vehicles; A planning module, configured to determine target paths of the underwater vehicles according to the graph structure data and a target graph attention network, where the target graph attention network is obtained by training an initial graph attention network through a multi-objective optimization reinforcement learning algorithm, and the reward function of the multi-objective optimization reinforcement learning algorithm includes multiple sub-reward functions: a total energy consumption sub-reward function, an energy distribution difference sub-reward function, a task scheduling constraint sub-reward function, and an information acquisition efficiency sub-reward function; A training module, where the initial graph attention network includes an initial graph attention encoder and an initial decoder. The training module is configured to generate multiple initial paths according to the graph structure data and the initial graph attention network; determine the cumulative reward value of each initial path according to the initial path and the reward function; train the initial graph attention encoder and the initial decoder according to each cumulative reward value and the policy gradient algorithm until a convergence condition is met.

7. The device according to claim 6, characterized in that, The target graph attention network includes a graph attention encoder and a graph decoder. The planning module is specifically configured to perform graph pooling processing and graph normalization processing on the graph structure data to determine target information; encode the target information according to the graph attention encoder to determine the encoding results of the nodes; and determine the target paths of the underwater vehicles according to the encoding results and the graph decoder.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Cooperative disinfection robot control method and system based on reinforcement learning

    CN115933639A