Marine cross-domain unmanned system cooperative path planning method, system and device

By decomposing the autonomous collaboration problem of cross-domain unmanned marine systems into two stages—task allocation and path planning—and employing reinforcement learning and optimization algorithms, the problems of high solution complexity and local optima in traditional methods are solved, achieving more efficient communication capacity and computational performance.

CN120506959BActive Publication Date: 2025-11-04TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511009492.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

Traditional rule-based decision-making methods for cross-domain unmanned marine systems suffer from high solution complexity, local optima, and low collaboration efficiency, making it difficult to effectively coordinate path planning among multiple agents.

Method used

The problem of autonomous collaboration of unmanned platforms is decomposed into a task allocation problem and a path planning problem. Through a two-stage cross-domain collaboration method, the task allocation results of the autonomous underwater vehicle cluster are first obtained, and then the target path is determined based on the state information of each unmanned vehicle and the path planning model. Reinforcement learning algorithm and optimization algorithm are used to maximize channel capacity.

Benefits of technology

It improves the collaborative path planning performance of marine cross-domain unmanned systems, with higher stability and lower solution difficulty, and achieves more efficient communication capacity and computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120506959B_ABST
    Figure CN120506959B_ABST
Patent Text Reader

Abstract

The application relates to a marine cross-domain unmanned system cooperative path planning method, system and device. The application belongs to the technical field of path planning. The method comprises the following steps: obtaining a task allocation result of an autonomous underwater vehicle cluster, wherein the task allocation result comprises information transmission frequency bands of respective autonomous underwater vehicles, and the information transmission frequency bands are used for representing communication links between the autonomous underwater vehicles and an unmanned ship cluster; obtaining state information of respective unmanned ships in the unmanned ship cluster, wherein the state information of the unmanned ships comprises position information of the unmanned ship cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned ships and target autonomous underwater vehicles, and the target autonomous underwater vehicles are autonomous underwater vehicles which establish communication links with the unmanned ships according to the task allocation result; and determining target paths of the respective unmanned ships according to the state information of the respective unmanned ships and a path planning model, wherein the target paths are used for representing paths when the channel capacity of the unmanned ships is maximized. The method can effectively improve the performance of the cooperative path planning method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of path planning, in particular to a marine cross-domain unmanned system cooperative path planning method, system and device. BACKGROUND

[0002] With the development of artificial intelligence, unmanned and clustered marine observation is a development trend of future intelligent marine industrialization deployment, and the decision-making ability of multiple agents is an important indicator of the autonomy of cross-domain unmanned systems. Taking a marine cross-domain unmanned system as an example, the system includes multiple unmanned surface vehicles (USVs) and multiple autonomous underwater vehicles (AUVs), wherein the AUVs are used to perform marine observation and data collection, and the USVs can provide stable and reliable communication services for the AUVs and assist in completing the detection tasks of the AUVs.

[0003] The traditional rule-based model decision-making method has problems such as high complexity, local optimization and low cooperation efficiency when dealing with the problem of multi-agent cross-domain cooperation, and the performance of the decision-making method needs to be improved. SUMMARY

[0004] Therefore, it is necessary to provide a marine cross-domain unmanned system cooperative path planning method, system and device capable of improving the performance of the cooperative path planning method.

[0005] In a first aspect, the application provides a marine cross-domain unmanned system cooperative path planning method, which is used for a marine cross-domain unmanned system cooperative path planning system, the marine cross-domain unmanned system cooperative path planning system includes an autonomous underwater vehicle cluster and an unmanned surface vehicle cluster, and the method includes the following steps.

[0006] Obtaining a task allocation result of the autonomous underwater vehicle cluster, the task allocation result including information transmission frequency bands of each autonomous underwater vehicle, the information transmission frequency band being used to represent a communication link between the autonomous underwater vehicle and the unmanned surface vehicle cluster;

[0007] Obtaining state information of each unmanned surface vehicle in the unmanned surface vehicle cluster, the state information of the unmanned surface vehicle including position information of the unmanned surface vehicle cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned surface vehicle and a target autonomous underwater vehicle, the target autonomous underwater vehicle being an autonomous underwater vehicle determined according to the task allocation result to establish a communication link with the unmanned surface vehicle;

[0008] Determining a target path of each unmanned surface vehicle according to the state information of each unmanned surface vehicle and a path planning model, the target path being used to represent a path when the channel capacity of the unmanned surface vehicle is maximized.

[0009] In a second aspect, the application further provides a marine cross-domain unmanned system cooperative path planning system, comprising an autonomous underwater vehicle cluster and an unmanned surface vehicle cluster.

[0010] The autonomous underwater vehicle cluster is configured to obtain a task allocation result of the autonomous underwater vehicle cluster, the task allocation result comprising information transmission frequency bands of the autonomous underwater vehicles, the information transmission frequency bands being used to represent communication links between the autonomous underwater vehicles and the unmanned surface vehicle cluster.

[0011] The unmanned surface vehicle cluster is configured to obtain state information of each unmanned surface vehicle in the unmanned surface vehicle cluster, the state information of the unmanned surface vehicle comprising position information of the unmanned surface vehicle cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned surface vehicle, and a target autonomous underwater vehicle, the target autonomous underwater vehicle being an autonomous underwater vehicle determined to establish a communication link with the unmanned surface vehicle according to the task allocation result; and determine a target path of each unmanned surface vehicle according to the state information of each unmanned surface vehicle and a path planning model, the target path being used to represent a path when the channel capacity of the unmanned surface vehicle is maximized.

[0012] In a third aspect, the application further provides an unmanned surface vehicle, comprising a memory and a processor, the memory storing a computer program, and the processor implementing steps of the method of any one of the first aspect when executing the computer program.

[0013] In a fourth aspect, the application further provides an autonomous underwater vehicle, comprising a memory and a processor, the memory storing a computer program, and the processor implementing steps of the method of any one of the first aspect when executing the computer program.

[0014] In a fifth aspect, the application further provides a computer readable storage medium, which stores a computer program, and the computer program implements steps of the method of any one of the first aspect when executed by a processor.

[0015] In a sixth aspect, the application further provides a computer program product, comprising a computer program, and the computer program implements steps of the method of any one of the first aspect when executed by a processor.

[0016] The method, system and device for cooperative path planning of the marine cross-domain unmanned system first obtain a task allocation result of the autonomous underwater vehicle cluster, the task allocation result including information transmission frequency bands of each autonomous underwater vehicle, the information transmission frequency band being used to represent a communication link between the autonomous underwater vehicle and the unmanned vehicle cluster; then, state information of each unmanned vehicle in the unmanned vehicle cluster is obtained, the state information of the unmanned vehicle including position information of the unmanned vehicle cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned vehicle and a target autonomous underwater vehicle, the target autonomous underwater vehicle being an autonomous underwater vehicle that establishes a communication link with the unmanned vehicle according to the task allocation result; and then, a target path of each unmanned vehicle is determined according to the state information of each unmanned vehicle and a path planning model, the target path being used to represent a path when the channel capacity of the unmanned vehicle is maximized. In the method, the autonomous cooperation problem of the unmanned platform is decomposed into a task allocation problem and a path planning problem, the target path of the unmanned vehicle is determined according to the task allocation result after the task allocation result is confirmed, and through the two-stage cross-domain cooperation method, the types of intelligent agents are the same and the tasks performed by the intelligent agents are the same in each stage, so that the solving process has higher stability, lower solving difficulty and lower calculation cost, and the performance of the cooperative path planning method can be effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without any creative effort.

[0018] Figure 1 An application environment diagram of the method for cooperative path planning of the marine cross-domain unmanned system in an embodiment;

[0019] Figure 2 A flowchart of the method for cooperative path planning of the marine cross-domain unmanned system in an embodiment;

[0020] Figure 3 A flowchart of the step of determining the task allocation result in an embodiment;

[0021] Figure 4 A flowchart of the training step of the task allocation model in an embodiment;

[0022] Figure 5 A flowchart of the training step of the path planning model in an embodiment;

[0023] Figure 6 A flowchart of the method for cooperative path planning of the marine cross-domain unmanned system in another embodiment;

[0024] Figure 7 A task allocation result diagram in an embodiment;

[0025] Figure 8 A reward curve convergence result comparison diagram in an embodiment when training in the second stage based on different reinforcement learning algorithms;

[0026] Figure 9 A three-dimensional trajectory diagram of an AUV and a SUV in an embodiment;

[0027] Figure 10 A communication capacity reward convergence result comparison diagram with training time when using different methods for collaborative path planning in an embodiment;

[0028] Figure 11 A communication capacity reward convergence result comparison diagram with training time when using different numbers of agents in an embodiment;

[0029] Figure 12 An internal structure diagram of an unmanned surface vehicle in an embodiment;

[0030] Figure 13 An internal structure diagram of an autonomous underwater vehicle in an embodiment. DETAILED DESCRIPTION

[0031] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0032] The ocean-crossing unmanned system collaborative path planning method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 The ocean-crossing unmanned system considers constructing a cluster of autonomous underwater vehicles and a cluster of unmanned surface vehicles, that is, N USVs and M AUVs. The AUVs are used to perform ocean observation and data collection tasks, and the USVs can provide stable and reliable communication services for the AUVs and assist in completing the detection tasks of the AUVs. As shown in Figure 1As shown, the motion of the AUV and the USV is considered ideal, the USV can realize free motion in a two-dimensional plane, and the AUV can complete three-dimensional motion underwater, and the maximum speed of the USV is higher than that of the AUV. In this application, ideal communication transmission can be realized between the USVs, that is, the communication rate meets the information transmission requirements, the AUV and the USV rely on underwater acoustic communication technology to transmit information, and the USV mainly cooperates by collecting data of the AUV. The AUV in the above-mentioned marine cross-domain unmanned system can be connected to any USV, each USV can be connected to k AUVs, and all AUVs need to be connected. Therefore, under the premise of meeting the stability of communication, the transmission interference between AUVs, the communication rate and other constraints need to be considered to construct the optimal trajectory planning scheme of the USV under multiple constraints.

[0033] Based on the above-mentioned marine cross-domain unmanned system, the process of constructing a spread spectrum multi-user communication model and a marine underwater acoustic channel model is described below.

[0034] Optionally, in the multi-user communication model, in order to maximize the information transmission efficiency between the USV and the AUV, while ensuring the stability of multiple communication links, a code division multiple access mode is used to realize multi-user access in a single subsystem, and a frequency division multiple access mode is used to avoid mutual interference for different subsystems. The spread spectrum underwater acoustic communication realizes robust multi-user data transmission from multiple AUVs to a single USV by assigning each AUV a unique and mutually orthogonal pseudo-random sequence. The baseband form of the transmission signal corresponding to the mth AUV is:

[0035]

[0036] wherein, is the ith data symbol sent by the mth AUV, is the symbol length, and satisfies , is the spreading factor, is the symbol width, is the corresponding spreading code sequence.

[0037] The received signal of the nth USV can be expressed as:

[0038]

[0039] wherein, is the number of AUVs connected to the nth USV, and , is the received signal amplitude, is the channel impulse response function from the mth AUV to the nth USV, is the received environmental noise.

[0040] Data demodulation can be achieved by performing matched filtering on the received signal and the spread spectrum sequence of the corresponding AUV. Taking the first AUV as an example, the matched filtering output of the i-th symbol is expressed as:

[0041]

[0042] in, Corresponding to the i-th symbol of the received signal, Let be the cross-correlation function between the spreading sequence corresponding to the 1st AUV and the spreading sequence corresponding to the mth AUV. Due to the orthogonality between different spreading sequences, when hour, The result is very small, ensuring accurate decoding of the data transmitted by each AUV.

[0043] Optionally, in the ocean acoustic channel model, due to the severe propagation attenuation of electromagnetic waves and light underwater, underwater acoustic communication is the most effective technology for long-distance information transmission between USVs and AUVs. The propagation loss of sound waves underwater includes spread loss and absorption loss, expressed as:

[0044]

[0045] in, Let n be the distance between the nth USV and the mth AUV. Let m be the depth of the m-th AUV. The absorption coefficient, measured in dB / km, is closely related to the sound wave frequency f and can be expressed as:

[0046]

[0047] Marine environmental noise is mainly caused by turbulence. Waves ,vessel and thermal noise Caused by, etc., therefore, marine environmental noise level It can be calculated as:

[0048]

[0049]

[0050]

[0051]

[0052]

[0053] in, This is the correlation coefficient for ship density.

[0054] According to the above, the signal-to-noise ratio of the nth USV receiving the underwater acoustic signal transmitted by the mth AUV can be expressed as:

[0055]

[0056] wherein, is the sound source level, indicating the sound pressure level of the sound source at a reference distance relative to the reference sound pressure, is the noise intensity.

[0057] The direct sequence spread spectrum technology can realize multi-user spread spectrum underwater acoustic communication with weak mutual interference. By designing a spread spectrum sequence with good orthogonality, the multiple access interference between users can be approximated as noise, and the signal-to-interference ratio (SIR) can be expressed as:

[0058]

[0059] Further, the signal-to-interference-plus-noise ratio (SINR) of the nth USV receiving the signal can be expressed as:

[0060]

[0061] Therefore, the communication capacity between the nth USV and the mth AUV received by the nth USV can be expressed as:

[0062]

[0063] wherein, is the communication bandwidth corresponding to the nth USV, is the spread spectrum factor.

[0064] The total communication capacity of the cross-domain cooperative system is:

[0065]

[0066] Optionally, in order to achieve efficient transmission of data of the ocean cross-domain unmanned system and reduce other interference as much as possible, the system optimization target can be expressed as MaxC.

[0067] Optionally, when the AUVs or the USVs share information, the transmission is a data file and does not include pictures, videos or other information, so the communication between the AUVs and the USVs is considered ideal, and the transmission data meets the channel capacity and the minimum communication rate requirement.

[0068] In addition, considering that the USV accesses at least one AUV, each AUV is ensured to be accessed, so there is the following constraint:

[0069]

[0070]

[0071] wherein, represents whether the USV is connected with the AUV, 1 represents connected, and 0 represents not connected, represents the ith USV, represents the jth AUV.

[0072] In an exemplary embodiment, as shown in Figure 2 , a marine cross-domain unmanned system cooperative path planning method is provided for a marine cross-domain unmanned system cooperative path planning system, the marine cross-domain unmanned system cooperative path planning system comprising an autonomous underwater vehicle cluster and an unmanned surface vehicle cluster, and the method comprises the following steps:

[0073] Step 201, obtaining a task allocation result of the autonomous underwater vehicle cluster.

[0074] The task allocation result comprises an information transmission frequency band of each autonomous underwater vehicle, and the information transmission frequency band is used to represent a communication link between the autonomous underwater vehicle and the unmanned surface vehicle cluster.

[0075] Optionally, the autonomous underwater vehicle cluster comprises a plurality of autonomous underwater vehicles AUVs, each AUV can correspond to an information transmission frequency band, and then a USV corresponding to the information transmission frequency band is selected for data backhaul.

[0076] The task allocation result can be a transmission frequency band in which the signal quality of each autonomous underwater vehicle satisfies the maximum signal-to-interference ratio in the signal transmission process. Therefore, the process of obtaining the task allocation result of the autonomous underwater vehicle cluster can be regarded as a solution to an optimization problem, and the objective function is the maximum signal-to-interference ratio.

[0077] In a possible implementation manner, the task allocation result of the autonomous underwater vehicle cluster can be determined according to a reinforcement learning algorithm.

[0078] Optionally, the optimization problem can be constructed as a Markov decision process (MDP), represented as . Wherein represents a state space, represents an action space, is a state transition function, is a reward function, represents a discount factor.

[0079] Wherein, for the autonomous underwater vehicle cluster, the corresponding state space is S A , and the state corresponding to each autonomous underwater vehicle can be defined as S A i =[ P a , P s , T a i , T a -i ] , respectively represent the position of the AUV cluster, the position of the USV cluster, the current AUV information transmission frequency band (target USV) and the transmission frequency band corresponding to other AUVs; the corresponding action space is P s(id) =[ f 1 , f 2 ,..., f n ] , that is, selecting the USV corresponding to the frequency band to return data.

[0080] Exemplarily, the reinforcement learning algorithm can include a multi-agent reinforcement learning algorithm (Multi-Agent Proximal Policy Optimization, MAPPO), and can also be a multi-agent deep deterministic policy gradient (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) or an improved algorithm thereof, etc., such as a multi-agent twin delayed deep deterministic policy gradient algorithm (Multi-Agent Twin Delayed Deep Deterministic policy gradient algorithm, MATD3), etc., which is not limited by the embodiments of the present application.

[0081] In another possible implementation manner, the task allocation result of the autonomous underwater vehicle cluster can be determined according to an optimization algorithm such as a particle swarm algorithm, an ant colony algorithm or a genetic algorithm.

[0082] Exemplarily, taking the determination of the task allocation result of the autonomous underwater vehicle cluster according to the genetic algorithm as an example, an initial population can be randomly generated, the population is a set composed of multiple individuals, each individual represents a potential solution, and the fitness value of each individual is calculated, that is, the pros and cons of the individual are evaluated according to the optimization target, then the individual is selected according to the fitness value to generate the next generation, the process of fitness evaluation, selection, crossover and mutation is repeated until the termination condition is met, the termination condition can be that the maximum iteration number is reached, the fitness value converges or a good enough solution is found, and the solution obtained by solving can be used as the task allocation result of the autonomous underwater vehicle cluster.

[0083] In step 202, state information of each USV in the USV cluster is acquired, and the state information of the USV includes position information of the USV cluster, position information of the AUV cluster, channel capacity of the USV, and a target AUV, which is an AUV determined to establish a communication link with the USV according to a task allocation result.

[0084] Optionally, after confirming the task allocation result, the USV selected by each AUV can be determined, and for each USV, an AUV establishing a communication link with the USV is determined.

[0085] Optionally, in the USV cluster, the position information can be exchanged among the USVs, or the position information of all USVs in the cluster and the position information of all AUVs in the AUV cluster can be acquired by a sensor or other intelligent agent.

[0086] Optionally, the channel capacity in the state information of the USV represents the channel capacity of the USV based on the current task allocation result, which can be calculated according to .

[0087] In step 203, a target path of each USV is determined according to the state information of each USV and a path planning model, and the target path is used to represent a path when the channel capacity of the USV is maximized.

[0088] Optionally, the problem of determining the target path of each USV can be regarded as an optimization problem, and the optimization objective is to maximize the total channel capacity.

[0089] Optionally, the path planning model can be determined according to a reinforcement learning algorithm, or can be determined according to an optimization algorithm such as a particle swarm algorithm, an ant colony algorithm, or a genetic algorithm, which is not limited in the embodiments of the application.

[0090] Optionally, when the path planning model is determined according to the reinforcement learning algorithm, for the USV cluster, the corresponding state space is , and the state corresponding to each USV can be defined as S P i =[ P a , P s , T i , C i ] , respectively represent the position of the AUV cluster, the position of the USV cluster, the target AUV of the current USV, and the channel capacity of the current USV; and the corresponding action space is A S i =[ v x , v y ] , respectively, represent the direction vectors of the x and y coordinate axes in the two-dimensional plane. Assuming that the USV control system can complete the motion solution of the current orientation, while ensuring approximate omnidirectional motion, the USV can be made to move at a constant speed in eight orientations by simplifying the USV dynamics, by outputting a specific orientation, and by letting the USV move at a constant speed. The control of the running track of the USV, that is, if the position coordinates (x, y) of the USV are both greater than 0, and The values of and can be if the position coordinates (x, y) of the USV are both less than 0, and The values of and can be if one of the position coordinates (x, y) of the USV is 0 and the other is greater than 0, and The values of and can be and 0, if one of the position coordinates (x, y) of the USV is 0 and the other is less than 0, and The values of and can be and 0.

[0091] Exemplarily, the reinforcement learning algorithm can include a multi-agent reinforcement learning algorithm (Multi-Agent Proximal Policy Optimization, MAPPO), and can also be a multi-agent deep deterministic policy gradient (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) or an improved algorithm thereof, etc., such as a multi-agent twin delayed deep deterministic policy gradient algorithm (Multi-Agent Twin Delayed Deep Deterministic policy gradient algorithm, MATD3), etc., which are not limited by the embodiments of the present application.

[0092] In another possible implementation manner, the task allocation result of the autonomous underwater vehicle cluster can be determined according to an optimization algorithm such as a particle swarm algorithm, an ant colony algorithm or a genetic algorithm.

[0093] It should be noted that in the prior art, when the multi-agent reinforcement learning algorithm is used to simultaneously plan the cooperation of the USV cluster and the AUV cluster, different types of agents may perform different tasks, i.e., the USV on the sea surface may provide an optimal path for the AUV to return data collected as a premise, but the AUV under the water selects a suitable frequency band to send data while completing the collection in the target area. At this time, the agents cannot adopt the same strategy network to realize the decision of multiple tasks, and as the number of agents increases, the new state brought by the interaction with the environment also increases, which easily causes the original algorithm to fall into a local optimal solution.

[0094] In the embodiment of the present application, first, the task allocation result of the AUV cluster is obtained, and the task allocation result includes the information transmission frequency band of each AUV, which is used to represent the communication link between the AUV and the USV cluster. Then, the state information of each USV in the USV cluster is obtained, which includes the position information of the USV cluster, the position information of the AUV cluster, the channel capacity of the USV, and the target AUV determined according to the task allocation result, which is the AUV establishing a communication link with the USV. Then, the target path of each USV is determined according to the state information of each USV and the path planning model, and the target path is used to represent the path when the channel capacity of the USV is maximized. In the above method, the autonomous cooperation problem of the unmanned platform is decomposed into a task allocation problem and a path planning problem. After confirming the task allocation result, the target path of the USV is determined according to the task allocation result. Through the two-stage cross-domain cooperation method, the types of agents are the same and the tasks performed by the agents are the same in each stage. Therefore, the solving process has higher stability, lower solving difficulty and lower computational overhead, and the performance of the cooperative path planning method can be effectively improved.

[0095] In one exemplary embodiment, as shown in Figure 3 the task allocation result of the AUV cluster is obtained, including:

[0096] Step 301, obtaining the state information of each AUV in the AUV cluster.

[0097] The state information of the AUV includes the position information of the AUV cluster, the position information of the USV cluster, and the information transmission frequency band of each AUV.

[0098] Optionally, in the AUV cluster, each AUV cluster can exchange position information with each other, or through sensors or other agents, each AUV can obtain the position information of all AUVs in the cluster and the position information of each USV in the USV cluster.

[0099] Optionally, the information transmission frequency band of each autonomous underwater vehicle can be a USV selected by the AUV in the current state. It can be understood that, as the positions of the AUV or the USV change, the signal quality corresponding to the information transmission frequency band does not necessarily meet the preset signal threshold.

[0100] In step 302, the task allocation result is determined according to the state information of each autonomous underwater vehicle and a task allocation model.

[0101] Optionally, the state information of each autonomous underwater vehicle can be input into the task allocation model, and the task allocation result can be determined according to the output of the task allocation model.

[0102] Optionally, the task allocation model can be determined based on a reinforcement learning algorithm, and the type of reinforcement learning is not limited in the embodiments of the present application.

[0103] Optionally, taking the determination of the task allocation model based on MAPPO as an example, in actual application, the task allocation model can include a second target policy network, and the determination of the task allocation result according to the state information of each autonomous underwater vehicle and the task allocation model can include the following steps:

[0104] The state information of each autonomous underwater vehicle is input into the second target policy network to obtain the task allocation result corresponding to each autonomous underwater vehicle.

[0105] Optionally, the second target policy network can be composed of a multilayer perceptron (MLP) and include an input layer, multiple hidden layers, and an output layer.

[0106] For example, taking the second target policy network including three fully connected layers as an example, the second target policy network can generally include an input layer, two hidden layers, and an output layer, and the input layer is used to receive data without calculation.

[0107] Optionally, the state space including the state information of each autonomous underwater vehicle can be taken as input data, received by the input layer, and output an action probability distribution. The action with the maximum probability can be determined as the task allocation result corresponding to each autonomous underwater vehicle.

[0108] The above method of obtaining the state information of each autonomous underwater vehicle in the autonomous underwater vehicle cluster and determining the task allocation result according to the state information of each autonomous underwater vehicle and the task allocation model can enable each agent in the autonomous underwater vehicle cluster to accurately determine the corresponding task allocation result based on the pre-trained task allocation model and the state information of the autonomous underwater vehicle cluster.

[0109] After determining the task result, the state information of each unmanned ship can be further determined, and the target path of each unmanned ship is determined. In an exemplary embodiment, the path planning model includes a first target strategy network and a first target evaluation network. The target path of each unmanned ship is determined according to the state information of each unmanned ship and the path planning model, including:

[0110] The state information of each unmanned ship is input into the first target strategy network to obtain the corresponding target path of each unmanned ship.

[0111] Optionally, the path planning model can be determined based on a reinforcement learning algorithm. The type of reinforcement learning is not limited in the embodiments of the present application.

[0112] Optionally, the first target strategy network and the second target strategy network are two different strategy networks, and the network parameters of the two are not the same.

[0113] Optionally, the first target strategy network can also be composed of a multilayer perceptron (MLP), including an input layer, multiple hidden layers and an output layer. For example, the hidden layer of the first target strategy network can be two layers.

[0114] Optionally, when determining the corresponding target path of each unmanned ship, that is, determining the motion direction and motion speed of each unmanned ship at the next moment, the state space including the state information of each unmanned ship as input data, received by the input layer, and outputting an action probability distribution, the action space is discrete, and the action with the maximum probability can be determined as the corresponding target path of each unmanned ship.

[0115] The training process of the task allocation model and the path planning model will be described below. Since the input of the path planning model includes the target autonomous underwater vehicle determined according to the task allocation result, the task allocation model needs to be trained first. After the network parameter update of the task allocation model is completed, the initial path planning model is trained to obtain the path planning model.

[0116] In an exemplary embodiment, as Figure 4 shown, the training process of the task allocation model includes:

[0117] Step 401, generating an initial task allocation result according to an initial task allocation model and initial state information of each autonomous underwater vehicle.

[0118] Optionally, the initial state information of each autonomous underwater vehicle also includes the position information of the autonomous underwater vehicle cluster, the position information of the unmanned ship cluster and the information transmission frequency band of each autonomous underwater vehicle.

[0119] Optionally, the initial state information can be input into a second initial policy network of the initial task allocation model to obtain an initial task allocation result.

[0120] At step 402, a second reward is determined according to each initial task allocation result and a second reward function.

[0121] The second reward function is determined according to a signal quality index value and a second preset weight coefficient, and the signal quality index value is used to represent the signal quality of the underwater acoustic signal emitted by the autonomous underwater vehicle to the unmanned surface vehicle.

[0122] Optionally, in the first stage, the AUV and the USV face the same task and have the same cooperation target, so it can be considered that the second reward functions corresponding to each AUV are the same and shared when the task allocation result is determined in the first stage, which can be represented by the following formula:

[0123]

[0124] Wherein, h represents the second preset weight coefficient, and taking the signal-to-interference ratio as an example of the signal quality index value, the task allocation result is used to represent the communication link established with the USV while satisfying the signal-to-interference ratio as large as possible for the current agent k.

[0125] At step 403, the initial task allocation model is trained according to the second reward and a policy gradient algorithm until a convergence condition is met.

[0126] The initial task allocation model includes a second initial policy network and a second initial evaluation network.

[0127] Optionally, training the initial task allocation model includes training the second initial policy network to update its network parameters to obtain a second target policy network; and training the second initial evaluation network to update its network parameters to obtain a second target evaluation network.

[0128] Optionally, the second initial policy network and the second initial evaluation network can be updated alternately, that is, the second initial policy network is updated based on the advantage function estimated by the second initial evaluation network with fixed parameters, and then the second initial evaluation network is updated using the data collected by the policy network with updated network parameters.

[0129] Optionally, the policy of the agent AUVi is re-expressed as , wherein represents the network parameters of the AUVi, represents the state information of the AUVi, represents the action of the AUVi; and the target function of the current agent AUV is is denoted as:

[0130]

[0131] where t denotes the time step or episode, and the optimal policy is obtained by maximizing the sum of the expected rewards for future time steps.

[0132] Optionally, the expected value of the current policy can be effectively utilized using importance sampling, i.e., by introducing a KL divergence, denoted as:

[0133]

[0134] where denotes the old policy distribution of agent i, denotes the new policy distribution, and the KL divergence can measure the difference between the new and old policy distributions. Further, the above distribution can be modified as follows:

[0135]

[0136] where denotes the estimated policy.

[0137] In order to further limit the update range of the policy while ensuring exploration performance, a truncated method is often used to limit the update range of the policy, and its loss function is denoted as:

[0138]

[0139] where clip denotes a clipping function that returns the upper and lower limits according to importance sampling. denotes the advantage function, which is used to improve learning efficiency and enhance training stability.

[0140] Each agent AUV is considered as a common task and reward, so the update method of the second initial policy network can be set to consistent update, and further, the second initial policy network is extended to multi-agent, and its loss function can be denoted as:

[0141]

[0142] where B denotes the batch size, N denotes the number of agents AUV, and S denotes the policy entropy, is a hyperparameter.

[0143] Similarly, the loss function of the second initial evaluation network can be denoted as:

[0144] L c φ = 1 BN ∑ i=1 B · ∑ k=1 N (max[ V φ i,k - R ̂ i 2 , clip V φ i,k , V φ ' i,k -ε, V φ ' i,k +ε - R ̂ i 2 ]

[0145] wherein represents a discounted reward, respectively represent new and old state value functions.

[0146] Optionally, the convergence condition can be that the number of iterations meets a preset iteration threshold, or that the average cumulative reward fluctuates less than a preset threshold within a certain number of iterations, or that the update amplitude of the policy network parameters is less than a threshold, and the training can be stopped when the convergence condition is met.

[0147] In an exemplary embodiment, as shown in Figure 5 the training process of the path planning model includes:

[0148] Step 501, after the task allocation model is trained, the task allocation results corresponding to each autonomous underwater vehicle are determined according to the task allocation model, and the target autonomous underwater vehicle in the initial state information of each unmanned surface vehicle is determined according to the task allocation results.

[0149] Step 502, an initial path is generated according to the initial path planning model and the updated initial state information of each unmanned surface vehicle.

[0150] Step 503, a first reward is determined according to each initial path and a first reward function, and the first reward function is determined according to the channel capacity and a first preset weight coefficient.

[0151] Optionally, in the second stage, the AUV and the USV face the same task and have the same cooperation target, so it can be considered that the first reward functions corresponding to each USV are the same and shared when the target path planning is determined in the second stage, which can be represented by the following formula:

[0152]

[0153] wherein, is the channel capacity, is the boundary reward, that is, the USV needs to collect data backhaul in the designated area and change its own trajectory on the basis of the existing allocated task result to maximize the channel capacity.

[0154] Step 504, the initial path planning model including the first initial policy network and the first initial evaluation network is trained according to the first reward and the policy gradient algorithm until the convergence condition is met.

[0155] Optionally, the training process of the path planning model is similar to the training process of the task allocation model. When the initial state information is determined, the corresponding target autonomous underwater vehicle of each unmanned surface vehicle is determined according to the trained task allocation model, and then the initial path is determined.

[0156] It should be noted that, for the training of the path planning model, since it depends on the task allocation result output by the task allocation model, and in the training process, the unmanned surface vehicle moves according to the initial target path determined by the initial path planning model, the position information of the unmanned surface vehicle will change constantly. With the iteration of the training times, the task allocation result is not the optimal strategy under the current environment, therefore, the state information of the current AUV and USV can be input into the task allocation model again to determine a new task allocation result, and the initial path planning model is trained based on the new task allocation result.

[0157] In a possible implementation manner, whether the task allocation result needs to be updated can be determined based on the training state of the initial path planning model.

[0158] For example, when the first reward continues to decrease in a certain number of iterations, the task allocation result of each autonomous underwater vehicle can be updated according to the task allocation model.

[0159] In another possible implementation manner, the task allocation result of each autonomous underwater vehicle can be updated according to the task allocation model when the training times meet the preset re-allocation frequency; the initial state information of each unmanned surface vehicle is updated according to the updated task allocation result; and the initial path planning model is continuously trained according to the updated initial state information.

[0160] As an optional implementation manner, as shown in Figure 6 The ocean cross-domain unmanned system cooperative path planning method provided in the embodiments of the present application can include the following specific steps:

[0161] Step 601: generating an initial task allocation result according to an initial task allocation model and initial state information of each autonomous underwater vehicle.

[0162] Step 602: determining a second reward according to each initial task allocation result and a second reward function.

[0163] The second reward function is determined according to a signal quality index value and a second preset weight coefficient, and the signal quality index value is used to represent the signal quality of the underwater acoustic signal emitted by the autonomous underwater vehicle to the unmanned surface vehicle.

[0164] In step 603, the initial task allocation model is trained according to the second reward and the policy gradient algorithm until a convergence condition is met, and a task allocation model is obtained, the initial task allocation model comprising a second initial policy network and a second initial evaluation network.

[0165] In step 604, a task allocation result corresponding to each autonomous underwater vehicle is determined according to the task allocation model, and initial state information of each unmanned ship is determined according to the task allocation result.

[0166] In step 605, an initial path is generated according to the initial path planning model and the updated initial state information of each unmanned ship.

[0167] In step 606, a first reward is determined according to each initial path and a first reward function.

[0168] The first reward function is determined according to a channel capacity and a first preset weight coefficient.

[0169] In step 607, the initial path planning model is trained according to the first reward and the policy gradient algorithm.

[0170] In step 608, when a preset redistribution frequency is met, the task allocation result corresponding to each autonomous underwater vehicle is updated according to the task allocation model.

[0171] In step 609, the initial state information of each unmanned ship is updated according to the updated task allocation result.

[0172] In step 610, the initial path planning model is continuously trained according to the updated initial state information until a convergence condition is met, and a path planning model is obtained, the initial path planning model comprising a first initial policy network and a first initial evaluation network.

[0173] In step 611, state information of each autonomous underwater vehicle in the autonomous underwater vehicle cluster is obtained.

[0174] The state information of the autonomous underwater vehicle comprises position information of the autonomous underwater vehicle cluster, position information of the unmanned ship cluster, and an information transmission frequency band of each autonomous underwater vehicle.

[0175] In step 612, the state information of each autonomous underwater vehicle is input into the second target policy network, and a task allocation result corresponding to each autonomous underwater vehicle is obtained.

[0176] The task allocation result comprises an information transmission frequency band of each autonomous underwater vehicle, and the information transmission frequency band is used to represent a communication link between the autonomous underwater vehicle and the unmanned ship cluster.

[0177] In step 613, state information of each unmanned ship in the unmanned ship cluster is obtained.

[0178] The state information of the unmanned surface vehicle includes position information of the unmanned surface vehicle cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned surface vehicle, and a target autonomous underwater vehicle, the target autonomous underwater vehicle being an autonomous underwater vehicle determined according to a task allocation result to establish a communication link with the unmanned surface vehicle.

[0179] In step 614, the state information of each unmanned surface vehicle is input into the first target strategy network to obtain a corresponding target path of each unmanned surface vehicle.

[0180] To verify the effectiveness of the method described in the embodiments of the present application, a cross-domain cooperation scenario containing more than 3 USVs and more than 9 AUVs can be set, where the moving range of the USV is [-500m, 500m], and the operating depth of the AUV is [0, 500m]. Based on the Torch2.0.1+cu117 deep learning framework, experiments are carried out on an Intel-Core(TM) i7-13700KF CPU and a GeForce RTX 4070 GPU. It is assumed that the AUV moves randomly in three-dimensional space at a constant speed, and an underwater camera is configured to collect information, and at the same time, local autonomous obstacle avoidance can be completed through a multi-beam sonar; in addition, since the operating range of the AUV is wide, the movement of the AUV in three-dimensional space is considered to be an approximate point model. In the design of the two-stage cross-domain cooperation task, the task allocation model and the path planning model are determined based on the MAPPO algorithm, for example, the maximum number of training steps is set to 1 million steps (first stage: task allocation) and 2 million steps (second stage: path planning) each time. In addition, in order to verify the advantages of the proposed method, the results of single-stage MAPPO and MADDPG algorithm and MATD3 algorithm in the second stage are compared.

[0181] As Figure 7 shown, the task allocation result of the AUV selecting the USV in a certain state is shown, and the cross-domain cooperation scenario includes three USVs USV0, USV1 and USV2 and nine AUVs, wherein the corresponding information transmission frequency bands of each USV receiving data back transmission are 5kHz, 7kHz and 9kHz respectively. The connection line between the AUV and the USV represents the transmission link, and Figure 7 As can be seen from the above table, some AUVs select the frequency band according to the distance, but also consider the number of other AUVs accessing to maximize the SIR result, reduce interference and ensure that the transmission distance is as close as possible. In addition, compared with other non-linear programming methods such as heuristic method, the calculation complexity and neural network structure of multi-agent reinforcement learning in deployment are related. For example, in the network adopted in the embodiments of the present application, only two hidden layers are included, and the network structure is simple, so the reasoning time is lower than that of the heuristic method, and at the same time, some difficulties in solving traditional planning methods are overcome.

[0182] As shown in Figure 8 , the reward curve convergence results of the second stage using MAPPO, MATD3 and MADDPG algorithms respectively are shown. Among them, the MATD3 and MADDPG methods use a continuous action space, and output the speed in the coordinate axis direction. Compared with the other two methods, MAPPO has a higher convergence rate, and the average reward remains at a high level after convergence. At the same time, as shown in Figure 8 , MATD3 and MADDPG do not have obvious convergence trends. On the one hand, data needs to be collected before training to meet the policy update, and on the other hand, there is no limit to the amplitude of the update, making it difficult to converge in scenarios with large exploration space.

[0183] In addition, for the training time of different reinforcement learning algorithms used in the second stage, MAPPO requires the shortest training time, and when training without segmentation using the same policy network (i.e. one-segment MAPPO) does not reach convergence under the same training time or longer training time. This indicates the superiority of MAPPO in trajectory planning compared to other methods and the efficiency of interacting with dynamic scenarios, as well as the effectiveness of the two-stage MAPPO method.

[0184] As shown in Figure 9 , the three-dimensional trajectory visualization results of AUV and USV in a collaborative path planning process are shown. Among them, the USV cluster is located on the ocean surface, i.e. in a two-dimensional plane with a depth of 0 to achieve free movement, and the AUV can complete three-dimensional movement underwater. 901-903 are the three-dimensional trajectories of the three USVs, and 904-912 are the three-dimensional trajectories of the nine AUVs.

[0185] As can be seen from Figure 9 , after the scene is initialized, the USVs are as close as possible to the AUV cluster center, and during the process, a new return node and trajectory are planned according to the distribution of AUVs and USVs to ensure maximum channel capacity, thereby demonstrating the effectiveness and advancement of the two-stage MAPPO adopted.

[0186] As shown in Figure 10 , the convergence of communication capacity reward with training time when using different methods for collaborative path planning is shown. One-segment MAPPO, the method in the present application, i.e. two-segment MAPPO (MAPPO2), and MADDPG algorithm in the second stage are used respectively.

[0187] As shown in Figure 10 , the three-dimensional trajectory visualization results of AUV and USV in a collaborative path planning process are shown. Among them, the USV cluster is located on the ocean surface, i.e. in a two-dimensional plane with a depth of 0 to achieve free movement, and the AUV can complete three-dimensional movement underwater. 901-903 are the three-dimensional trajectories of the three USVs, and 904-912 are the three-dimensional trajectories of the nine AUVs.It can be seen that the average communication capacity of the two-stage method is higher than that of the one-stage method. This is mainly because the addition of the first-stage task allocation reduces the invalid exploration space of reinforcement learning in the one-stage method, which helps the algorithm converge and solve for the optimal result. In addition, MAPPO performs better than the MADDPG algorithm, which is mainly due to its clip strategy in policy update and its high sample utilization.

[0188] Furthermore, such as Figure 11 As shown, taking connections of 3 USVs and 9 AUVs (3-9) and 2 USVs and 6 AUVs (2-6) as examples, the convergence of communication capacity reward with training time is analyzed when the two-stage MAPPO is applied to different numbers of agents. Figure 11 It can be seen that both show significant changes in the early stages of training; however, as convergence approaches, the smaller number of agents results in a higher average communication rate reward. This is because a smaller number of agents reduces the dimensionality of the state space, thereby accelerating reinforcement learning convergence and resulting in better performance.

[0189] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0190] Based on the same inventive concept, this application also provides a marine cross-domain unmanned system cooperative path planning system for implementing the above-mentioned marine cross-domain unmanned system cooperative path planning method. The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the marine cross-domain unmanned system cooperative path planning system provided below can be found in the limitations of the marine cross-domain unmanned system cooperative path planning method described above, and will not be repeated here.

[0191] In one exemplary embodiment, such as Figure 1 As shown, a collaborative path planning system for cross-domain unmanned systems in the ocean is provided, which includes autonomous underwater vehicle swarms and unmanned surface vessel swarms.

[0192] The autonomous underwater vehicle cluster is used to obtain a task allocation result of the autonomous underwater vehicle cluster, and the task allocation result includes an information transmission frequency band of each autonomous underwater vehicle, and the information transmission frequency band is used to represent a communication link between the autonomous underwater vehicle and the unmanned ship cluster.

[0193] The unmanned ship cluster is used to determine state information of each unmanned ship in the unmanned ship cluster according to the task allocation result, and the state information of the unmanned ship includes position information of the unmanned ship cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned ship, and a target autonomous underwater vehicle, and the target autonomous underwater vehicle is an autonomous underwater vehicle that establishes a communication link with the unmanned ship; and a target path of each unmanned ship is determined according to the state information of each unmanned ship and a path planning model, and the target path is used to represent a path when the channel capacity of the unmanned ship is maximized.

[0194] Each module in the above-mentioned ocean cross-domain unmanned system cooperative path planning system can be realized by software, hardware, or a combination thereof, in whole or in part. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to each module.

[0195] In an exemplary embodiment, a kind of unmanned ship is provided, and its internal structure diagram can be as shown in Figure 12 The unmanned ship includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected by a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the unmanned ship is used to provide computing and control capabilities. The memory of the unmanned ship includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the unmanned ship is used to store data. The input / output interface of the unmanned ship is used to exchange information between the processor and external devices. The communication interface of the unmanned ship is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement an ocean cross-domain unmanned system cooperative path planning method.

[0196] In an exemplary embodiment, an autonomous underwater vehicle is provided, and its internal structure diagram can be as shown in Figure 13The autonomous underwater vehicle includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the autonomous underwater vehicle is configured to provide computing and control capabilities. The memory of the autonomous underwater vehicle includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the autonomous underwater vehicle is configured to store data. The input / output interface of the autonomous underwater vehicle is configured to exchange information between the processor and external devices. The communication interface of the autonomous underwater vehicle is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a method for cooperative path planning of a marine cross-domain unmanned system.

[0197] Those skilled in the art can understand that, Figure 12 and Figure 13 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the equipment to which the scheme of the present application is applied. A specific unmanned ship or autonomous underwater vehicle can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0198] In an exemplary embodiment, an unmanned ship is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the method embodiments described above.

[0199] In an exemplary embodiment, an autonomous underwater vehicle is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the method embodiments described above.

[0200] In an embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program is executed by a processor to implement the steps of any of the method embodiments described above.

[0201] In an embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the steps of any of the method embodiments described above.

[0202] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0203] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0204] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific manner, but should not be construed as limiting the scope of the patent of the present application. It should be noted that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A cooperative path planning method for a cross-domain unmanned marine system, characterized in that, A collaborative path planning system for cross-domain unmanned systems in the ocean, the system comprising autonomous underwater vehicle swarms and unmanned surface vessel swarms, the method comprising: Obtain the task allocation result of the autonomous underwater vehicle cluster, the task allocation result including the information transmission frequency band of each autonomous underwater vehicle, the information transmission frequency band being used to characterize the communication link between the autonomous underwater vehicle and the unmanned surface vessel cluster; The status information of each unmanned surface vessel (USV) in the USV cluster is obtained. The status information of the USV includes the location information of the USV cluster, the location information of the autonomous underwater vehicle (AUV) cluster, the channel capacity of the USV, and the target AUV. The target AUV is the AUV that establishes a communication link with the USV, as determined by the task allocation result. The target path of each unmanned surface vessel is determined based on the state information and path planning model of each unmanned surface vessel. The target path is used to characterize the path that maximizes the channel capacity of the unmanned surface vessel. The training process of the path planning model includes: After the task allocation model is trained, the task allocation result corresponding to each of the autonomous underwater vehicles is determined according to the task allocation model, and the target autonomous underwater vehicle in the initial state information of each of the unmanned surface vessels is determined according to the task allocation result. The initial path is generated based on the initial path planning model and the updated initial state information of each unmanned surface vessel. The first reward is determined based on each initial path and the first reward function, wherein the first reward function is determined based on the channel capacity and the first preset weighting coefficient. The initial path planning model is trained according to the first reward and policy gradient algorithm until the convergence condition is met. The initial path planning model includes a first initial policy network and a first initial evaluation network.

2. The method according to claim 1, characterized in that, The process of obtaining the task allocation results of the autonomous underwater vehicle cluster includes: The status information of each autonomous underwater vehicle in the autonomous underwater vehicle cluster is obtained. The status information of the autonomous underwater vehicle includes the location information of the autonomous underwater vehicle cluster, the location information of the unmanned surface vessel cluster, and the information transmission frequency band of each autonomous underwater vehicle. The task allocation result is determined based on the status information of each autonomous underwater vehicle and the task allocation model.

3. The method according to claim 1, characterized in that, The path planning model includes a first target policy network. Determining the target path for each unmanned surface vessel (USV) based on its state information and the path planning model includes: The state information of each unmanned surface vessel is input into the first target strategy network to obtain the target path corresponding to each unmanned surface vessel.

4. The method according to claim 2, characterized in that, The task allocation model includes a second target policy network. Determining the task allocation result based on the state information of each autonomous underwater vehicle and the task allocation model includes: The status information of each autonomous underwater vehicle is input into the second target policy network to obtain the task allocation result corresponding to each autonomous underwater vehicle.

5. The method according to claim 1, characterized in that, The training process of the path planning model includes, but is not limited to: When the number of training sessions meets the preset redistribution frequency, the task allocation results corresponding to each autonomous underwater vehicle are updated according to the task allocation model. The initial state information of each of the unmanned surface vessels is updated according to the updated task allocation results; Based on the updated initial state information, the initial path planning model is trained again.

6. The method according to claim 4, characterized in that, The training process of the task allocation model includes: The initial task allocation result is generated based on the initial task allocation model and the initial state information of each autonomous underwater vehicle. The second reward is determined based on the initial task allocation results and the second reward function. The second reward function is determined based on the signal quality index value and the second preset weighting coefficient. The signal quality index value is used to characterize the signal quality of the underwater acoustic signal emitted by the autonomous underwater vehicle to the unmanned surface vessel. The initial task allocation model is trained according to the second reward and policy gradient algorithm until the convergence condition is met. The initial task allocation model includes a second initial policy network and a second initial evaluation network.

7. A collaborative path planning system for a cross-domain unmanned marine system, characterized in that, The marine cross-domain unmanned system collaborative path planning system includes autonomous underwater vehicle swarms and unmanned surface vessel swarms; The autonomous underwater vehicle (AUV) cluster is used to obtain the task allocation result of the AUV cluster. The task allocation result includes the information transmission frequency band of each AUV. The information transmission frequency band is used to characterize the communication link between the AUV and the unmanned surface vessel (USV) cluster. The unmanned surface vessel (USV) swarm is used to acquire the status information of each USV in the USV swarm. The status information of the USVs includes the position information of the USV swarm, the position information of the autonomous underwater vehicle (AUV) swarm, the channel capacity of the USVs, and the target AUV. The target AUV is the AUV that establishes a communication link with the USVs, as determined by the task allocation result. The system also determines the target path of each USV based on its status information and a path planning model. The target path represents the path that maximizes the channel capacity of the USVs. The training process of the path planning model includes: After the task allocation model is trained, the task allocation result corresponding to each autonomous underwater vehicle is determined according to the task allocation model, and the target autonomous underwater vehicle in the initial state information of each unmanned surface vessel is determined according to the task allocation result; the initial path is generated according to the initial path planning model and the updated initial state information of each unmanned surface vessel. A first reward is determined based on each initial path and a first reward function, wherein the first reward function is determined based on channel capacity and a first preset weight coefficient; the initial path planning model is trained based on the first reward and a policy gradient algorithm until the convergence condition is met, wherein the initial path planning model includes a first initial policy network and a first initial evaluation network.

8. An unmanned surface vessel, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. An autonomous underwater vehicle, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Unmanned aerial vehicle path planning method for emergency information collection and transmission

    CN113050672A

  • Cross-domain collaborative hunting method, device and system

    CN116166034A