Cooperative path planning method, system and equipment for ocean cross-domain unmanned system

By dismantling the autonomous collaboration problem of the ocean cross-domain unmanned system into task allocation and path planning problems, a two-stage method is adopted to solve the problems of high complexity and low collaboration efficiency in the traditional method, and more efficient collaborative path planning is achieved.

CN120506959AActive Publication Date: 2025-08-19TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202511009492.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-08-19
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

The traditional ocean cross-domain unmanned system decision-making method based on rules model has problems such as high complexity, local optimization and low collaboration efficiency in multi-agent collaboration problems.

Method used

The autonomous cooperation problem of the unmanned platform is broken down into task allocation problems and path planning problems. Through the two-stage cross-domain collaboration method, the task allocation results of the autonomous submarine cluster and the status information of the unmanned boat cluster are first obtained, and the target path of the unmanned boat is determined according to the path planning model to maximize channel capacity.

Benefits of technology

The performance of ocean cross-domain unmanned system collaborative path planning is improved, with higher stability and lower solution difficulty, and more efficient collaborative path planning is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120506959A_ABST
    Figure CN120506959A_ABST
Patent Text Reader

Abstract

The invention relates to a collaborative path planning method, system and equipment for an ocean cross-domain unmanned system. The method is used in the path planning technology field. The method comprises: obtaining a task allocation result of an autonomous underwater vehicle cluster, the task allocation result comprising an information transmission frequency band of each autonomous underwater vehicle, the information transmission frequency band being used for representing a communication link between the autonomous underwater vehicle and the unmanned ship cluster; acquiring state information of each unmanned ship in the unmanned ship cluster, the state information of the unmanned ship including position information of the unmanned ship cluster, position information of the autonomous underwater vehicle cluster, channel capacity of the unmanned ship and a target autonomous underwater vehicle; the target autonomous underwater vehicle is an autonomous underwater vehicle which is determined according to a task allocation result and establishes a communication link with the unmanned ship; and determining a target path of each unmanned ship according to the state information of each unmanned ship and a path planning model, wherein the target path is used for representing a path when the channel capacity of the unmanned ship is maximized. By adopting the method, the performance of the collaborative path planning method can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of path planning technology, and in particular to a method, system and device for collaborative path planning of an ocean cross-domain unmanned system. Background Art

[0002] With the development of artificial intelligence, unmanned and clustered ocean observation is a future trend in the industrial deployment of smart oceans. The decision-making ability of multiple agents is a key indicator of the autonomy of cross-domain unmanned systems. For example, a cross-domain ocean unmanned system consists of multiple unmanned surface vehicles (USVs) and autonomous underwater vehicles (AUVs). The AUVs perform tasks such as ocean observation and data collection, while the USVs provide stable and reliable communication services for the AUVs and assist them in completing their detection missions.

[0003] Traditional rule-based decision-making methods have problems such as high solution complexity, local optimality and low collaboration efficiency when dealing with multi-agent cross-domain collaboration problems. The performance of decision-making methods needs to be improved. Summary of the Invention

[0004] Based on this, it is necessary to provide a collaborative path planning method, system and equipment for ocean cross-domain unmanned systems that can improve the performance of collaborative path planning methods in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for collaborative path planning of an ocean cross-domain unmanned system, which is used in an ocean cross-domain unmanned system collaborative path planning system. The ocean cross-domain unmanned system collaborative path planning system includes an autonomous underwater vehicle swarm and an unmanned boat swarm. The method includes:

[0006] Obtaining the task allocation results of the autonomous underwater vehicle cluster, which include the information transmission frequency bands of each autonomous underwater vehicle. The information transmission frequency bands are used to represent the communication links between the autonomous underwater vehicles and the unmanned vehicle cluster;

[0007] Obtaining status information of each unmanned vehicle in the unmanned vehicle swarm, including the location information of the unmanned vehicle swarm, the location information of the autonomous underwater vehicle swarm, the channel capacity of the unmanned vehicles, and the target autonomous underwater vehicle. The target autonomous underwater vehicle is the autonomous underwater vehicle that establishes a communication link with the unmanned vehicle determined based on the task allocation result;

[0008] The target path of each unmanned boat is determined according to the state information of each unmanned boat and the path planning model. The target path is used to represent the path that maximizes the channel capacity of the unmanned boat.

[0009] In a second aspect, the present application also provides a marine cross-domain unmanned system collaborative path planning system, which includes an autonomous underwater vehicle cluster and an unmanned boat cluster;

[0010] Autonomous underwater vehicle cluster, used to obtain the task allocation results of the autonomous underwater vehicle cluster, the task allocation results include the information transmission frequency band of each autonomous underwater vehicle, and the information transmission frequency band is used to represent the communication link between the autonomous underwater vehicle and the unmanned vehicle cluster;

[0011] The unmanned boat cluster is used to obtain the status information of each unmanned boat in the unmanned boat cluster. The status information of the unmanned boat includes the location information of the unmanned boat cluster, the location information of the autonomous underwater vehicle cluster, the channel capacity of the unmanned boat and the target autonomous underwater vehicle. The target autonomous underwater vehicle is the autonomous underwater vehicle that establishes a communication link with the unmanned boat based on the task allocation result; the target path of each unmanned boat is determined based on the status information of each unmanned boat and the path planning model. The target path is used to represent the path when the channel capacity of the unmanned boat is maximized.

[0012] In a third aspect, the present application also provides an unmanned boat, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the methods of the first aspect above when executing the computer program.

[0013] In a fourth aspect, the present application also provides an autonomous underwater vehicle, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the methods of the first aspect above when executing the computer program.

[0014] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods of the first aspect above.

[0015] In a sixth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of any one of the methods of the first aspect above.

[0016] The above-mentioned ocean cross-domain unmanned system collaborative path planning method, system and equipment first obtain the task allocation result of the autonomous submersible cluster. The task allocation result includes the information transmission frequency band of each autonomous submersible, and the information transmission frequency band is used to characterize the communication link between the autonomous submersible and the unmanned boat cluster; then, the status information of each unmanned boat in the unmanned boat cluster is obtained. The status information of the unmanned boat includes the position information of the unmanned boat cluster, the position information of the autonomous submersible cluster, the channel capacity of the unmanned boat and the target autonomous submersible. The target autonomous submersible is the autonomous submersible that establishes a communication link with the unmanned boat determined according to the task allocation result; then, the target path of each unmanned boat is determined according to the status information of each unmanned boat and the path planning model. The target path is used to characterize the path when the channel capacity of the unmanned boat is maximized. In the above method, the autonomous collaboration problem of the unmanned platform is decomposed into a task allocation problem and a path planning problem. After confirming the task allocation result, the target path of the unmanned boat is determined based on the task allocation result. Through a two-stage cross-domain collaboration method, in each stage, the type of intelligent agent is the same and the tasks performed are the same. In this way, the solution process has higher stability, lower solution difficulty and computational overhead, which can effectively improve the performance of the collaborative path planning method. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a diagram of an application environment of a method for collaborative path planning of an ocean cross-domain unmanned system in one embodiment;

[0019] Figure 2 1 is a flow chart of a method for collaborative path planning of an ocean cross-domain unmanned system in one embodiment;

[0020] Figure 3 A flowchart of the step of determining a task allocation result in one embodiment;

[0021] Figure 4 A schematic diagram of a flow chart of the training steps of a task allocation model in one embodiment;

[0022] Figure 5 Schematic diagram of a flow chart of the training steps of a path planning model in one embodiment;

[0023] Figure 6 A flowchart of a method for collaborative path planning for an ocean cross-domain unmanned system in another embodiment;

[0024] Figure 7 A schematic diagram of a task allocation result in one embodiment;

[0025] Figure 8 A comparison chart of reward curve convergence results during the second phase of training based on different reinforcement learning algorithms in one embodiment;

[0026] Figure 9 Schematic diagram of the three-dimensional trajectory of an AUV and an SUV in one embodiment;

[0027] Figure 10 A comparison diagram of the convergence results of communication capacity reward over training time when using different methods for collaborative path planning in one embodiment;

[0028] Figure 11 A graph comparing the convergence results of communication capacity rewards for different numbers of agents over training time in one embodiment;

[0029] Figure 12 This is a diagram of the internal structure of an unmanned boat in one embodiment;

[0030] Figure 13 1 is a diagram of the internal structure of an autonomous underwater vehicle in one embodiment. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0032] The ocean cross-domain unmanned system collaborative path planning method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown in the figure, the ocean cross-domain unmanned system considers building a cluster of autonomous underwater vehicles and unmanned boats, that is, including N USVs and M AUVs. AUVs are used to perform tasks such as ocean observation and data collection. USVs can provide stable and reliable communication services for AUVs and assist in completing the detection tasks of AUVs. Figure 1As shown, the movement of AUV and USV is considered to be ideal. USV can achieve free movement in a two-dimensional plane, while AUV can complete three-dimensional movement underwater, and the maximum speed of USV is higher than the maximum speed of AUV. In this application, ideal communication transmission can be achieved between USVs, that is, the communication rate meets the information transmission requirements, and information is transmitted between AUV and USV by underwater acoustic communication technology, and USV mainly collaborates by collecting data from AUV. The AUV in the above-mentioned ocean cross-domain unmanned system can be accessed by any USV, where each USV may be assigned to access k AUVs, and all AUVs need to be assigned. Therefore, under the premise of meeting the stability of communication, it is necessary to consider the constraints such as transmission interference and communication rate between AUVs to construct the optimal trajectory planning scheme for USV under multiple constraints.

[0033] Based on the above-mentioned ocean cross-domain unmanned system, the process of constructing a spread spectrum multi-user communication model and an ocean underwater acoustic channel model is explained below.

[0034] Optionally, in the multi-user communication model, in order to maximize the information transmission efficiency between USVs and AUVs while ensuring the stability of multiple communication links, code division multiple access is used within a single subsystem to achieve multi-user access, and frequency division multiple access is used for different subsystems to avoid mutual interference. Spread spectrum underwater acoustic communication achieves robust multi-user data transmission from multiple AUVs to a single USV by assigning unique and mutually orthogonal pseudo-random sequences to each AUV. Among them, the baseband form of the transmission signal corresponding to the mth AUV is:

[0035]

[0036] in, is the i-th data symbol sent by the m-th AUV, is the symbol length and satisfies , is the spreading factor, is the code element width, is the corresponding spreading code sequence.

[0037] The nth USV receiving signal can be expressed as:

[0038]

[0039] in, is the number of AUVs connected to the nth USV, and , is the received signal amplitude, is the channel impulse response function from the mth AUV to the nth USV, is the received ambient noise.

[0040] Data demodulation can be achieved by matching the received signal with the spread spectrum sequence of the corresponding AUV. Taking the first AUV as an example, the matched filtering output of the i-th symbol is expressed as:

[0041]

[0042] in, Corresponding to the i-th symbol of the received signal, is the cross-correlation function between the spreading sequence corresponding to the 1st AUV and the spreading sequence corresponding to the mth AUV. Due to the orthogonality between different spreading sequences, when hour, The result is very small, ensuring the accurate decoding of the data sent by each AUV.

[0043] Alternatively, in the ocean acoustic channel model, since electromagnetic waves and light have severe propagation attenuation underwater, underwater acoustic communication is the most effective technology for long-distance information transmission between USV and AUV. The propagation loss of sound waves underwater includes expansion loss and absorption loss, which can be expressed as:

[0044]

[0045] in, is the distance between the nth USV and the mth AUV, is the depth of the mth AUV, is the absorption coefficient, the unit is dB / km, which is closely related to the sound wave frequency f and can be expressed as:

[0046]

[0047] The ocean ambient noise is mainly caused by turbulence , waves ,vessel and thermal noise Therefore, the marine environmental noise level It can be calculated as:

[0048]

[0049]

[0050]

[0051]

[0052]

[0053] in, is the ship density correlation coefficient.

[0054] According to the above content, the signal-to-noise ratio of the underwater acoustic signal transmitted by the mth AUV received by the nth USV can be expressed as:

[0055]

[0056] in, is the sound source level, which represents the sound pressure level of the sound source at the reference distance relative to the reference sound pressure. is the noise intensity.

[0057] Direct sequence spread spectrum technology can realize multi-user spread spectrum underwater acoustic communication with weak mutual interference. By designing spreading sequences with good orthogonality, the multiple access interference between users can be approximated as noise. The signal-to-interference ratio (SIR) can be expressed as:

[0058]

[0059] Furthermore, the signal-to-interference-plus-noise ratio (SINR) of the n-th USV received signal can be expressed as:

[0060]

[0061] Therefore, the communication capacity between the nth USV and the mth AUV is expressed as:

[0062]

[0063] in, is the communication bandwidth corresponding to the nth USV, is the spreading factor.

[0064] The total communication capacity of the cross-domain collaborative system is:

[0065]

[0066] Optionally, with the goal of efficient transmission of ocean cross-domain unmanned system data while minimizing other interference, the system optimization target can be expressed as MaxC.

[0067] Optionally, the information transmitted between AUVs or USVs during information sharing is all data files, which do not contain information such as pictures and videos. Therefore, it is considered that the communication between AUVs and between USVs when transmitting information internally is ideal, and the transmitted data meets the channel capacity and minimum communication rate requirements.

[0068] In addition, considering that a USV is connected to at least one AUV, each AUV is guaranteed to be connected, so the following constraints exist:

[0069]

[0070]

[0071] in, Indicates whether the USV and AUV are connected, 1 for connected and 0 for unconnected. represents the i-th USV, represents the j-th AUV.

[0072] In an exemplary embodiment, Figure 2 As shown, a method for collaborative path planning of an ocean cross-domain unmanned system is provided, which is used in an ocean cross-domain unmanned system collaborative path planning system. The ocean cross-domain unmanned system collaborative path planning system includes an autonomous underwater vehicle cluster and an unmanned boat cluster. The method includes the following steps:

[0073] Step 201: Obtain the task allocation result of the autonomous underwater vehicle cluster.

[0074] Among them, the task allocation results include the information transmission frequency bands of each autonomous submersible, which are used to characterize the communication link between the autonomous submersible and the unmanned boat cluster.

[0075] Optionally, the autonomous underwater vehicle cluster includes multiple autonomous underwater vehicles (AUVs), each AUV may correspond to an information transmission frequency band, and then selects a USV corresponding to the information transmission frequency band to transmit data back.

[0076] The task assignment result can be the transmission frequency band where the signal quality of each autonomous underwater vehicle meets the maximum signal-to-interference ratio during signal transmission. Therefore, the process of obtaining the task assignment results for the autonomous underwater vehicle cluster can be viewed as solving an optimization problem, whose objective function is to maximize the signal-to-interference ratio.

[0077] In one possible implementation, the task allocation results of the autonomous underwater vehicle cluster can be determined based on a reinforcement learning algorithm.

[0078] Alternatively, the optimization problem can be formulated as a Markov decision process (MDP), expressed as .in Represents the state space, represents the action space, is the state transition function, is the reward function, Represents the discount factor.

[0079] Among them, for the autonomous underwater vehicle cluster, the corresponding state space is S A , the state corresponding to each autonomous underwater vehicle can be defined as S A i =[ P a , P s , T a i , T a -i ] , respectively represent the position of the AUV cluster, the position of the USV cluster, the information transmission frequency band of the current AUV (target USV), and the corresponding transmission frequency bands of other AUVs; the corresponding action space is P s(id) =[ f 1 , f 2 ,..., f n ] , that is, select the USV corresponding to the frequency band for data backtransmission.

[0080] Exemplarily, the reinforcement learning algorithm may include a multi-agent reinforcement learning algorithm (Multi-Agent Proximal Policy Optimization, MAPPO), or a multi-agent deep deterministic policy gradient (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) or its improved algorithm, such as the multi-agent twin delayed deep deterministic policy gradient (Multi-Agent Twin Delayed Deep Deterministic policy gradient algorithm, MATD3), etc., and the embodiments of the present application are not limited to this.

[0081] In another possible implementation, the task allocation result of the autonomous underwater vehicle cluster can be determined based on an optimization algorithm such as a particle swarm algorithm, an ant colony algorithm, or a genetic algorithm.

[0082] For example, taking the determination of the task allocation result of a cluster of autonomous underwater vehicles based on a genetic algorithm as an example, an initial population can be randomly generated. The population is a set of multiple individuals, each individual represents a potential solution, and the fitness value of each individual is calculated, that is, the pros and cons of the individuals are evaluated according to the optimization goal, and then individuals are selected according to the fitness value to generate the next generation. The process of fitness evaluation, selection crossover and mutation is repeated until the termination condition is met. The termination condition can be reaching the maximum number of iterations, the fitness value convergence, or finding a sufficiently good solution. The solution obtained can be used as the task allocation result of the cluster of autonomous underwater vehicles.

[0083] Step 202: Obtain status information of each unmanned vehicle in the unmanned vehicle cluster. The status information of the unmanned vehicle includes the location information of the unmanned vehicle cluster, the location information of the autonomous underwater vehicle cluster, the channel capacity of the unmanned vehicle, and the target autonomous underwater vehicle. The target autonomous underwater vehicle is the autonomous underwater vehicle that establishes a communication link with the unmanned vehicle determined according to the task allocation result.

[0084] Optionally, after confirming the task allocation result, the unmanned boat selected by each autonomous underwater vehicle can be determined. For each unmanned boat, the autonomous underwater vehicle that establishes a communication link with the unmanned boat is also determined.

[0085] Optionally, in an unmanned boat cluster, each unmanned boat can exchange position information with each other, or through sensors or other intelligent agents, an unmanned boat can obtain the position information of all unmanned boats in its cluster, as well as the position information of all autonomous underwater vehicles in an autonomous underwater vehicle cluster.

[0086] Optionally, the channel capacity in the status information of the unmanned boat indicates that the channel capacity of the unmanned boat can be calculated based on the current task allocation result. Calculated.

[0087] Step 203: Determine the target path of each unmanned boat based on the state information of each unmanned boat and the path planning model. The target path is used to represent the path that maximizes the channel capacity of the unmanned boat.

[0088] Optionally, the problem of determining the target path of each unmanned boat can be regarded as an optimization problem, whose optimization goal is to maximize the total channel capacity.

[0089] Optionally, the path planning model can be determined based on a reinforcement learning algorithm, or based on an optimization algorithm such as a particle swarm algorithm, an ant colony algorithm, or a genetic algorithm, and this embodiment of the present application does not limit this.

[0090] Optionally, when determining the path planning model based on the reinforcement learning algorithm, for the unmanned boat cluster, the corresponding state space is , the state corresponding to each unmanned boat can be defined as S P i =[ P a , P s , T i , C i ] , respectively represent the position of the AUV cluster, the position of the USV cluster, the target autonomous underwater vehicle of the current USV, and the channel capacity of the current USV; the corresponding action space is A S i =[ v x , v y ] , respectively represent the direction vectors along the x and y coordinate axes in the two-dimensional plane. Assuming that the USV control system can complete the motion solution of the current orientation and ensure approximate omnidirectional motion, it is possible to simplify the USV dynamics and output a specific orientation to make the USV move at a constant speed in eight orientations. Control the trajectory of the USV, that is, if the position coordinates (x, y) of the USV are both greater than 0, and The value of can be If the position coordinates (x, y) of the USV are both less than 0, and The value of can be − If one of the USV position coordinates (x, y) is 0 and the other is greater than 0, and The values of can be If one of the USV position coordinates (x, y) is 0 and the other is less than 0, and The values can be - and 0.

[0091] Exemplarily, the reinforcement learning algorithm may include a multi-agent reinforcement learning algorithm (Multi-Agent Proximal Policy Optimization, MAPPO), or a multi-agent deep deterministic policy gradient (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) or its improved algorithm, such as the multi-agent twin delayed deep deterministic policy gradient (Multi-Agent Twin Delayed Deep Deterministic policy gradient algorithm, MATD3), etc., and the embodiments of the present application are not limited to this.

[0092] In another possible implementation, the task allocation result of the autonomous underwater vehicle cluster can be determined based on an optimization algorithm such as a particle swarm algorithm, an ant colony algorithm, or a genetic algorithm.

[0093] It's important to note that existing methods for collaborative planning of swarms of unmanned watercraft and autonomous underwater vehicles (AUVs) based on multi-agent reinforcement learning algorithms involve different tasks for different agents to perform when collaborating. For example, a USV on the surface might provide the optimal path for the AUV to collect data, while an AUV underwater might select the appropriate frequency band to transmit data while completing data collection within the target area. In this scenario, the agents cannot use the same policy network to make decisions across multiple tasks. Furthermore, as the number of agents increases, the number of new states introduced by interactions with the environment also increases, making it easy for the original algorithm to fall into a local optimum.

[0094] In the embodiment of the present application, the task allocation result of the autonomous underwater vehicle cluster is first obtained. The task allocation result includes the information transmission frequency band of each autonomous underwater vehicle, which is used to represent the communication link between the autonomous underwater vehicle and the unmanned vehicle cluster. Then, the status information of each unmanned vehicle in the unmanned vehicle cluster is obtained. The status information of the unmanned vehicle includes the location information of the unmanned vehicle cluster, the location information of the autonomous underwater vehicle cluster, the channel capacity of the unmanned vehicle, and the target autonomous underwater vehicle. The target autonomous underwater vehicle is the autonomous underwater vehicle that establishes a communication link with the unmanned vehicle determined according to the task allocation result. Then, the target path of each unmanned vehicle is determined based on the status information of each unmanned vehicle and the path planning model. The target path is used to represent the path that maximizes the channel capacity of the unmanned vehicle. In the above method, the autonomous collaboration problem of the unmanned platform is decomposed into a task allocation problem and a path planning problem. After confirming the task allocation result, the target path of the unmanned vehicle is determined based on the task allocation result. Through the two-stage cross-domain collaboration method, in each stage, the type of intelligent agent is the same and the tasks performed are the same. In this way, the solution process has higher stability, lower solution difficulty and computational overhead, and can effectively improve the performance of the collaborative path planning method.

[0095] In an exemplary embodiment, Figure 3 As shown, the task allocation results of the autonomous underwater vehicle cluster are obtained, including:

[0096] Step 301: Obtain status information of each master underwater vehicle in the autonomous underwater vehicle cluster.

[0097] Among them, the status information of the autonomous underwater vehicle includes the location information of the autonomous underwater vehicle cluster, the location information of the unmanned boat cluster and the information transmission frequency band of each autonomous underwater vehicle.

[0098] Optionally, in an autonomous underwater vehicle cluster, each master underwater vehicle cluster can exchange position information with each other, or through sensors or other intelligent agents, each master underwater vehicle can obtain the position information of all autonomous underwater vehicles in the cluster, as well as the position information of each unmanned boat in the unmanned boat cluster.

[0099] Optionally, the information transmission frequency band of each autonomous underwater vehicle may be the USV to which the AUV selects information transmission in the current state. It is understandable that as the position of the AUV or USV changes, the signal quality corresponding to the information transmission frequency band may not necessarily meet the preset signal threshold.

[0100] Step 302: Determine a task allocation result based on the status information of each master submersible and the task allocation model.

[0101] Optionally, the status information of each master submersible may be input into a task allocation model, and the task allocation result may be determined according to the output of the task allocation model.

[0102] Optionally, the task allocation model may be determined based on a reinforcement learning algorithm. The embodiment of the present application does not limit the type of reinforcement learning.

[0103] Optionally, taking the determination of the task allocation model based on MAPPO as an example, in actual application, the task allocation model may include a second target strategy network, and determining the task allocation result according to the status information of each main underwater vehicle and the task allocation model may include the following steps:

[0104] The status information of each main submersible is input into the second target strategy network to obtain the task allocation results corresponding to each main submersible.

[0105] Optionally, the second target policy network may be composed of a multilayer perceptron (MLP), including an input layer, multiple hidden layers, and an output layer.

[0106] For example, taking the second target strategy network including three fully connected layers as an example, the second target strategy network may generally include an input layer, two hidden layers, and an output layer, where the input layer is used to receive data and does not perform calculations.

[0107] Optionally, the state space including the state information of each master underwater vehicle can be As input data, it is received by the input layer and outputs the action probability distribution. The action with the highest probability can be determined as the task allocation result corresponding to each main submersible.

[0108] By obtaining the status information of each main submersible in the autonomous submersible cluster and determining the task allocation result according to the status information of each main submersible and the task allocation model, each intelligent agent in the autonomous submersible cluster can accurately determine the corresponding task allocation result based on the status information of the autonomous submersible cluster based on the pre-trained task allocation model.

[0109] After determining the mission results, the status information of each unmanned vehicle can be further determined, and the target path of each unmanned vehicle can be determined. In an exemplary embodiment, the path planning model includes a first target strategy network and a first target evaluation network. The target path of each unmanned vehicle is determined based on the status information of each unmanned vehicle and the path planning model, including:

[0110] The status information of each unmanned boat is input into the first target strategy network to obtain the target path corresponding to each unmanned boat.

[0111] Optionally, the path planning model can be determined based on a reinforcement learning algorithm. The embodiments of the present application do not limit the type of reinforcement learning.

[0112] Optionally, the first target policy network and the second target policy network are two different policy networks, and their network parameters are different.

[0113] Optionally, the first target policy network may be composed of a multilayer perceptron (MLP), including an input layer, multiple hidden layers, and an output layer. For example, the first target policy network may have two hidden layers.

[0114] Optionally, when determining the target path corresponding to each unmanned boat, that is, when determining the movement direction and movement speed of each unmanned boat at the next moment, the state space including the state information of each unmanned boat can be As input data, it is received by the input layer and outputs the action probability distribution. Its action space is discrete, and the action with the highest probability can be determined as the target path corresponding to each unmanned boat.

[0115] The following describes the training process for the task assignment model and the path planning model. Since the path-scale model's input includes the target AUV determined based on the task assignment results, the task assignment model must be trained first. After the task assignment model's network parameters are updated, the initial path planning model is trained to obtain the path-scale model.

[0116] In an exemplary embodiment, Figure 4 As shown in Figure 2, the training process of the task allocation model includes:

[0117] Step 401: Generate an initial task allocation result according to the initial task allocation model and the initial state information of each master submersible.

[0118] Optionally, the initial status information of each main submersible also includes the position information of the autonomous submersible cluster, the position information of the unmanned boat cluster and the information transmission frequency band of each main submersible.

[0119] Optionally, the initial state information may be input into a second initial strategy network of the initial task allocation model to obtain an initial task allocation result.

[0120] Step 402: Determine a second reward based on the initial task assignment results and the second reward function.

[0121] The second reward function is determined according to the signal quality index value and the second preset weight coefficient, and the signal quality index value is used to characterize the signal quality of the underwater acoustic signal transmitted by the autonomous underwater vehicle to the unmanned boat.

[0122] Optionally, in the first stage, the AUV and USV face the same task and have the same collaborative goal. Therefore, it can be considered that when determining the task allocation result in the first stage, the second reward function corresponding to each AUV is the same and shared, which can be expressed by the following formula:

[0123]

[0124] Wherein, h represents the second preset weight coefficient. Taking the signal quality index value as the signal-to-interference ratio as an example, the task allocation result is used to represent the communication link established with the USV for the current agent k when the signal-to-interference ratio is as large as possible.

[0125] Step 403: Train the initial task allocation model according to the second reward and policy gradient algorithm until the convergence condition is met.

[0126] The initial task allocation model includes a second initial strategy network and a second initial evaluation network.

[0127] Optionally, training the initial task allocation model includes training the second initial strategy network and updating its network parameters. , get the second target strategy network; train the second initial evaluation network and update its network parameters , and obtain the second target evaluation network.

[0128] Exemplarily, the second initial policy network and the second initial evaluation network can be updated alternately, first updating the second initial policy network based on the advantage function estimated by the second initial evaluation network with fixed parameters, and then updating the second initial evaluation network using data collected by the policy network after updating the network parameters.

[0129] Optionally, set the strategy of the agent AUVi Re-expressed as ,in represents the network parameters of AUVi, Indicates the status information of AUVi, Represents the action of AUVi; then the objective function of the current intelligent agent AUV is is represented as:

[0130]

[0131] Where t represents the time or number of steps, and the optimal strategy estimate is obtained by maximizing the sum of expected rewards at future moments.

[0132] Alternatively, importance sampling can be used to effectively utilize the expectation of the current strategy, that is, by introducing KL divergence, expressed as:

[0133]

[0134] in represents the old policy distribution of agent i, Then represents the new strategy distribution, and the KL divergence can measure the difference between the new and old strategy distributions. Furthermore, the above distribution can be corrected to the following form:

[0135]

[0136] in Represents the estimation strategy.

[0137] In order to further limit the update range of the strategy while ensuring the exploration performance, a truncation method is often used to limit the strategy update range. The loss function is expressed as:

[0138]

[0139] Among them, clip represents the truncation function, which returns the upper and lower limits of importance sampling. Represents the advantage function, which is used to improve learning efficiency and enhance training stability.

[0140] Each intelligent AUV is regarded as a common task and reward, so the update method of the second initial policy network can be set to consistent update. Furthermore, the second initial policy network is expanded to multiple intelligent agents, and its loss function can be expressed as:

[0141]

[0142] Where B represents the batch size, N represents the number of intelligent AUVs, and S represents the policy entropy. is a hyperparameter.

[0143] Similarly, the loss function of the second initial evaluation network can be expressed as:

[0144] L c φ = 1 BN ∑ i=1 B · ∑ k=1 N (max[ V φ i,k - R ̂ i 2 , clip V φ i,k , V φ ' i,k -ε, V φ ' i,k +ε - R ̂ i 2 ]

[0145] in Represents a discount reward, Represent the new and old state value functions respectively.

[0146] Optionally, the convergence condition can be satisfied when the number of iterations meets a preset iteration threshold, or when the average cumulative reward fluctuates less than a preset threshold within a certain number of iterations, or when the update amplitude of the policy network parameters is less than a threshold, the convergence condition can be considered to be satisfied and training can be stopped.

[0147] In an exemplary embodiment, Figure 5 As shown in Figure 2, the training process of the path planning model includes:

[0148] Step 501: After the task allocation model training is completed, the task allocation results corresponding to each autonomous underwater vehicle are determined according to the task allocation model, and the target autonomous underwater vehicle in the initial state information of each unmanned vehicle is determined according to the task allocation results.

[0149] Step 502: Generate an initial path based on the initial path planning model and the updated initial state information of each unmanned boat.

[0150] Step 503: Determine a first reward based on each initial path and a first reward function, where the first reward function is determined based on the channel capacity and a first preset weight coefficient.

[0151] Optionally, in the second stage, the AUV and USV have the same mission and the same collaborative goal. Therefore, it can be considered that when determining the target path planning in the second stage, the first reward function corresponding to each USV is the same and shared, which can be expressed by the following formula:

[0152]

[0153] in, is the channel capacity, For boundary rewards, the USV needs to collect data backhaul within the specified area and change its own trajectory based on the results of the assigned tasks to maximize the channel capacity.

[0154] Step 504: Train the initial path planning model according to the first reward and policy gradient algorithm until the convergence condition is met. The initial path planning model includes a first initial policy network and a first initial evaluation network.

[0155] Optionally, the training process of the path planning model is similar to the training process of the task allocation model. When determining the initial state information, it is necessary to determine the target autonomous underwater vehicle corresponding to each unmanned boat based on the trained task allocation model, and then determine the initial path.

[0156] It should be noted that the training of the path planning model depends on the task allocation results output by the task allocation model. During the training process, the unmanned boat moves according to the initial target path determined by the initial path planning model, and its position information will continue to change. With the iteration of the training times, the task allocation result is not the optimal strategy under the current environment. Therefore, based on the current status information of the AUV and USV, the task allocation model can be input again to determine the new task allocation result, and then the initial path planning model can be trained based on the new task allocation result.

[0157] In one possible implementation, whether the task allocation result needs to be updated may be determined based on the training status of the initial path planning model.

[0158] For example, when the first reward continues to decrease within a certain number of iterations, the task allocation results corresponding to the respective master submersibles may be updated according to the task allocation model.

[0159] In another possible implementation method, when the number of training times meets the preset reallocation frequency, the task allocation results corresponding to each main submersible can be updated according to the task allocation model; the initial state information of each unmanned boat can be updated according to the updated task allocation results; and the initial path planning model can continue to be trained according to the updated initial state information.

[0160] As an optional implementation, Figure 6 As shown, the ocean cross-domain unmanned system collaborative path planning method provided in the embodiment of the present application may include the following specific steps:

[0161] Step 601: Generate an initial task allocation result according to the initial task allocation model and the initial state information of each master submersible.

[0162] Step 602: Determine a second reward based on the initial task assignment results and the second reward function.

[0163] The second reward function is determined according to the signal quality index value and the second preset weight coefficient, and the signal quality index value is used to characterize the signal quality of the underwater acoustic signal transmitted by the autonomous underwater vehicle to the unmanned boat.

[0164] Step 603: Train the initial task allocation model according to the second reward and policy gradient algorithm until the convergence condition is met, and obtain the task allocation model. The initial task allocation model includes the second initial policy network and the second initial evaluation network.

[0165] Step 604: Determine the task allocation results corresponding to each master underwater vehicle according to the task allocation model, and determine the initial state information of each unmanned vehicle according to the task allocation results.

[0166] Step 605: Generate an initial path based on the initial path planning model and the updated initial state information of each unmanned boat.

[0167] Step 606: Determine a first reward based on each initial path and the first reward function.

[0168] The first reward function is determined according to the channel capacity and the first preset weight coefficient.

[0169] Step 607: Train the initial path planning model according to the first reward and the policy gradient algorithm.

[0170] Step 608: When the number of training times meets the preset reallocation frequency, the task allocation results corresponding to each master submersible are updated according to the task allocation model.

[0171] Step 609: Update the initial state information of each unmanned boat according to the updated task allocation result.

[0172] Step 610: Continue training the initial path planning model according to the updated initial state information until a convergence condition is met, thereby obtaining a path planning model. The initial path planning model includes a first initial policy network and a first initial evaluation network.

[0173] Step 611: Obtain status information of each master underwater vehicle in the autonomous underwater vehicle cluster.

[0174] Among them, the status information of the autonomous underwater vehicle includes the location information of the autonomous underwater vehicle cluster, the location information of the unmanned boat cluster and the information transmission frequency band of each autonomous underwater vehicle.

[0175] Step 612: Input the status information of each master submersible into the second target strategy network to obtain the task allocation result corresponding to each master submersible.

[0176] The task allocation results include the information transmission frequency bands of each autonomous underwater vehicle, which are used to characterize the communication link between the autonomous underwater vehicle and the unmanned vehicle cluster;

[0177] Step 613: Obtain status information of each unmanned boat in the unmanned boat cluster.

[0178] Among them, the status information of the unmanned boat includes the location information of the unmanned boat cluster, the location information of the autonomous underwater vehicle cluster, the channel capacity of the unmanned boat and the target autonomous underwater vehicle. The target autonomous underwater vehicle is the autonomous underwater vehicle that establishes a communication link with the unmanned boat based on the task allocation results.

[0179] Step 614: Input the status information of each unmanned boat into the first target strategy network to obtain the target path corresponding to each unmanned boat.

[0180] For example, to verify the effectiveness of the methods described in the embodiments of this application, a cross-domain collaborative scenario involving more than three USVs and more than nine AUVs can be set up, where the USV movement range is [-500m, 500m] and the AUV operating depth is [0, 500m]. Experiments were conducted based on the Torch 2.0.1 + cu117 deep learning framework, using an Intel Core(TM) i7-13700KF CPU and a GeForce RTX 4070 GPU. The AUVs were assumed to move randomly in three-dimensional space at a constant speed, equipped with underwater cameras for information collection, and capable of local autonomous obstacle avoidance using multi-beam sonar. Furthermore, due to the wide operating range of the AUVs, their movement in three-dimensional space was considered an approximate point model. In the design of a two-stage cross-domain collaborative task, the task allocation model and path planning model were determined based on the MAPPO algorithm. The maximum number of training steps for each task was set to 1 million (for the first stage, task allocation) and 2 million (for the second stage, path planning), respectively. In addition, in order to verify the advantages of the proposed method, the results are compared with the single-stage MAPPO and the results of the MADDPG algorithm and MATD3 algorithm in the second stage.

[0181] like Figure 7 The above shows the task allocation result of AUV selecting USV in a certain state. The cross-domain collaboration scenario includes three USVs, USV0, USV1 and USV2, and 9 AUVs. The corresponding information transmission frequency bands of the data received and transmitted back by each USV are 5kHz, 7kHz and 9kHz respectively. The connection line between AUV and USV represents the transmission link. Figure 7 As can be seen in the figure, some AUVs complete the selection of frequency bands according to distance, but at the same time take into account the number of other connected AUVs to maximize the SIR result, reduce interference and ensure that the transmission distance is as close as possible. In addition, compared with other nonlinear programming methods such as heuristics, the computational complexity of multi-agent reinforcement learning during deployment is related to the neural network structure. For example, the network used in the embodiment of this application contains only two hidden layers. Its network structure is simple, so its inference time is lower than the solution time of the heuristic method, while overcoming some difficulties that are difficult to solve with traditional planning methods.

[0182] like Figure 8 As shown in the figure, the reward curve convergence results of the second stage using MAPPO, MATD3 and MADDPG algorithms are shown. Among them, MATD3 and MADDPG methods use continuous action space and output the speed in the coordinate axis direction. Compared with the other two methods, MAPPO has a higher convergence rate and the average reward remains at a high level after convergence. At the same time, Figure 8 As shown in the figure, MATD3 and MADDPG have no obvious convergence trend. On the one hand, data needs to be collected in the early stage of training to satisfy the strategy update. On the other hand, there is no limit on the update amplitude, which makes it difficult to converge in scenarios with a large exploration space.

[0183] Furthermore, MAPPO achieved the shortest training time for the second phase using different reinforcement learning algorithms. However, when trained without segmentation using the same policy network (i.e., one-stage MAPPO), convergence was not achieved at the same or even longer training times. This demonstrates MAPPO's superiority over other methods in trajectory planning and its efficiency in interacting with dynamic scenarios, while also demonstrating the effectiveness of the two-stage MAPPO approach.

[0184] like Figure 9 As shown, the three-dimensional trajectory visualization results of AUV and USV during a collaborative path planning process, where the USV cluster is located on the ocean surface, that is, it can move freely in a two-dimensional plane with a depth of 0, while the AUV can complete three-dimensional movement underwater. 901-903 are the three-dimensional trajectories of three USVs, and 904-912 are the three-dimensional trajectories of nine AUVs.

[0185] From Figure 9 It can be seen from the figure that after the scene is initialized, the USV is as close to the AUV cluster center as possible. During the process, a new return node and trajectory are planned through a task redistribution based on the distribution of AUVs and USVs, thereby maximizing the channel capacity, reflecting the effectiveness and advancement of the adopted two-stage MAPPO.

[0186] like Figure 10 As shown, the convergence of communication capacity reward with training time when different methods are used for collaborative path planning, namely, one-stage MAPPO, the method in the embodiment of the present application, namely two-stage MAPPO (MAPPO2), and the MADDPG algorithm in the second stage.

[0187] like Figure 10It can be seen that the average communication capacity of the two-stage method is higher than that of the one-stage method. This is mainly because the addition of the first-stage task allocation reduces the ineffective exploration space of reinforcement learning in the one-stage method, thereby helping the algorithm converge and solve to the optimal result. In addition, MAPPO performs better than the MADDPG algorithm, which is also mainly due to its clip strategy in policy update and its high sample utilization.

[0188] Further, such as Figure 11 As shown in the figure, taking the connection of 3 USVs and 9 AUVs (3-9) and the connection of 2 USVs and 6 AUVs (2-6) as examples, the convergence of the communication capacity reward with training time when the two-stage MAPPO acts on different numbers of agents is analyzed. Figure 11 As can be seen, both agents show significant changes in the early stages of training; however, as convergence approaches, agents with a smaller number of agents experience a higher average communication rate reward. This is because a smaller number of agents reduces the state space dimension, which in turn accelerates reinforcement learning convergence, resulting in better performance.

[0189] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0190] Based on the same inventive concept, the embodiment of the present application also provides a marine cross-domain unmanned system collaborative path planning system for implementing the above-mentioned marine cross-domain unmanned system collaborative path planning method. The implementation solution provided by this system is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations in one or more marine cross-domain unmanned system collaborative path planning system embodiments provided below can be found in the above-mentioned limitations on the marine cross-domain unmanned system collaborative path planning method, and will not be repeated here.

[0191] In an exemplary embodiment, Figure 1 As shown, a collaborative path planning system for an ocean cross-domain unmanned system is provided, which includes an autonomous underwater vehicle cluster and an unmanned boat cluster;

[0192] Autonomous underwater vehicle cluster, used to obtain the task allocation results of the autonomous underwater vehicle cluster, the task allocation results include the information transmission frequency band of each autonomous underwater vehicle, and the information transmission frequency band is used to represent the communication link between the autonomous underwater vehicle and the unmanned vehicle cluster;

[0193] The unmanned boat cluster is used to determine the status information of each unmanned boat in the unmanned boat cluster based on the task allocation results. The status information of the unmanned boat includes the location information of the unmanned boat cluster, the location information of the autonomous underwater vehicle cluster, the channel capacity of the unmanned boat and the target autonomous underwater vehicle. The target autonomous underwater vehicle is the autonomous underwater vehicle that establishes a communication link with the unmanned boat; the target path of each unmanned boat is determined based on the status information of each unmanned boat and the path planning model. The target path is used to represent the path when the channel capacity of the unmanned boat is maximized.

[0194] Each module in the above-mentioned ocean cross-domain unmanned system collaborative path planning system can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device's memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0195] In an exemplary embodiment, an unmanned boat is provided, the internal structure of which can be shown as follows: Figure 12 As shown. The unmanned boat includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the unmanned boat is used to provide computing and control capabilities. The memory of the unmanned boat includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the unmanned boat is used to store data. The input / output interface of the unmanned boat is used to exchange information between the processor and external devices. The communication interface of the unmanned boat is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for collaborative path planning of an ocean cross-domain unmanned system is realized.

[0196] In an exemplary embodiment, an autonomous underwater vehicle is provided, the internal structure of which can be shown as follows: Figure 13As shown. The autonomous underwater vehicle includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the autonomous underwater vehicle is used to provide computing and control capabilities. The memory of the autonomous underwater vehicle includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the autonomous underwater vehicle is used to store data. The input / output interface of the autonomous underwater vehicle is used to exchange information between the processor and an external device. The communication interface of the autonomous underwater vehicle is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for collaborative path planning of an ocean cross-domain unmanned system is implemented.

[0197] Those skilled in the art will understand that Figure 12 and Figure 13 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the equipment to which the scheme of the present application is applied. The specific unmanned boat or autonomous underwater vehicle may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0198] In an exemplary embodiment, an unmanned boat is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps described in any of the above method embodiments when executing the computer program.

[0199] In an exemplary embodiment, an autonomous underwater vehicle is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps described in any of the above method embodiments when executing the computer program.

[0200] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps described in any of the above method embodiments are implemented.

[0201] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of any of the above method embodiments when executed by a processor.

[0202] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0203] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0204] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for collaborative path planning of an ocean cross-domain unmanned system, characterized in that: A method for a collaborative path planning system for an ocean cross-domain unmanned system, wherein the collaborative path planning system includes an autonomous underwater vehicle cluster and an unmanned boat cluster, includes: Obtaining a task assignment result for the autonomous underwater vehicle cluster, the task assignment result including an information transmission frequency band for each autonomous underwater vehicle, the information transmission frequency band being used to characterize a communication link between the autonomous underwater vehicle and the unmanned vehicle cluster; Obtaining status information of each of the unmanned vehicles in the unmanned vehicle cluster, the unmanned vehicle status information including location information of the unmanned vehicle cluster, location information of the autonomous underwater vehicle cluster, channel capacity of the unmanned vehicles, and a target autonomous underwater vehicle, where the target autonomous underwater vehicle is an autonomous underwater vehicle that establishes a communication link with the unmanned vehicle, as determined by the task assignment result; A target path for each unmanned boat is determined according to the state information of each unmanned boat and a path planning model, wherein the target path is used to represent a path that maximizes the channel capacity of the unmanned boat.

2. The method according to claim 1, characterized in that The obtaining of the task allocation result of the autonomous underwater vehicle cluster includes: Acquire status information of each of the autonomous underwater vehicles in the autonomous underwater vehicle cluster, where the status information of the autonomous underwater vehicles includes location information of the autonomous underwater vehicle cluster, location information of the unmanned vehicle cluster, and information transmission frequency bands of each of the autonomous underwater vehicles; The task allocation result is determined according to the status information of each autonomous underwater vehicle and a task allocation model.

3. The method according to claim 1, characterized in that The path planning model includes a first target strategy network, and determining the target path of each unmanned boat according to the state information of each unmanned boat and the path planning model includes: The status information of each unmanned boat is input into the first target strategy network to obtain the target path corresponding to each unmanned boat.

4. The method according to claim 2, characterized in that The task allocation model includes a second target strategy network, and determining the task allocation result according to the state information of each autonomous underwater vehicle and the task allocation model includes: The status information of each autonomous underwater vehicle is input into the second target strategy network to obtain the task allocation result corresponding to each autonomous underwater vehicle.

5. The method according to claim 3, characterized in that The training process of the path planning model includes: After the task allocation model training is completed, determining the task allocation results corresponding to each of the autonomous underwater vehicles according to the task allocation model, and determining the target autonomous underwater vehicle in the initial state information of each of the unmanned vehicles according to the task allocation results; generating an initial path according to the initial path planning model and the updated initial state information of each of the unmanned boats; Determining a first reward according to each of the initial paths and a first reward function, wherein the first reward function is determined according to a channel capacity and a first preset weight coefficient; The initial path planning model is trained according to the first reward and policy gradient algorithm until a convergence condition is met, wherein the initial path planning model includes a first initial policy network and a first initial evaluation network.

6. The method according to claim 5, characterized in that The training process of the path planning model includes: When the number of training times meets the preset reallocation frequency, updating the task allocation results corresponding to each of the autonomous underwater vehicles according to the task allocation model; Updating the initial state information of each of the unmanned boats according to the updated task allocation result; The initial path planning model continues to be trained according to the updated initial state information.

7. The method according to claim 4, characterized in that The training process of the task allocation model includes: generating an initial task allocation result according to the initial task allocation model and the initial state information of each of the autonomous underwater vehicles; determining a second reward based on each of the initial task assignment results and a second reward function, wherein the second reward function is determined based on a signal quality index value and a second preset weight coefficient, wherein the signal quality index value is used to represent the signal quality of the underwater acoustic signal transmitted by the autonomous underwater vehicle to the unmanned vehicle; The initial task allocation model is trained according to the second reward and policy gradient algorithm until a convergence condition is met, wherein the initial task allocation model includes a second initial policy network and a second initial evaluation network.

8. A collaborative path planning system for ocean cross-domain unmanned systems, characterized by: The ocean cross-domain unmanned system collaborative path planning system includes an autonomous underwater vehicle cluster and an unmanned boat cluster; The autonomous underwater vehicle cluster is used to obtain a task allocation result of the autonomous underwater vehicle cluster, wherein the task allocation result includes an information transmission frequency band of each autonomous underwater vehicle, and the information transmission frequency band is used to represent a communication link between the autonomous underwater vehicle and the unmanned vehicle cluster; The unmanned boat cluster is used to obtain the status information of each unmanned boat in the unmanned boat cluster, the status information of the unmanned boat includes the position information of the unmanned boat cluster, the position information of the autonomous underwater vehicle cluster, the channel capacity of the unmanned boat and the target autonomous underwater vehicle, and the target autonomous underwater vehicle is an autonomous underwater vehicle that establishes a communication link with the unmanned boat determined according to the task allocation result; the target path of each unmanned boat is determined according to the status information of each unmanned boat and the path planning model, and the target path is used to represent the path when the channel capacity of the unmanned boat is maximized.

9. An unmanned boat, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An autonomous underwater vehicle comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle path planning method for emergency information collection and transmission

    CN113050672A

  • Cross-domain collaborative hunting method, device and system

    CN116166034A

  • Unmanned cluster spectrum management method based on multi-agent reinforcement learning

    CN118301759A

  • AUV (Autonomous Underwater Vehicle) comprehensive path planning method for communication efficiency and obstacle avoidance

    CN118915762A

  • Sea-air integrated cross-domain heterogeneous collaborative planning method

    CN119717868A

Cited By

  • Dynamic task allocation and path planning method for unmanned ship cluster

    CN121091894A

  • Inspection boat autonomous collaborative awareness and dynamic planning system and method

    CN122284679A