Unmanned cross-domain platform collaborative task planning method and system

By constructing a multimodal sensor network, game theory, and reinforcement learning optimization framework, combined with federated learning and an integrated air-space-sea communication network, the problems of high complexity in centralized decision-making and low efficiency in heterogeneous platform collaboration in maritime search and rescue are solved. This enables efficient and dynamic task planning and resource sharing, and improves the system's adaptability and privacy protection.

CN120952391APending Publication Date: 2025-11-14WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511033878.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies in maritime search and rescue suffer from problems such as high computational complexity in centralized decision-making, low efficiency in collaboration between heterogeneous platforms, insufficient adaptability to dynamic environments, and low efficiency in experience sharing.

Method used

By constructing a multimodal sensor network for data synchronization and cleaning, employing game theory and reinforcement learning optimization frameworks for task allocation, introducing a federated learning mechanism to achieve efficient data interaction and resource sharing across domain platforms, designing an integrated air-space-sea communication network architecture, and dynamically selecting cooperative or competitive modes to optimize task planning.

Benefits of technology

It significantly reduces the computational burden of centralized decision-making, improves the collaboration efficiency and dynamic environment adaptability between heterogeneous platforms, enhances the efficiency of experience sharing, and ensures data privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952391A_ABST
    Figure CN120952391A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned cross-domain platform collaborative task planning method and system, and the method comprises the steps: collecting original data through a multi-mode sensor, and processing the original data to obtain time-space consistency data; dynamically selecting a cooperation mode or a competition mode according to the cooperation gain value; in the cooperation mode, a hierarchical task allocation system is constructed based on the game theory, the unmanned aerial vehicle is used as a leader, a Nash equilibrium solution is obtained by using a reverse induction method according to time-space consistency data to issue a task strategy, and the unmanned surface vessel and the unmanned underwater vessel are used as executors to receive allocated tasks and optimize the task execution strategy of the unmanned surface vessel and the unmanned underwater vessel; in the competition mode, task information is issued through the unmanned aerial vehicle, the task information is received through the unmanned surface vessel and the unmanned underwater vessel, autonomous quotation is performed based on real-time environment parameters to participate in task auction, and task allocation is performed by adopting a first price auction mechanism. According to the method, the problems of high complexity of centralized decision calculation and low cooperation efficiency among heterogeneous platforms can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of modern transportation technology, specifically relating to a collaborative task planning method and system for unmanned cross-domain platforms. Background Technology

[0002] With the rapid development of artificial intelligence and unmanned technologies, collaborative mission planning for unmanned systems in maritime search and rescue faces new challenges. Maritime search and rescue scenarios are becoming increasingly complex, demanding greater adaptability to dynamic environments. This is especially true in cross-domain collaboration and resource-constrained environments, where multi-platform collaboration efficiency becomes a key limiting factor. Existing technologies, under centralized decision-making models, experience exponential growth in computational complexity as mission scale expands, making it difficult to meet the demands of rapid response. Furthermore, differences in capabilities, communication latency, and energy limitations between heterogeneous platforms render traditional optimization methods inadequate for practical application requirements.

[0003] With the development of intelligent technologies for unmanned platforms in the field of maritime search and rescue, the definition of mission scenarios, the delineation of functional boundaries, and the application requirements are becoming increasingly clear. Existing algorithms mainly rely on perception algorithms and white-box algorithms, which are highly dependent on predefined mission scenarios. However, the maritime search and rescue environment is complex and ever-changing, and target requirements and environmental conditions (such as wind, waves, and visibility) are difficult to predict. Existing algorithms struggle to achieve dynamic decision-making and planning under various constraints. Furthermore, resource competition, path conflicts, and execution uncertainties are common problems in multi-platform collaboration. How to achieve resource sharing and competitive balance through dynamic strategy adjustments has become a current research focus.

[0004] Furthermore, the development of intelligent scheduling technology has provided new solutions for maritime search and rescue. Federated learning, as a distributed optimization framework, allows platforms to train models locally and share experience through parameter aggregation, ensuring data security and improving the system's intelligence level. However, existing federated learning methods still have bottlenecks in multi-agent scheduling, such as insufficient efficiency of experience sharing under limited communication conditions and the need to improve model adaptability in dynamic environments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for collaborative task planning of unmanned cross-domain platforms in order to overcome the shortcomings of the existing technology. This method can solve the problems of high computational complexity of centralized decision-making and low efficiency of collaboration between heterogeneous platforms.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for collaborative task planning on an unmanned cross-domain platform is provided, comprising: A dynamic environment model is constructed by using a multimodal sensor network to collect raw data from unmanned aerial vehicles (UAVs), unmanned surface vessels (USVs), and unmanned underwater vehicles (UUVs). The raw data is then synchronized in time, aligned in coordinates, and cleaned to obtain spatiotemporally consistent data, ensuring the accuracy of the data. The collaboration gain value is calculated based on the team's overall collaboration capability value, average task urgency, competition loss degree, average risk level, and environmental response coefficient. Based on the collaboration gain value, the collaboration mode or competition mode is dynamically selected for task planning. In the collaborative mode, a hierarchical task allocation system is constructed based on game theory. The UAV acts as the leader and uses backward induction to obtain the Nash equilibrium solution based on the spatiotemporal consistency data to issue task strategies. The USV and UUV act as executors to receive the assigned tasks and optimize their own task execution strategies. In the competition mode, mission information is published by the unmanned aerial vehicle (UAV), and the unmanned surface vessel (USV) and the unmanned underwater vehicle (UUV) receive the mission information and participate in the mission auction by autonomously bidding based on real-time environmental parameters. The mission is allocated by the UAV using a first-price auction mechanism.

[0007] The above plan also includes: A two-layer reinforcement learning optimization framework is established, in which the upper-layer decision module uses reinforcement learning combined with task scheduling optimization to perform global task allocation, and the lower-layer execution module performs path planning and obstacle avoidance optimization, and outputs the optimal path. Based on federated learning, a decentralized learning experience-sharing mechanism is formed for multiple unmanned platforms. The unmanned aerial vehicles (UAVs) lead the global model aggregation and update, while the unmanned surface vessels (USVs) and unmanned underwater vehicles (UUVs) perform local training and encrypted parameter uploading. Construct an integrated air-space-sea hierarchical communication network architecture to achieve efficient data interaction across heterogeneous platforms.

[0008] In the above scheme, the step of dynamically selecting the cooperation mode or the competition mode based on the cooperation gain value includes: If the cooperation gain value is not less than 0, select the cooperation mode; if the cooperation gain value is less than 0, select the competition mode. In the above scheme, the real-time environmental parameters include: the unmanned surface vessel (USV) and the unmanned underwater vehicle (UUV) collect their own status parameters and mission parameters in real time; the own status parameters include spatial coordinates, motion vectors, energy reserves, and communication quality, and the mission parameters include operating area, time window, and environmental constraints.

[0009] In the above scheme, the raw data of the unmanned aerial vehicle (UAV), unmanned surface vessel (USV) and unmanned underwater vehicle (UUV) includes: flight status data of the unmanned aerial vehicle (UAV), surface motion data of the unmanned surface vessel (USV), underwater motion status data of the unmanned underwater vehicle (UUV) and environmental information provided by various sensors; The time synchronization and coordinate alignment of the raw data includes: using high-precision timestamps to synchronize the status data of the unmanned aerial vehicle (UAV), the unmanned surface vessel (USV), and the unmanned underwater vehicle (UUV), and using geographic coordinate transformation to ensure that the data of different platforms can be expressed in the same coordinate system; The data cleaning includes: after acquiring the original data, performing data cleaning on prominent features, checking for missing values ​​and abnormal data; interpolating and filling missing data, using Kalman filtering to smooth the time series data, and removing abnormal data that exceeds the range. The flight status data of the unmanned aerial vehicle (UAV) includes flight altitude, speed, heading, remaining battery power, etc., which are collected by optical cameras, lidar, INS, and GPS; wherein the optical cameras are used for target recognition and geographic environment perception, the lidar is used to generate three-dimensional terrain information, and the INS and GPS are combined to provide high-precision location information. The surface motion data of the unmanned surface vessel (USV) includes real-time position, speed, heading, fuel remaining, communication status, and information about the surrounding environment, which are collected by radar, AIS, GNSS, inertial measurement unit, and ultrasonic sensors. The underwater motion status data of the unmanned underwater vehicle (UUV) includes underwater depth, speed, heading, remaining battery power, sonar detection information, and underwater environmental data, which are collected through sonar system, multibeam echo sounder, inertial navigation system, Doppler velocimeter and environmental sensors.

[0010] In the above scheme, the UAV, acting as the leader, uses backward induction to obtain the Nash equilibrium solution based on the spatiotemporal consistency data and issues a task strategy. The USV and UUV, acting as executors, receive the assigned tasks and optimize their own task execution strategies, including: The leader uses backward induction to predict the executor's optimal response, solves an optimization problem containing the executor's objective function, and obtains a Nash equilibrium solution; the UAV issues tasks, specifies the execution platform for each task, and satisfies the conditions of task uniqueness constraint, time constraint, and resource constraint. The unmanned aerial vehicle (UAV) acts as the leader, responsible for global scheduling and task allocation; the unmanned surface vessel (USV) acts as the executor, responsible for path optimization; and the unmanned underwater vehicle (UUV) acts as the executor, responsible for sonar obstacle avoidance. The USV uses the Theta* algorithm for path planning, and the UUV uses sonar technology for obstacle avoidance. The executor receiving the assigned task and optimizing its own task execution strategy includes: the unmanned surface vessel (USV) and the unmanned underwater vehicle (UUV) selecting an execution strategy according to the assigned task to maximize their own benefits under Nash equilibrium conditions; The leader's method of obtaining the Nash equilibrium solution using backward induction includes: The game theory mentioned adopts Stackelberg game theory, and the Nash equilibrium solution is a Stackelberg equilibrium solution. The unmanned aerial vehicles (UAVs) are deployed in multiple clusters, with one UAV serving as the core master control node. This UAV is responsible for global task planning and resource scheduling, while the other UAVs are responsible for data collection and command issuance. To enhance system robustness, a dynamic leader switching mechanism is added. When the UAV serving as the core master control node fails or is in poor condition, a new core master control node is dynamically elected from the UAV cluster to address the risk of single-point failure in complex cross-domain environments.

[0011] In the above scheme, the task auction includes: the UAV receiving task feedback and confirming the task type, priority, and time limit, and acting as the auctioneer to publish task information for the USV and UUV to bid; the USV and UUV acting as bidders determining their respective bid prices; after receiving the bids, the auctioneer conducts a qualification review of the bidders and sets time and resource restrictions for participating in the bidding; the USV allocates the task based on the highest bid price, and the corresponding USV or UUV obtains the right to execute the task.

[0012] In the above scheme, under the collaborative mode, the core objective is to maximize the global task benefits. The scheme adopts a centralized collaborative allocation method dominated by the UAV, uses a consensus algorithm for resource allocation, employs high-frequency information sharing for communication, and focuses on path collaboration and data fusion for strategy optimization, and makes unified strategy arrangements. In the competitive mode, with the core objective of maximizing individual task benefits, the unmanned surface vessel (USV) and the unmanned underwater vehicle (UUV) participate in the auction by autonomously bidding. Resource allocation adopts the first-price auction mechanism. Communication only occurs during the bidding phase. The strategy optimization focuses on individual efficiency optimization, uses algorithms for autonomous obstacle avoidance, and leverages a federated learning mechanism to achieve experience sharing.

[0013] In the above scheme, the upper-layer decision module of the two-layer reinforcement learning optimization framework includes an environment state module, an Actor network, a Critic network, and an experience replay pool; the lower-layer execution module of the upper-layer decision module of the two-layer reinforcement learning optimization framework includes a path optimization module. The environment state module acquires location coordinates, remaining energy, and task level information as input to the Actor network and the Critic network. The Actor network consists of an Online policy network and a Target policy network, used to generate action policies based on the environment state. The Critic network consists of an Online Q network and a Target Q network, used to evaluate the action policies of the Actor network and provide feedback rewards. The experience replay pool stores historical interaction data and provides samples for network training through a sampling mechanism. The path optimization module combines global path planning and obstacle avoidance optimization to ultimately output the optimal path.

[0014] In the above scheme, the multi-unmanned platform experience sharing mechanism based on federated learning to form decentralized learning includes: a global layer, a communication layer, and an edge layer; The global layer, led by the UAV, is responsible for aggregating and updating the global model, and setting the global optimization objective, loss function, and weights. The communication layer ensures the security and privacy of data transmission by employing secure transmission protocols and differential privacy technology. The edge layer, including the USV and UUV, is responsible for local training, involving data preprocessing, local model updates, and parameter compression and encryption. The edge layer uploads the encrypted parameters to the global layer for global aggregation and optimization of the global model. The construction of the integrated air-space-sea layered communication network architecture includes: A layered communication network architecture is constructed, in which underwater communication adopts underwater acoustic communication technology and blue-green laser communication, and suppresses underwater multipath effects and signal attenuation through adaptive modulation and coding (AMC) technology; the surface layer deploys a 5G private network and a LoRa wide area network to support wide-area coverage and dynamic bandwidth allocation between the unmanned surface vessel (USV) and buoy nodes; the air layer communication integrates low-orbit satellite and 5G millimeter wave, with the unmanned aerial vehicle (UAV) as the air base station, and directional transmission is achieved through beamforming technology.

[0015] Furthermore, to achieve the above objectives, this invention also proposes an unmanned cross-domain platform collaborative task planning system, comprising: The global situational awareness module constructs a dynamic environment model through a multimodal sensor network, collects raw data from UAVs, unmanned surface vessels, and unmanned underwater vehicles, and performs time synchronization, coordinate alignment, and data cleaning on the raw data to obtain spatiotemporally consistent data. The dynamic selection module calculates a collaboration gain value based on the team's overall collaboration capability, average task urgency, competition loss, average risk level, and environmental response coefficient. Based on the collaboration gain value, it dynamically selects either a collaboration mode or a competition mode for task planning. The collaboration module is used to construct a hierarchical task allocation system based on game theory. The UAV acts as the leader and uses backward induction to obtain the Nash equilibrium solution based on the spatiotemporal consistency data to issue task strategies. The unmanned surface vessel and the unmanned underwater vessel act as executors to receive the assigned tasks and optimize their own task execution strategies. The Nash equilibrium solution is obtained by the leader using backward induction. The competition module is used in the competition mode to publish mission information through the UAV, receive the mission information through the unmanned surface vessel and the unmanned underwater vessel and participate in the mission auction by autonomously bidding based on real-time environmental parameters, and allocate the mission through the UAV using a first price auction mechanism.

[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) This invention provides a collaborative task planning method for unmanned cross-domain platforms, which can solve the problems of high computational complexity of centralized decision-making, low collaboration efficiency between heterogeneous platforms, and insufficient adaptability to dynamic environments. This method constructs a master-slave hierarchical control model based on game theory. The UAV is responsible for global task scheduling as the master node, and the USV and UUV are autonomously optimizing local strategies as execution nodes, which significantly reduces the computational burden of the central node. It introduces a competitive auction mechanism driven by task timeliness and quickly allocates tasks through the first-price auction to improve decision-making efficiency. This method constructs a task allocation model based on master-slave game theory. The UAV matches tasks according to the platform's capabilities to form a "capability-task" optimization strategy. It introduces a two-layer path planning framework. The upper layer uses reinforcement learning to optimize the global path, and the lower layer uses a deep deterministic strategy gradient algorithm to achieve local obstacle avoidance and dynamically adjust the navigation trajectory to avoid conflicts. Through the cooperation-competition game mechanism, the cooperation or competition mode is dynamically switched according to the urgency of the task and the resource status to balance resource allocation.

[0017] (2) The collaborative task planning method for unmanned cross-domain platforms provided by this invention can also solve the problems of insufficient adaptability to dynamic environments, low efficiency of experience sharing, and insufficient protection of data privacy. This method collects environmental data in real time through a multimodal sensor network, constructs a dynamic environment model, and provides spatiotemporal consistency support for decision-making. The experience sharing module utilizes federated learning to update the global reinforcement learning strategy through parameter aggregation, thereby improving the platform's adaptability and task generalization performance in dynamic environments. This method designs a decentralized federated reinforcement learning framework, in which UAVs act as aggregation nodes to integrate the local task experience of USVs / UUVs through parameter fusion, without the need to directly share the original data. It introduces a cross-domain semantic interoperability index and a heterogeneous cognitive consistency threshold evaluation model to quantify the efficiency of experience sharing, ensure model compatibility and optimization effect, and combine differential privacy technology to protect the data privacy of each platform through gradient perturbation, thereby reducing the risk of sensitive information leakage. Attached Figure Description

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference figures denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating a collaborative task planning method for an unmanned cross-domain platform according to Embodiment 1 of the present invention.

[0019] Figure 2 This is a schematic diagram of the unmanned cross-domain collaborative system in Embodiment 1 of the present invention.

[0020] Figure 3 This is a flowchart of the data processing of the unmanned cross-domain collaborative system in Embodiment 1 of the present invention.

[0021] Figure 4 This is a schematic diagram of the collaboration mode in Embodiment 1 of the present invention.

[0022] Figure 5 This is a schematic diagram of the two-layer reinforcement learning optimization framework in Embodiment 2 of the present invention.

[0023] Figure 6 This is a schematic diagram of the collaborative task planning framework for heterogeneous unmanned platforms based on federated learning in Embodiment 2 of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0025] It should be understood that the sequence number of each step in the embodiment does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0026] Example 1 This application provides a method for collaborative task planning on an unmanned cross-domain platform. Please refer to [link / reference]. Figure 1 , Figure 1 This is a flowchart illustrating a collaborative task planning method for an unmanned cross-domain platform according to an embodiment of this application. The collaborative task planning method for an unmanned cross-domain platform according to an embodiment of this application includes: S1 constructs a dynamic environment model through a multimodal sensor network, collects raw data from UAVs, unmanned surface vessels, and unmanned underwater vehicles, and performs time synchronization, coordinate alignment, and data cleaning on the raw data to obtain spatiotemporally consistent data.

[0027] In this embodiment, as Figure 2 As shown, a cross-domain unmanned collaborative system is formed through the coordinated operation of unmanned aerial vehicles (UAVs), unmanned surface vessels (USVs), and unmanned underwater vehicles (UUVs) to jointly complete maritime search and rescue missions. Specifically, UAVs operate in the airspace, responsible for mission issuance and feedback; USVs operate on the surface, interacting with airspace and underwater equipment and performing surface tasks; and UUVs operate underwater, performing underwater tasks and communicating with surface and airspace equipment. Figure 2 The information flow indicated by the middle arrow reflects the task allocation, data sharing, and resource coordination mechanism among UAVs, USVs, and UUVs, achieving efficient cross-platform collaboration through a layered communication network.

[0028] Specifically, in this embodiment, such as Figure 3As shown, UAVs collect data such as flight altitude, speed, heading, and remaining battery power using optical cameras, lidar, INS, and GPS. Optical cameras are used for target identification and geographic environment perception, lidar is used to generate 3D terrain information, and INS and GPS combined provide high-precision location information. USVs collect real-time position, speed, heading, remaining fuel, communication status, and surrounding environment information through radar, AIS, GNSS, inertial measurement units, and ultrasonic sensors. UUVs rely on sonar systems, multibeam echo sounders, inertial navigation systems, Doppler velocimeters, and environmental sensors to collect underwater depth, speed, heading, remaining battery power, sonar detection information, and environmental data.

[0029] The raw data acquired above includes flight status data of UAVs, surface motion data of USVs, underwater motion data of UUVs, and environmental information provided by various sensors. To ensure data accuracy, the raw data must first be synchronized in time and aligned in coordinates. This involves using high-precision timestamps to synchronize the data from UAVs, USVs, and UUVs, and using geographic coordinate transformation to ensure that data from different platforms can be expressed in the same coordinate system.

[0030] After obtaining the raw data, data cleaning is required to identify key features and check for missing values ​​and outliers. Missing data is filled using interpolation, and the time-series data is smoothed using Kalman filtering, while outliers that are outside the acceptable range are removed.

[0031] S2 calculates the collaboration gain value based on the team's overall collaboration capability value, average task urgency, competition loss degree, average risk level, and environmental response coefficient. Based on the collaboration gain value, it dynamically selects either the collaboration mode or the competition mode for task planning.

[0032] In this embodiment, dynamically selecting the cooperation mode or competition mode based on the cooperation gain value includes: If the cooperation gain value is not less than 0, select the cooperation mode; if the cooperation gain value is less than 0, select the competition mode. For maritime search and rescue missions, the cooperation gain value Based on the team's overall collaboration ability score Average urgency of tasks Competition loss Average risk level Environmental response coefficient The calculation yields the following results:

[0033] like If the benefits of collaboration outweigh the losses from competition, then the cooperative mode is adopted; if If the losses from competition outweigh the benefits from cooperation, then the competitive mode is adopted.

[0034]

[0035] The formula for calculating task urgency is:

[0036] in, , The closer a value is to 0, the higher the urgency of the task. It is to perform tasks Time required; Among them, task weight coefficient The calculation formula is:

[0037] in, P i Task i Priority; T limit,i For the task i The maximum allowed completion time.

[0038] Average urgency of tasks It is the urgency of all tasks. The mean.

[0039] Teamwork ability The calculation formula is:

[0040] in, EFF k For unmanned platforms A k The execution efficiency index Q k For communication quality index, E k For unmanned platforms A k Energy consumption for performing the task; G coop This measures the team's overall collaboration ability; a higher value indicates stronger collaboration efficiency.

[0041] Competition loss The calculation formula is:

[0042] in, C k For unmanned platforms Ak Task execution cost; B k Bidding prices for unmanned platforms; ν This refers to the communication loss factor. L comp This reflects the intensity of task competition between unmanned platforms; a higher value indicates more severe competition losses.

[0043] Based on past experience and expert evaluation, the average risk level is determined according to relevant indicators. and environmental response coefficient The details are as follows: Average risk level Threatened density Obstacle density Platform health status and communication reliability To unify the dimensions of the various indicators and normalize them, the following measures are taken:

[0044] in, and These represent the upper and lower limits of the corresponding indicators. , , and The corresponding indicator is after normalization.

[0045]

[0046] in, , , and The corresponding weights.

[0047] Environmental response coefficient With sea state index ,visibility Weather level and underwater noise It's related to the indicators.

[0048] Similarly, in order to unify the dimensions of the indicators, they are normalized.

[0049]

[0050] in, , , and The corresponding indicator is after normalization.

[0051]

[0052] in, , , and The corresponding weights.

[0053] S3, in the collaborative mode, constructs a hierarchical task allocation system based on game theory. The drone acts as the leader and uses backward induction to obtain the Nash equilibrium solution based on spatiotemporal consistency data to issue task strategies. The unmanned surface vessel and unmanned underwater vehicle act as executors to receive the assigned tasks and optimize their own task execution strategies.

[0054] Specifically, in this embodiment, the game theory adopted is Stackelberg game theory, and the Nash equilibrium solution is the Stackelberg equilibrium solution.

[0055] In this embodiment, as Figure 4 As shown, in collaborative mode, the UAV acts as the leader, responsible for global scheduling and task allocation, while the USV and UUV act as executors, responsible for path optimization and sonar obstacle avoidance, respectively. The system achieves efficient task distribution and optimization through a task allocation module. The USV uses the Theta* algorithm for path planning, while the UUV uses sonar technology for obstacle avoidance; the two collaborate through an information-sharing mechanism. The entire system aims to maximize task utility, achieving an organic integration of global scheduling and local execution through a layered architecture.

[0056] In task optimization based on master-slave game theory, a hierarchical task planning system is established with the UAV as the leader and the USV and UUV as executors. After the UAV assigns tasks, the USV and UUV need to optimize their own task execution strategies to maximize their own gains. The specific process of this hierarchical decision-making model is as follows: S31, Leader Action: UAV releases task allocation strategy, clarifies the execution platform for each task, and meets constraints such as task uniqueness, time, and resources.

[0057] S32, Executor Response: Each USV / UUV selects an execution strategy based on its assigned task, maximizing its own benefit under Nash equilibrium conditions.

[0058] S33: The leader predicts the executor's optimal response through backward induction, solves the optimization problem containing the executor's objective function, and obtains the Stackelberg equilibrium solution.

[0059] S34, Iterative Loop: Each platform shares the execution results and updates the global model through federated learning, and the UAV adjusts the allocation for the next stage accordingly.

[0060] Collection of unmanned platforms: ; Task Collection: .

[0061] UAVs act as leaders, deployed in multiple clusters, but one UAV is the core master node, responsible for global task planning and resource scheduling, while the other UAVs are responsible for data collection and command issuance. To enhance system robustness, a dynamic leader switching mechanism is added, allowing a new leader to be dynamically elected from the UAV cluster when a UAV fails or is in poor condition, thus addressing the risk of single points of failure in complex cross-domain environments. USVs and UUVs act as executors, deployed in multiple executor clusters, collaboratively executing search and rescue tasks issued by the leader.

[0062] As the leader, the UAV optimizes search and rescue mission allocation to maximize overall benefits, provided that the executors adopt the best strategies.

[0063] in Assign variables to the task; i It is the executor's ID; if a task is assigned to an executor, then... If the task is not assigned to an executor, then , U i It is the first i The rewards for each executor's task execution; x i It is the first i The tasks that each executor needs to perform It refers to the strategy adopted to perform the task.

[0064] The leader's task allocation constraint formula is: Task uniqueness constraint:

[0065] Time constraints:

[0066] Resource constraints:

[0067] in, It is to perform tasks Time required (seconds) The maximum time limit for executing the task. The unmanned platform needs to perform tasks Current energy reserves at that time To meet the minimum energy requirements for performing the mission, It is a dynamic time window.

[0068] The fitness formula for dynamically changing leaders is:

[0069] in, S j For the platform j The leader fitness, with a value range of [0,1]; E j The remaining energy of the platform; E max Maximize the platform's potential; q c For communication quality index; C j For platform computing power; C max This represents the maximum computing power in the system. η k For the task k Completion level; N k For the platform j Total number of tasks executed; w E , w q , w C , w η These are the weighting coefficients.

[0070] Specifically, in this embodiment, the leader UAV prioritizes assigning surface search and rescue missions to USVs and underwater debris detection missions to UUVs, forming a "capability-mission" matching game strategy. After receiving mission information, the USV uses path planning to determine its route, performs search and rescue missions on the water's surface, and interacts with the UAV and UUV based on surrounding environmental information to dynamically plan new trajectories. After obtaining the mission coordinates from the UAV, the UUV uses sonar to construct underwater terrain and shares relevant information with other unmanned platforms.

[0071] After tasks are assigned to UAVs, each USV and UUV needs to optimize its own task execution strategy to maximize its own benefits. The benefit function formula for task execution by the executor is as follows:

[0072] in, U For the executor's benefits in performing the task, R i For task completion rate, C i For task execution costs, E i The energy consumption is denoted as α, β, and γ, which are the corresponding weighting coefficients. Task risk,i The inherent risk value of the task. ρi The risk level is determined by the task.

[0073] Its optimization objective is:

[0074]

[0075] Among them, st U These are constraints that ensure that the executor will not unilaterally change its strategy under the optimal strategy, so that each executor's strategy is the optimal choice given the strategies of other executors. x i It is the first i The tasks that each executor needs to perform y ( x i ) is the strategy adopted to perform the task.

[0076] Task completion rate R i It can be represented as:

[0077] in η i Describes the uncertainty of the task execution environment, with a value range of [0,1]. η i =0 indicates an ideal environment. η i The larger the value, the higher the level of environmental risk.

[0078] Execution cost C i The calculation formula is:

[0079] in, L i This represents the total length of the route that the USV / UUV needs to travel. λ i Represents the difficulty of the task. ρ i The risk level is determined by the task.

[0080]

[0081] Task energy consumption E i The calculation formula is:

[0082] in, P i This represents the power consumption per unit time for USV / UUV.

[0083] In a cyclic game, when the change in payoff is less than a threshold, the game converges and ends. The convergence conditions are as follows:

[0084] in The convergence threshold, Indicates the first Phase Platform The benefits; here It is a platform The profit value in a certain iteration; Indicates the first Phase Platform The utility reflects the new benefits after an update.

[0085] In S4, in competitive mode, mission information is released via drones, and mission information is received via unmanned surface vessels and unmanned underwater vehicles. Based on real-time environmental parameters, they autonomously bid and participate in mission auctions. Missions are allocated via drones using a first-price auction mechanism.

[0086] In this embodiment, the UAV receives a maritime search and rescue mission and confirms the mission type, priority, and time limit. Subsequently, the UAV publishes the mission information, and the USV and UUV determine their respective bidding prices based on real-time environmental parameters. After bidding, the UAV conducts a qualification review and sets time and resource restrictions for participating in the bidding. Through auction allocation, the USV / UUV obtains the mission execution right. The entire process forms a closed-loop feedback mechanism, ensuring the efficiency and adaptability of mission allocation, and is suitable for multi-platform collaborative mission planning in dynamic environments.

[0087] The bidding mechanism in this embodiment is as follows: Bidders participate in the bidding process. Bidder entities (unmanned surface vessels and unmanned underwater vehicles) collect their own state parameters in real time. These parameters include spatial coordinates, motion vectors, energy reserves, and communication quality. Task parameters include the operating area, time window, and environmental constraints. A multi-objective optimization algorithm framework is constructed based on this framework, which integrates fuzzy decision theory and game theory models. Utilizing an adaptive weight allocation strategy, it generates a Pareto-optimal bid solution under constraints such as task completion, resource consumption rate, and risk coefficient.

[0088] The tendering party (drone) needs to conduct a qualification review of the bidders' status, construct a set of constraints including core indicators such as spatial accessibility, energy sustainability, and communication reliability, and set thresholds to filter out bidders who cannot perform the task well. Based on spatiotemporal accessibility analysis, a task suitability assessment system for participating entities is established to evaluate the bidders' own status. When the status parameters of a participating entity exceed the constraint threshold, the system automatically triggers a qualification filtering mechanism to terminate its participation, thereby improving the decision-making efficiency of the tendering process.

[0089] bidders A k You need to calculate your optimal bid price. B k for:

[0090] ,

[0091] in, B k bidder A k The submitted bid price, C k It is the task execution cost, determined by the path length. L k , λ i It is the complexity factor of the task; E k It is the energy consumption of task execution, affected by power consumption. P k and execution time T exec,k Influence: P k It is the power consumption per unit time of the unmanned platform; λ k , μ k As an adjustment factor, it controls the weight of energy consumption in the bidding price. EFF k This is a task execution efficiency index.

[0092] After the initial auction concludes, the leader UAV will conduct multiple rounds of bidding, with the executor determining the highest bid price. Multiple rounds of price adjustments were conducted, with the following steps: S41, Calculate the initial bid for the current round.

[0093] S42, get the current highest price.

[0094] S43, calculate the adjustment range.

[0095] S44, Update Bidding Price.

[0096] For bidders A k The formula for adjusting the bid price is:

[0097] in It is the first r The adjustment range of the wheel can be defined as:

[0098] Where α is the adjustment coefficient; EFF max This is the most efficient ratio across all platforms, used to normalize the adjustment range.

[0099] In this embodiment, the auction uses a first-price auction mechanism via drone. Based on the bids of each bidder, an optimal bid price is determined; that is, the final winning bid price is the highest bid. The final winning bid price is:

[0100] In this embodiment, under the collaborative mode, the core objective is to maximize the overall task benefits, and a centralized collaborative allocation method dominated by UAVs is adopted. Resource allocation uses a consensus algorithm, communication is conducted through high-frequency information sharing, and strategy optimization focuses on path coordination and data fusion, with unified strategy arrangements.

[0101] In competitive mode, with maximizing individual task rewards as the core objective, USVs / UUVs participate in auctions through autonomous bidding. Resource allocation employs a first-price auction mechanism, and communication occurs only during the bidding phase. Strategy optimization focuses on individual performance, employing algorithms for autonomous obstacle avoidance and leveraging federated learning for experience sharing. Ultimately, both modes proceed to the task execution phase, completing the task planning process.

[0102] In summary, the embodiments of this application provide a method for collaborative task planning of unmanned cross-domain platforms, which can solve the problems of high computational complexity of centralized decision-making and low efficiency of collaboration between heterogeneous platforms.

[0103] Example 2 This application provides a method for collaborative task planning on an unmanned cross-domain platform, wherein the method of this embodiment is basically the same as that of embodiment one, except that this method further includes: A two-layer reinforcement learning optimization framework is established. The upper-layer decision-making module of the two-layer reinforcement learning optimization framework uses reinforcement learning combined with task scheduling optimization method to perform global task allocation, while the lower-layer execution module performs path planning and obstacle avoidance optimization and outputs the optimal path.

[0104] Specifically, in this embodiment, such as Figure 5 As shown, the two-layer reinforcement learning optimization framework includes an upper-layer decision-making module and a lower-layer execution module; the upper-layer decision-making module includes an environment state module, an experience replay pool module, an Actor network module, and a Critic network module, while the lower-layer execution module includes a path optimization module.

[0105] In this embodiment, the environment state module acquires information such as location coordinates, remaining energy, and task level, which serves as input to the Actor and Critic networks. The Actor network, composed of an Online policy network and a Target policy network, is responsible for generating action policies based on the environment state. The Critic network, composed of an Online Q network and a Target Q network, evaluates the Actor network's action policies and provides feedback rewards. The experience replay pool stores historical interaction data and provides samples for network training through a sampling mechanism, improving the stability and efficiency of learning. The path optimization module combines global path planning and obstacle avoidance optimization to ultimately output the optimal path. The entire module forms a closed-loop optimization process through continuous sampling, storage, feedback rewards, and state inputs.

[0106] In this embodiment, the upper-layer optimization module is responsible for global task scheduling, and uses reinforcement learning combined with task scheduling optimization methods to allocate global tasks; the lower-layer execution module is responsible for path planning and obstacle avoidance.

[0107] The objective function for task scheduling optimization in the upper-level optimization module is:

[0108] We have also set the following constraint mechanisms for this framework: Task uniqueness constraint:

[0109] Time constraints:

[0110] Resource constraints: .

[0111] The objective function for path optimization in the lower-level execution module is:

[0112] in, C i It is to perform tasks i The cost, Ei It is to perform tasks i energy consumption P It is the set of tasks traversed along the path.

[0113] In this embodiment, the method also includes: a multi-unmanned platform experience sharing mechanism based on federated learning to form decentralized learning, in which the unmanned aerial vehicle (UAV) leads the global model aggregation and update, and the unmanned surface vessel (USV) and unmanned underwater vehicle (UUV) perform local training and encrypted parameter uploading.

[0114] Specifically, in this embodiment, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a heterogeneous unmanned platform collaborative task planning framework based on federated learning. The global layer, led by the UAV, is responsible for aggregating and updating the global model, setting the global optimization objective, loss function, and weights. The communication layer ensures the security and privacy of data transmission, employing secure transmission protocols and differential privacy technology. The edge layer, including USVs and UUVs, is responsible for local training, involving data preprocessing, local model updates, and parameter compression and encryption. Finally, the edge layer uploads the encrypted parameters to the global layer for global aggregation, optimizing the global model. The entire framework achieves data privacy protection and efficient model updates through a layered architecture, making it suitable for multi-platform collaborative task planning in dynamic environments.

[0115] Federated learning employs a decentralized learning approach, where each unmanned platform trains its model locally and optimizes it through parameter aggregation, eliminating the need to directly share raw data and thus protecting the data privacy of each platform. Within this federated learning framework, the master server is a UAV (Unmanned Aerial Vehicle), responsible for updating the global model; execution nodes are USV (Unmanned Virtual Vehicle) / UUV (Unmanned Virtual Vehicle), responsible for optimizing local tasks and training local models.

[0116] Several optimization objective functions were set for the federated learning model.

[0117] (1) Maximize the global task benefits The optimization objective formula for task scheduling in federated learning is:

[0118] in, U total This contributes to the overall mission benefits of the entire unmanned system.

[0119] (2) The formula for minimizing the loss function in task scheduling in federated learning is:

[0120] in, These are the parameters of the global task optimization model; K The number of unmanned platforms participating in federated learning;w k It is an unmanned platform A k The weight.

[0121] in w k The formula for calculating the weight is:

[0122] in, For unmanned platforms A k The amount of task data possessed; It is an unmanned platform A k The local loss function for training.

[0123] In federated learning, the formula for minimizing the task execution loss function is:

[0124] (3) Minimize privacy budget

[0125] in, For gradient sensitivity, C represents the relaxation term, and C represents the proportion of the executing party. R comm For communication rounds, Add Gaussian noise during local training The variance in.

[0126] In this embodiment, the method further includes: constructing an integrated air-space-sea hierarchical communication network architecture to achieve efficient data interaction across heterogeneous platforms.

[0127] Specifically, in this embodiment, a layered communication network architecture is first constructed. Underwater layer communication employs underwater acoustic communication technology and blue-green laser communication, and uses adaptive modulation and coding (AMC) technology to suppress underwater multipath effects and signal attenuation. The surface layer deploys a 5G private network and a LoRa wide area network, supporting wide-area coverage and dynamic bandwidth allocation between USVs and buoy nodes. The air layer communication integrates low-Earth orbit satellites and 5G millimeter waves, with UAVs acting as airborne base stations, achieving directional transmission through beamforming technology.

[0128] Secondly, a multi-mode relay and protocol conversion mechanism is constructed. Multi-protocol gateways are deployed in buoys, USVs, and UAVs to parse data packet headers and complete protocol conversion in real time. Data caching and scheduling strategies are implemented, using the LRU caching algorithm and QoS classification mechanism to reduce data forwarding latency through edge computing nodes. An anti-interference transmission model is established, constructing a channel state prediction model based on a long short-term memory network. Input parameters include signal-to-noise ratio, Doppler shift, and noise power spectral density, dynamically adjusting transmit power and modulation scheme.

[0129] Finally, a secure and reliable interactive system is constructed, using quantum encryption for transmission. Quantum keys are generated based on the BB84 protocol to perform end-to-end encryption on sensitive data.

[0130] Another aspect of this application provides an unmanned cross-domain platform collaborative task planning system, including: The global situational awareness module constructs a dynamic environment model through a multimodal sensor network, collects raw data from UAVs, unmanned surface vessels, and unmanned underwater vehicles, and performs time synchronization, coordinate alignment, and data cleaning on the raw data to obtain spatiotemporally consistent data. The dynamic selection module calculates the collaboration gain value based on the team's overall collaboration capability, average task urgency, competition loss, average risk level, and environmental response coefficient. Based on the collaboration gain value, it dynamically selects either collaboration mode or competition mode for task planning. The collaboration module is used to construct a hierarchical task allocation system based on game theory. The drone acts as the leader and uses backward induction to obtain the Nash equilibrium solution based on spatiotemporal consistency data to issue task strategies. The unmanned surface vessels and unmanned underwater vessels act as executors to receive the assigned tasks and optimize their own task execution strategies. The leader obtains the Nash equilibrium solution through backward induction. The competition module is used in competition mode to publish mission information via drones, receive mission information via unmanned surface vessels and unmanned underwater vehicles and participate in mission auctions by autonomously bidding based on real-time environmental parameters, and allocate missions via drones using a first-price auction mechanism.

[0131] In this embodiment, a privacy-preserving collaborative learning mechanism is established. Addressing the issue of sensitive information potentially involved in knowledge sharing among multiple agents in maritime search and rescue, a decentralized federated reinforcement learning framework is designed. By constructing a cross-domain semantic mapping matrix, a knowledge transfer evaluation model based on an interoperability index is proposed to effectively quantify the experience compatibility of heterogeneous agents. Furthermore, differential privacy technology is combined to achieve gradient perturbation protection, enabling each unmanned platform to share search and rescue experience and optimization strategies without disclosing its own sensor data and mission details, thereby improving the intelligence level and collaborative efficiency of the entire search and rescue system. This embodiment also constructs an integrated air-ground-space communication architecture. To ensure reliable information interaction among cross-domain heterogeneous platforms, a hierarchical adaptive communication network model is proposed. This architecture includes: an underwater layer employing hybrid underwater acoustic-blue-green laser communication; a surface layer deploying multi-standard gateways integrating 5G and LoRa protocol stacks; and an air layer constructing a space-ground fusion network, achieving wide-area coverage through millimeter-wave beamforming and low-orbit satellite relay transmission. Meanwhile, a multimodal communication performance evaluation function was proposed, and an anti-interference transmission optimization model was established that includes multiple objective constraints such as time delay sensitivity, data integrity rate, and channel stability. This ensures the stability and continuity of communication in all directions under complex marine environments, enabling timely issuance of search and rescue commands and real-time transmission of search and rescue data, thereby improving the coordination efficiency and decision-making accuracy of maritime search and rescue.

[0132] It should be noted that, depending on the implementation needs, the various steps described in this application can be broken down into more steps, or two or more steps or parts of the steps can be combined into new steps to achieve the purpose of this invention.

[0133] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for collaborative task planning on an unmanned cross-domain platform, characterized in that, include: A dynamic environment model is constructed by using a multimodal sensor network to collect raw data from UAVs, unmanned surface vessels, and unmanned underwater vehicles. The raw data is then synchronized in time, aligned in coordinates, and cleaned to obtain spatiotemporally consistent data. The collaboration gain value is calculated based on the team's overall collaboration capability value, average task urgency, competition loss degree, average risk level, and environmental response coefficient. Based on the collaboration gain value, the collaboration mode or competition mode is dynamically selected for task planning. In the collaborative mode, a hierarchical task allocation system is constructed based on game theory. The UAV acts as the leader and uses backward induction to obtain the Nash equilibrium solution based on the spatiotemporal consistency data to issue task strategies. The unmanned surface vessel and the unmanned underwater vessel act as executors to receive the assigned tasks and optimize their own task execution strategies. In the competition mode, the drone publishes mission information, the unmanned surface vessel and the unmanned underwater vehicle receive the mission information and participate in the mission auction by autonomously bidding based on real-time environmental parameters, and the drone uses a first-price auction mechanism to allocate the mission.

2. The method for collaborative task planning of an unmanned cross-domain platform according to claim 1, characterized in that, Also includes: A two-layer reinforcement learning optimization framework is established, in which the upper-layer decision module uses reinforcement learning combined with task scheduling optimization to perform global task allocation, and the lower-layer execution module performs path planning and obstacle avoidance optimization, and outputs the optimal path. Based on federated learning, a decentralized learning multi-unmanned platform experience sharing mechanism is formed, in which the UAV leads the global model aggregation and update, and the unmanned surface vessel and unmanned underwater vessel perform local training and encrypted parameter uploading. Construct an integrated hierarchical communication network architecture that combines air, space, and sea to enable data interaction across heterogeneous platforms.

3. The method for collaborative task planning on an unmanned cross-domain platform according to claim 1, characterized in that, The dynamic selection of cooperation mode or competition mode based on cooperation gain value includes: If the cooperation gain value is not less than 0, select the cooperation mode; if the cooperation gain value is less than 0, select the competition mode. The real-time environmental parameters include: the unmanned surface vessel and the unmanned underwater vessel collect their own status parameters and mission parameters in real time; the own status parameters include spatial coordinates, motion vectors, energy reserves, and communication quality, and the mission parameters include operating area, time window, and environmental constraints.

4. The method for collaborative task planning on an unmanned cross-domain platform according to claim 1, characterized in that, The raw data from the UAV, unmanned surface vessel, and unmanned underwater vehicle includes: flight status data of the UAV, surface motion data of the unmanned surface vessel, underwater motion status data of the unmanned underwater vehicle, and environmental information provided by various sensors. The time synchronization and coordinate alignment of the raw data includes: using high-precision timestamps to synchronize the status data of the UAV, the unmanned surface vessel and the unmanned underwater vehicle, and using geographic coordinate transformation to ensure that the data of different platforms can be expressed in the same coordinate system; The data cleaning includes: after acquiring the original data, performing data cleaning on prominent features, checking for missing values ​​and abnormal data; interpolating and filling missing data, using Kalman filtering to smooth the time series data, and removing abnormal data that exceeds the range. The flight status data of the UAV includes flight altitude, speed, heading, and remaining battery power, which are collected through optical cameras, lidar, INS, and GPS; wherein the optical cameras are used for target recognition and geographic environment perception, the lidar is used to generate three-dimensional terrain information, and the INS and GPS are combined to provide high-precision location information. The surface motion data of the unmanned surface vessel includes real-time position, speed, heading, fuel remaining, communication status, and information about the surrounding environment, which are collected by radar, AIS, GNSS, inertial measurement unit, and ultrasonic sensors. The underwater motion status data of the unmanned underwater vehicle includes underwater depth, speed, heading, remaining battery power, sonar detection information, and underwater environmental data, which are collected through a sonar system, multibeam echo sounder, inertial navigation system, Doppler velocimeter, and environmental sensors.

5. The method for collaborative task planning of an unmanned cross-domain platform according to claim 1, characterized in that, The drone, acting as the leader, uses backward induction to derive a Nash equilibrium solution based on the spatiotemporal consistency data and issues a task strategy. The unmanned surface vessel and the unmanned underwater vehicle, acting as executors, receive the assigned tasks and optimize their own task execution strategies, including: The leader uses backward induction to predict the optimal response of the executor, solves an optimization problem containing the executor's objective function, and obtains a Nash equilibrium solution; the drone, as the leader, issues tasks, specifies the execution platform for each task, and satisfies the conditions of task uniqueness constraint, time constraint, and resource constraint. The unmanned surface vessel acts as the executor responsible for path optimization, and the unmanned underwater vessel acts as the executor responsible for sonar obstacle avoidance; wherein the unmanned surface vessel uses the Theta* algorithm for path planning, and the unmanned underwater vessel uses sonar technology to achieve obstacle avoidance. The leader's method of obtaining the Nash equilibrium solution using backward induction includes: The game theory mentioned adopts Stackelberg game theory, and the Nash equilibrium solution is a Stackelberg equilibrium solution. The drones, acting as leaders, are deployed in multiple clusters. The core master control node is one of the drones, which is responsible for global task planning and resource scheduling. The other drones are responsible for collecting data and issuing instructions. When the drone acting as the core master control node malfunctions or is in poor condition, a new core master control node is dynamically elected from the drone cluster.

6. The method for collaborative task planning of an unmanned cross-domain platform according to claim 1, characterized in that, The mission auction includes: the UAV receiving mission feedback and confirming the mission type, priority, and time limit; acting as the auctioneer, the UAV publishing mission information for bidding on the unmanned surface vessel and the unmanned underwater vehicle; the unmanned surface vessel and the unmanned underwater vehicle determining their respective bid prices; after receiving the bids, the auctioneer conducting a qualification review of the bidders and setting time and resource restrictions for participating in the bidding; the unmanned surface vessel allocating the mission based on the highest bid price, with the corresponding unmanned surface vessel or unmanned underwater vehicle obtaining the mission execution right.

7. The method for collaborative task planning of an unmanned cross-domain platform according to claim 1, characterized in that: In the aforementioned collaborative mode, with the core objective of maximizing global task benefits, a centralized collaborative allocation method led by the UAV is adopted. Resource allocation uses a consensus algorithm, communication is conducted through high-frequency information sharing, and strategy optimization focuses on path collaboration and data fusion, with unified strategy arrangements. In the competitive mode, with the core objective of maximizing individual task benefits, the unmanned surface vessels and unmanned underwater vessels participate in the auction by autonomously bidding. Resource allocation adopts a first-price auction mechanism. Communication only occurs during the bidding phase. Strategy optimization focuses on individual performance optimization, employs algorithms for autonomous obstacle avoidance, and leverages a federated learning mechanism to achieve experience sharing.

8. The method for collaborative task planning of an unmanned cross-domain platform according to claim 2, characterized in that, The upper-level decision-making module of the two-layer reinforcement learning optimization framework includes an environment state module, an Actor network, a Critic network, and an experience replay pool; the lower-level execution module of the upper-layer decision-making module of the two-layer reinforcement learning optimization framework includes a path optimization module. The environment state module acquires location coordinates, remaining energy, and task level information as input to the Actor network and the Critic network. The Actor network consists of an Online policy network and a Target policy network, used to generate action policies based on the environment state. The Critic network consists of an Online Q network and a Target Q network, used to evaluate the action policies of the Actor network and provide feedback rewards. The experience replay pool stores historical interaction data and provides samples for network training through a sampling mechanism. The path optimization module combines global path planning and obstacle avoidance optimization to ultimately output the optimal path.

9. The method for collaborative task planning of an unmanned cross-domain platform according to claim 2, characterized in that, The multi-unmanned platform experience sharing mechanism based on federated learning to form decentralized learning includes: a global layer, a communication layer, and an edge layer; The global layer, led by the UAV, is responsible for aggregating and updating the global model, and setting the global optimization objective, loss function, and weights. The communication layer ensures the security and privacy of data transmission by employing secure transmission protocols and differential privacy technology. The edge layer, including the unmanned surface vessel and the unmanned underwater vehicle, is responsible for local training, involving data preprocessing, local model updates, and parameter compression and encryption. The edge layer uploads the encrypted parameters to the global layer for global aggregation and optimization of the global model. The construction of the integrated air-space-sea layered communication network architecture includes: A layered communication network architecture is constructed, in which underwater communication adopts underwater acoustic communication technology and blue-green laser communication, and suppresses underwater multipath effect and signal attenuation through adaptive modulation and coding technology; the surface layer deploys 5G private network and LoRa wide area network to support wide area coverage and dynamic bandwidth allocation between the unmanned surface vessel and buoy nodes; the air layer communication integrates low-orbit satellite and 5G millimeter wave, with the UAV as the air base station, and directional transmission is achieved through beamforming technology.

10. A collaborative task planning system for an unmanned cross-domain platform, characterized in that, include: The global situational awareness module constructs a dynamic environment model through a multimodal sensor network, collects raw data from UAVs, unmanned surface vessels, and unmanned underwater vehicles, and performs time synchronization, coordinate alignment, and data cleaning on the raw data to obtain spatiotemporally consistent data. The dynamic selection module calculates a collaboration gain value based on the team's overall collaboration capability, average task urgency, competition loss, average risk level, and environmental response coefficient. Based on the collaboration gain value, it dynamically selects either a collaboration mode or a competition mode for task planning. The collaboration module is used to construct a hierarchical task allocation system based on game theory. The UAV acts as the leader and uses backward induction to obtain the Nash equilibrium solution based on the spatiotemporal consistency data to issue task strategies. The unmanned surface vessel and the unmanned underwater vessel act as executors to receive the assigned tasks and optimize their own task execution strategies. The Nash equilibrium solution is obtained by the leader using backward induction. The competition module is used in the competition mode to publish mission information through the UAV, receive the mission information through the unmanned surface vessel and the unmanned underwater vessel and participate in the mission auction by autonomously bidding based on real-time environmental parameters, and allocate the mission through the UAV using a first price auction mechanism.