Task on-demand dynamic cooperation method and device for heterogeneous multi-agent

Through the master-slave formation mode of leading drones and multi-source information fusion, the problems of mission adaptability, computing bottlenecks and dynamic environment adaptability of multi-UAV collaborative systems are solved, and efficient and low-cost mission collaboration and resource optimization are achieved.

CN120780018AInactive Publication Date: 2025-10-14SHANGHAI JUSHI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510909191.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing multi-UAV collaborative systems have deficiencies in mission adaptability, computing and communication bottlenecks, dynamic environment adaptability, and heterogeneous intelligent agent collaboration efficiency, making it difficult to collaborate efficiently in complex environments.

Method used

Adopting the master-slave formation mode of the leading drone, through global task planning, dynamic allocation and multi-source information fusion, and taking advantage of the heterogeneous capabilities of the leading and following drones, dynamic task collaboration on demand is achieved, including cross-device adaptation of initial semantic segmentation and high-precision semantic segmentation.

Benefits of technology

It improves the efficiency of task execution, reduces hardware and computing costs, enhances adaptability to complex dynamic environments and resource sharing capabilities, and improves the collaborative performance of multiple UAV systems in complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780018A_ABST
    Figure CN120780018A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous multi-agent-oriented task on-demand dynamic cooperation method and device, belongs to the technical field of heterogeneous multi-unmanned aerial vehicle systems, and particularly relates to heterogeneous multi-unmanned aerial vehicle task on-demand dynamic cooperation. The problems of insufficient task adaptability, bottleneck in calculation and communication, poor dynamic environment adaptability, low heterogeneous unmanned aerial vehicle cooperation efficiency and task scheduling and decision lagging in existing multi-unmanned aerial vehicle cooperation are solved. The method comprises the following steps: receiving a high-precision semantic segmentation result returned by any flight-following unmanned aerial vehicle by a flight-leading unmanned aerial vehicle, and performing multi-source information fusion on the high-precision semantic segmentation result and an initial semantic segmentation result to obtain a final semantic segmentation result. The heterogeneous multi-agent-oriented task on-demand dynamic cooperation method and device are suitable for heterogeneous multi-agent systems such as an unmanned aerial vehicle (UAV) and the like, and efficient task cooperation and resource optimization can be realized in a dynamic flight environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of heterogeneous multi-unmanned aerial vehicle systems, and particularly relates to task on-demand dynamic cooperation of heterogeneous multi-unmanned aerial vehicles. BACKGROUND

[0002] In existing multi-agent systems, especially in the field of unmanned aerial vehicles (UAVs), individual agents have significant limitations in terms of perception range, detection angle, endurance, and computing resources, which severely restrict their ability to complete tasks independently in complex environments. Therefore, the current research trend tends to adopt multi-agent cooperation to overcome these limitations. The following are some existing technical solutions that demonstrate the application of multi-UAV cooperation in video surveillance, search and rescue, target detection, and other fields.

[0003] 1) Multi-UAV video surveillance system SkyStitch: The SkyStitch system solves the problem of limited field of view of onboard cameras through multi-UAV cooperation. The system generates a panoramic video stream by stitching multiple aerial video streams, providing users with a wider field of view. To reduce the workload of the ground station, SkyStitch adopts a distributed feature extraction technique. In addition, the system uses flight controller prompts to improve the efficiency of video stitching and reduces video jitter through a Kalman filter-based state estimation model, thereby improving the speed and quality of video stitching.

[0004] 2) Multi-UAV cooperative search platform: To address the limited capabilities of individual micro-UAVs, researchers designed a multi-UAV cooperative search experimental platform. The platform uses image detection technology to search for targets in image frames and achieves autonomous cooperative control of multiple UAVs through a master-slave formation mode. A group of three micro-UAVs was used in the experiment to improve the efficiency and coverage of target detection through cooperative search.

[0005] 3) Multi-UAV cooperative system SkyNet: The SkyNet system aims to address the poor visibility and limited onboard computing resources of individual UAVs. The system calculates the three-dimensional position of a person by cross-searching from multiple views, enabling accurate and real-time personnel identification and positioning. SkyNet improves recognition accuracy through fusion of aerial images and reduces processing delay through task scheduling between edge devices and cloud servers.

[0006] 4) Edge-assisted multi-UAV network Air-CAD: The Air-CAD system leverages air-ground collaboration to achieve rapid and accurate crowd anomaly detection. The system, which includes two phases, person detection and multi-feature analysis, improves detection accuracy and reduces inference latency. By leveraging the synergy of edge computing and cloud computing, Air-CAD overcomes the issues of individual drones' perspective inaccuracies and limited computing and network capabilities.

[0007] Although existing multi-agent collaborative systems have made some progress in areas such as video surveillance, search and rescue, and target detection, these technologies still have the following significant shortcomings, which limit their widespread application and performance optimization in complex environments: 1) Insufficient Task Adaptability: Existing systems are mostly designed for specific tasks (such as video stitching, object search, or person location) and lack support for on-demand dynamic collaboration among heterogeneous multi-agent tasks. This task-specific nature makes it difficult for the systems to flexibly adapt to complex tasks and fails to fully utilize the heterogeneous capabilities of multiple agents.

[0008] 2) Computational and communication bottlenecks: Existing systems face significant bottlenecks in computing and communication. For example, while edge computing has reduced the pressure on ground stations, the issues of computing task allocation and communication latency in multi-agent collaboration remain unresolved, especially in large-scale heterogeneous agent networks.

[0009] 3) Poor adaptability to dynamic environments: Existing technologies have poor adaptability to complex dynamic environments. For example, while the Kalman filter-based state estimation model can reduce video jitter, it is still difficult to ensure real-time performance and accuracy in rapidly changing environments (such as sudden target movement or drastic changes in ambient lighting).

[0010] 4) Low efficiency in heterogeneous agent collaboration: Existing systems often assume that agents have homogeneous capabilities and fail to fully consider the collaborative optimization of heterogeneous agents (such as drones with different sensors, computing power, and endurance). This homogeneous assumption limits the system's performance in complex tasks.

[0011] 5) Task scheduling and decision lag: Although existing technologies (such as SkyNet) have been optimized in task scheduling, the real-time performance of task scheduling and decision-making is still insufficient in dynamic multi-agent collaboration. Especially when task requirements change rapidly, the system's response speed and decision-making accuracy are difficult to meet actual needs.

[0012] In summary, existing technologies for multi-agent collaboration still suffer from issues such as insufficient task adaptability, computational and communication bottlenecks, poor adaptability to dynamic environments, low efficiency in heterogeneous agent collaboration, and delayed task scheduling and decision-making. These issues limit the widespread application and performance optimization of multi-agent systems in complex environments. A method for on-demand dynamic collaboration of heterogeneous multi-agent systems is urgently needed to address these shortcomings. Summary of the Invention

[0013] The present invention proposes a method and device for on-demand dynamic collaboration of tasks for heterogeneous multi-agents, which solves the problems of insufficient task adaptability, computing and communication bottlenecks, poor adaptability to dynamic environments, low efficiency of heterogeneous UAV collaboration, and delayed task scheduling and decision-making in existing multi-UAV collaboration.

[0014] The method for on-demand dynamic collaboration of heterogeneous multi-agent tasks according to the present invention comprises a leading UAV and multiple following UAVs, which fly in a master-slave formation flight mode with the leading UAV as the main body. The method comprises the following steps: Global mission planning step: The leading UAV formulates a global mission plan based on mission requirements and its own perception capabilities; The step of acquiring a large-scale image: the pilot drone acquires low-resolution large-scale images according to the global mission plan; Dynamic allocation step: The leading UAV uses the initial semantic segmentation model to perform initial semantic segmentation on the low-definition large-scale image, obtains the initial semantic segmentation result, identifies the area requiring high-precision processing, and dynamically allocates the coordinate information of the area requiring high-precision processing to the following UAV; Multi-source information fusion step: the leading UAV receives the high-precision semantic segmentation result returned by any following UAV, and performs multi-source information fusion on it with the initial semantic segmentation result to obtain the final semantic segmentation result; Adaptation step during testing of the Leading Flying UAV: ​​The Leading Flying UAV dynamically updates the initial semantic segmentation model by updating the statistics of the new images it collects in real time to adapt during cross-device testing.

[0015] Furthermore, a preferred embodiment is provided, in the dynamic allocation step, identifying areas requiring high-precision processing is as follows: The initial semantic segmentation model divides the initial semantic segmentation model into multiple image blocks and extracts the attention score of each; Multiple image blocks are sorted from high to low according to the attention scores, and the image blocks that are ranked in front by a given number are regarded as areas that need high-precision processing.

[0016] Furthermore, a preferred embodiment is provided, in the multi-source information fusion step, the multi-source information fusion is performed with the initial semantic segmentation result as follows: The initial semantic segmentation result is a coarse-grained segmentation result; The high-precision semantic segmentation result is a fine-grained segmentation result; The coarse-grained segmentation results are combined with the fine-grained segmentation results by direct coverage or probability comparison to generate the final semantic segmentation results.

[0017] Furthermore, a preferred embodiment is provided, wherein the adaptation step during the pilot drone test includes: Layer normalization statistics calculation steps: Iteratively calculate the runtime mean and variance of new images collected by the pilot drone in the channel and spatial dimensions; Statistics update step: Based on the runtime mean and variance of the new images collected by the Lingfei UAV in the channel and spatial dimensions, the exponential moving average method is used to update the statistics in real time.

[0018] The present invention also proposes a task-on-demand dynamic collaboration method for heterogeneous multi-agents, wherein the heterogeneous multi-agents include a leading drone and multiple following drones, and fly in a master-slave formation flight mode with the leading drone as the main body; the method includes the following steps: Steps for receiving coordinate information: any follower drone receives the coordinate information of the area requiring high-precision processing assigned by the leader drone; Steps for obtaining small-scale images: any following drone obtains high-definition small-scale images of the area that needs high-precision processing based on the coordinate information of the area that needs high-precision processing; High-precision semantic segmentation step: Any of the following drones uses a high-precision semantic segmentation model to perform high-precision semantic segmentation on the high-definition small-scale images of the area that requires high-precision processing, and obtains high-precision semantic segmentation results; Returning high-precision results: Any follower drone returns the high-precision semantic segmentation results to the leader drone. Adaptation step during follow-up drone testing: Any follow-up drone dynamically updates the high-precision semantic segmentation model by fusing the statistics of newly collected images from other follow-up drones in real time to adapt to cross-device testing.

[0019] Furthermore, a preferred embodiment is provided, wherein the high-precision semantic segmentation model is a CNN-based semantic segmentation model.

[0020] Furthermore, a preferred embodiment is provided, wherein the adaptation step during the following drone test includes: Batch normalization statistics calculation: Iteratively calculate the runtime mean and variance of each channel in the spatial dimension for each new image collected by the following drone; Each following drone sends the runtime mean and variance of each channel of the collected new image obtained by its iterative calculation in the spatial dimension to other following drones; Statistics sharing and fusion: each follow-up unmanned aerial vehicle receives the running mean and variance of the collected new images for each channel in the spatial dimension sent by other follow-up unmanned aerial vehicles, and stores them into a memory pool; based on the data stored in the memory pool, the statistics are updated using an exponential moving average method.

[0021] The application also proposes a task on-demand dynamic cooperation device for a heterogeneous multi-agent, which includes a leading unmanned aerial vehicle and a plurality of follow-up unmanned aerial vehicles, and flies in a master-slave formation flight mode with the leading unmanned aerial vehicle as the master; the device includes the following modules: Global task planning module: the leading unmanned aerial vehicle formulates a global task plan according to task requirements and its own perception ability; Acquisition of large-range images module: the leading unmanned aerial vehicle acquires low-definition large-range images collected according to the global task plan; Dynamic allocation module: the leading unmanned aerial vehicle performs initial semantic segmentation on the low-definition large-range images using an initial semantic segmentation model, obtains initial semantic segmentation results, and identifies areas that need high-precision processing, and dynamically allocates coordinate information of the areas that need high-precision processing to the follow-up unmanned aerial vehicles; Multi-source information fusion module: the leading unmanned aerial vehicle receives high-precision semantic segmentation results returned by any one of the follow-up unmanned aerial vehicles, and performs multi-source information fusion on the high-precision semantic segmentation results and the initial semantic segmentation results to obtain final semantic segmentation results; Leading unmanned aerial vehicle test adaptation module: the leading unmanned aerial vehicle dynamically updates the initial semantic segmentation model by updating the statistics of the new images collected in real time to adapt to cross-device testing.

[0022] The application also proposes a task on-demand dynamic cooperation device for a heterogeneous multi-agent, which includes a leading unmanned aerial vehicle and a plurality of follow-up unmanned aerial vehicles, and flies in a master-slave formation flight mode with the leading unmanned aerial vehicle as the master; the device includes the following modules: Receiving coordinate information module: any one of the follow-up unmanned aerial vehicles receives the coordinate information of the areas that need high-precision processing allocated by the leading unmanned aerial vehicle; Acquisition of small-range images module: any one of the follow-up unmanned aerial vehicles acquires high-definition small-range images of the areas that need high-precision processing collected according to the coordinate information of the areas that need high-precision processing; High-precision semantic segmentation module: any one of the follow-up unmanned aerial vehicles performs high-precision semantic segmentation on the high-definition small-range images of the areas that need high-precision processing using a high-precision semantic segmentation model to obtain high-precision semantic segmentation results; High-precision result transmission module: any one of the follow-up unmanned aerial vehicles returns the high-precision semantic segmentation results to the leading unmanned aerial vehicle; The follow-up UAV test time adaptation module: any follow-up UAV dynamically updates a high-precision semantic segmentation model by fusing statistical quantities of images newly collected by other follow-up UAVs in real time, to adapt to cross-device testing.

[0023] The application further provides an electronic device, comprising: one or more processors; a storage device configured to store one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the task-on-demand dynamic cooperation method for heterogeneous multi-agent as described in any of the above embodiments.

[0024] The application has the following beneficial effects: The core innovation of the application is to provide a task-on-demand dynamic cooperation method for heterogeneous multi-agent, which provides an efficient, low-cost and adaptable solution for the practical application requirements of heterogeneous multi-agent systems such as unmanned aerial vehicles (UAVs) in complex dynamic environments. 1. The task-on-demand dynamic cooperation method for heterogeneous multi-agent significantly improves task execution efficiency: through the task-on-demand dynamic cooperation mechanism, the task division of multi-agent can be allocated and adjusted in real time according to task requirements and environmental changes, avoiding resource waste and task conflicts, and significantly improving task execution efficiency; in tasks such as UAV cooperative semantic segmentation, target detection and search and rescue, complex tasks can be quickly responded to and completed, meeting the demand for efficiency in practical applications.

[0025] 2. The task-on-demand dynamic cooperation method for heterogeneous multi-agent reduces hardware and computing costs: the use of low-cost sensors and lightweight model design significantly reduces hardware costs and computing resource consumption, making the technology widely applicable in resource-constrained scenarios; by assigning high-precision semantic segmentation tasks to follow-up UAVs, the computing burden of lead UAVs is reduced, further reducing overall energy consumption and costs.

[0026] 3. The task-on-demand dynamic cooperation method for heterogeneous multi-agent achieves adaptation to complex dynamic environments: through the cross-device testing time adaptation method and dynamic environment optimization mechanism, model parameters can be adjusted in real time in complex dynamic environments such as target movement and light changes, maintaining high-precision segmentation and target detection performance; this feature makes the UAV system more adaptable and practical in real-world search and rescue, disaster monitoring and other practical scenarios.

[0027] 4. The on-demand dynamic task collaboration method for heterogeneous multi-agents described in this invention enables multi-device collaboration and resource sharing: by sharing and fusing statistics across multiple devices (i.e., drones), it can simulate the effect of increasing batch size, improve model adaptability, and avoid the additional consumption of video memory resources. This feature enables multi-drone systems to efficiently utilize resources in collaborative tasks, improving overall performance.

[0028] The method and device for on-demand dynamic task collaboration for heterogeneous multi-agents described in the present invention are applicable to heterogeneous multi-agent systems such as unmanned aerial vehicles (UAVs), and can achieve efficient task collaboration and resource optimization in a dynamic flight environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 A schematic diagram of a process for on-demand dynamic collaboration of heterogeneous multi-agent tasks in one embodiment of the present invention; Figure 2 A schematic diagram of the architecture of the SegFormer-B1 model in one embodiment of the present invention; Figure 3 A schematic diagram of a multi-source information fusion process in one embodiment of the present invention; Figure 4 FIG. 1 is a schematic diagram of an architecture adapted for cross-device testing in one embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to make the technical solutions and advantages of the present invention more clearly described, the specific embodiments of the present invention will be further described in detail and completely in conjunction with the accompanying drawings. The various embodiments described below are only part of the preferred embodiments of the present invention, rather than all implementation plans; the various embodiments described below are intended to explain the present invention and cannot be understood as limiting the present invention; the reasonable combination of the technical features defined in the various embodiments of the present invention, as well as all other implementation plans obtained by ordinary technicians in this field based on the embodiments of the present invention without making creative work, all fall within the scope of protection of the present invention.

[0032] In a first embodiment, a method for on-demand dynamic collaboration of tasks for heterogeneous multi-agents is provided. The heterogeneous multi-agents include a leading drone and multiple following drones, which fly in a master-slave formation flight mode with the leading drone as the main body. The method includes the following steps: Global mission planning step: The leading UAV formulates a global mission plan based on mission requirements and its own perception capabilities; The step of obtaining a large-scale image: the pilot drone obtains the collected low-definition large-scale image according to the global mission plan; Dynamic allocation step: The leading UAV uses the initial semantic segmentation model to perform initial semantic segmentation on the low-definition large-scale image, obtains the initial semantic segmentation result, identifies the area requiring high-precision processing, and dynamically allocates the coordinate information of the area requiring high-precision processing to the following UAV; Multi-source information fusion step: the leading UAV receives the high-precision semantic segmentation result returned by any following UAV, and performs multi-source information fusion on it with the initial semantic segmentation result to obtain the final semantic segmentation result; Adaptation step during the pilot drone test: The pilot drone dynamically updates the initial semantic segmentation model by updating the statistics of the new images it collects in real time to adapt during cross-device testing.

[0033] In this embodiment, the heterogeneous multi-agent refers to multiple heterogeneous drones, including a leading drone and multiple following drones, which can be called a heterogeneous multi-drone collaboration framework.

[0034] In this embodiment, the low-definition large-scale image and the high-definition small-scale image described below are not absolute numerical descriptions, that is, they do not limit the definition of high-definition or the range of large-scale. They are relative expressions, namely: Compared with low-definition large-scale images, high-definition small-scale images are clearer (i.e., they can clearly display more details) and have a smaller shooting range (i.e., they cover a smaller field of view).

[0035] In this embodiment, the area requiring high-precision processing is not described by an absolute value, but a relative expression, namely: In a low-definition, large-scale image, a region that requires higher definition than other regions.

[0036] In this embodiment, the high-precision semantic segmentation result is obtained by performing high-precision semantic segmentation on a high-definition small-scale image of an area requiring high-precision processing using a high-precision semantic segmentation model by the following drone.

[0037] The high-precision semantic segmentation result may also be referred to as an improved segmentation result, or simply an improved result.

[0038] Similarly, the initial semantic segmentation result can be referred to as the initial segmentation result or initial result; the final semantic segmentation result can be referred to as the final segmentation score or final result.

[0039] In this embodiment, high-precision semantic segmentation, high-precision semantic segmentation results, and high-precision semantic segmentation models are not absolute numerical descriptions, but are relative expressions, namely: Compared with the initial semantic segmentation, the high-precision semantic segmentation has higher segmentation accuracy.

[0040] Compared with the initial semantic segmentation results, the high-precision semantic segmentation results have higher segmentation accuracy.

[0041] Compared with the initial semantic segmentation model, the high-precision semantic segmentation model has higher segmentation accuracy (or more fine-grained).

[0042] In this embodiment, a camera is installed on the Lingfei UAV for collecting low-definition and large-scale images.

[0043] In this embodiment, the coordinate information of the area requiring high-precision processing sent by the leading drone is GPS position information of the area requiring high-precision processing.

[0044] In this embodiment, the leading UAV collects low-definition and large-scale images through high-altitude image capture.

[0045] The following drone captures low-altitude images and collects high-definition small-scale images based on coordinate information (i.e., GPS location information).

[0046] In this embodiment, the leading UAV is equipped with a high-performance computing module (such as NVIDIA Jetson Xavier NX), and any one of the following UAVs is equipped with a medium-performance computing module (such as NVIDIA Jetson TX2).

[0047] The (high-performance or medium-performance) computing module is used to implement a task-on-demand dynamic collaboration method for heterogeneous multi-agents.

[0048] The performance of the high-performance computing module carried on the leading UAV is stronger than that of the medium-performance computing module carried on any of the following UAVs.

[0049] The Lingfei UAV is equipped with a high-performance computing module to process low-definition and large-scale images.

[0050] Although the following drone needs to process higher-precision images, it does not require higher performance within the range of the images it processes, and can save costs by being equipped with a medium-performance computing module.

[0051] In this embodiment, the method can realize a dynamic task collaboration mechanism based on demand. Based on the different task requirements (such as coverage area), the leading drone formulates a global task plan and then collects the corresponding low-resolution large-scale imagery according to the global task plan.

[0052] In this embodiment, the leading UAV is also responsible for global navigation, controlling the leading UAV and multiple following UAVs to execute the global mission planning task.

[0053] In this embodiment, the leading UAV is also responsible for formation control, controlling the leading UAV and multiple following UAVs to execute a master-slave formation flight mode with the leading UAV as the main body.

[0054] In this embodiment, the following drone dynamically adjusts its flight path according to the instructions of the leading drone and goes to the designated area to perform tasks (such as collecting high-definition small-area images).

[0055] In this embodiment, a formation control strategy is used to enable the leading UAV and multiple following UAVs to maintain relative positions, and dynamic allocation (or dynamic task allocation) is performed at the same time to ensure efficient collaboration between the leading UAV and multiple following UAVs in a complex environment.

[0056] In this embodiment, dynamic allocation means that during multi-UAV formation collaboration, the lead UAV coordinates and allocates subtasks such as perception (e.g., image acquisition), computation, and communication to each UAV in real time, based on mission requirements, environmental changes, and UAV status, to achieve globally optimal collaborative efficiency and robustness. For example, certain UAVs can be temporarily assigned as relay nodes, responsible for data transfer and coordination, ensuring stable information transmission in complex terrain.

[0057] The leading drone dynamically adjusts the cruising path, flight speed or collaborative position of certain drones based on mission progress and changes in environmental obstacles to adapt to current mission requirements.

[0058] When a drone exits a mission due to a malfunction, insufficient battery, etc., the leading drone will reallocate its mission to other drones in real time to maintain the continuity and stability of the overall mission.

[0059] In this embodiment, the coordinate information of the area requiring high-precision processing is dynamically allocated to the following drone. By dynamically allocating computing tasks, the computing power of heterogeneous drones is fully utilized, avoiding resource waste.

[0060] In this embodiment, the purpose of adapting during cross-device testing is to improve the adaptability and robustness of the airborne (i.e., drone-mounted) semantic segmentation model in a dynamic flight environment.

[0061] In this embodiment, the "cross-device" in the cross-device test adaptation refers to "between multiple drones." The "test adaptation" in the cross-device test adaptation is a special term.

[0062] Test-time adaptation (TTA) is a technique for lightweight adjustments to models (such as semantic segmentation models) during inference to address the problem of inconsistent input data distribution (i.e., distribution shift) during training, thereby improving the robustness and generalization ability of the model in real-world environments.

[0063] Without retraining the model, TTA adaptively adjusts the input or internal state of the model during the test phase, allowing the model to maintain good performance in the new environment.

[0064] Traditional TTA is for a single device.

[0065] In this implementation, cross-device (or cross-drone) test-time adaptation (TTA) is proposed, extending the TTA mechanism to multi-drone collaboration scenarios. This aims to address the performance degradation of models (i.e., semantic segmentation models, initial semantic segmentation models, and high-precision semantic segmentation models) when the distribution of perception data from multiple drones is inconsistent, thereby improving overall robustness, accuracy, and collaboration.

[0066] In this implementation, through cross-device test-time adaptation and statistical sharing and updating among multiple drones, efficient dynamic adjustment of the model (semantic segmentation model) is achieved during the test-time adaptation phase without updating model parameters, thereby significantly improving the real-time performance and segmentation accuracy of the model in complex environments.

[0067] In this embodiment, if Figure 4 The figure shows a schematic diagram of the architecture adapted for cross-device testing. BN layer (Batch Normalization): is a method of regularization and accelerated training widely used in deep neural networks, and is a structure within the semantic segmentation model.

[0068] Memory: refers to the unit on the drone that stores data.

[0069] In this embodiment, due to the limited hardware resources of the UAV, the onboard model is adapted by updating statistics (ie, no model parameters need to be updated).

[0070] In this embodiment, in order to adapt to the dynamic flight environment, a dynamic update of the semantic segmentation model is proposed. The leading UAV and the following UAV update the statistics of the semantic segmentation model through new image data collected in real time to improve the adaptability and segmentation accuracy of the model.

[0071] In this embodiment, the balance between low cost and high efficiency is achieved by: Low-cost sensors: All drones are equipped with low-cost cameras that collect low-definition, large-scale images for onboard processing, reducing hardware costs.

[0072] Computational task separation: By assigning high-precision semantic segmentation tasks to the following drone, the computational burden of the leading drone is reduced while improving the energy efficiency of the overall system.

[0073] Communication optimization: Through local task allocation and information fusion, the communication overhead between drones is reduced and the real-time performance of the system is improved.

[0074] In addition, in one embodiment, low-cost cameras (or low-cost sensors) are installed on the leading UAV and the following UAV to capture images at a relatively low cost.

[0075] In addition, in one embodiment, since the sensors used to collect images by the drone are all low-cost sensors (i.e., low-cost cameras), the initial semantic segmentation results and the high-precision semantic segmentation model obtained are both low-resolution results; therefore: The Lingfei UAV upsamples the low-resolution initial semantic segmentation results to high-resolution results to obtain high-resolution initial semantic segmentation results; The leading drone upsamples the low-resolution, high-precision semantic segmentation model returned by the following drone to a high-resolution result to obtain a high-resolution, high-precision semantic segmentation model, and performs multi-source information fusion on it with the high-resolution initial semantic segmentation result to obtain the final semantic segmentation result.

[0076] In the second embodiment, in the dynamic allocation step, the area requiring high-precision processing is identified as follows: The initial semantic segmentation model divides the initial semantic segmentation model into multiple image blocks and extracts the attention score of each; Multiple image blocks are sorted from high to low according to the attention scores, and the image blocks that are ranked in front by a given number are regarded as areas that need high-precision processing.

[0077] In this embodiment, the leading drone sends the coordinate information of the area requiring high-precision processing to the following drone based on the information collected by the GPS.

[0078] In this implementation, attention scores are extracted for each image patch, helping the pilot drone focus on areas with greater information or detail. This allows the model to identify image patches of high importance (or "important patches") in the semantic segmentation task during inference. These attention scores indicate the model's level of attention to different regions and can be used to identify areas requiring further refinement (i.e., supplementary detail).

[0079] In this embodiment, the image block with a high attention score represents the location of the area in the low-definition large-scale image that needs to be supplemented with detailed information.

[0080] In this embodiment, the leading drone sends the coordinate information of the area that requires high-precision processing to the following drone, which performs high-precision semantic segmentation. The leading drone then performs multi-source information fusion, realizing shared model parameters without adding additional computing and storage burdens.

[0081] In addition, in one embodiment, the initial semantic segmentation model is a Transformer-based semantic segmentation model.

[0082] In addition, in one embodiment, the Transformer-based semantic segmentation model is a SegFormer-B1 model.

[0083] In this implementation, the SegFormer-B1 model mainly consists of an encoder (such as a Transformer encoder) and a decoder.

[0084] The decoder gradually restores the multi-scale features extracted by the encoder into a high-resolution pixel-level prediction map, that is, it outputs a semantic segmentation result of the same size as the input image.

[0085] In this implementation, the SegFormer-B1 model is a lightweight model based on the Transformer architecture, which realizes efficient and accurate attention score extraction and region positioning functions.

[0086] In this implementation, due to the limited computing power of the drone edge device, a lightweight SegFormer-B1 model is used to extract the attention scores of image blocks.

[0087] In this embodiment, the SegFormer-B1 model includes a Patch Embedding layer, which is used to divide the input low-definition large-scale image into multiple image blocks.

[0088] In this embodiment, the SegFormer-B1 model adopts a layered approach to capture information of different scales in the image through step-by-step downsampling and convolution operations; After feature extraction, the SegFormer-B1 model further processes the extracted features using a Transformer encoder to capture long-range dependencies between features.

[0089] In this embodiment, as shown in the following figure: Figure 2 The SegFormer-B1 model includes 4 stages, each stage including 2 Transformer blocks, each Transformer block including an attention module and a multi-layer perceptron (MLP).

[0090] The attention module is used to model long-range dependencies in space through self-attention mechanisms to capture global context information of the image.

[0091] The multi-layer perceptron is used to perform nonlinear transformation of the channel dimension at each spatial location to enhance the expressive power and discriminability of the features.

[0092] In this embodiment, the low-resolution large-scale image input into the SegFormer-B1 model is 512x512x3. Note that "512x512x3" means "512 pixels x 512 pixels x 3 channels".

[0093] The Patch Embedding layer divides the input low-resolution large-scale image into 128x128 image blocks.

[0094] Each image block is embedded into a 64-dimensional feature space through convolution operations.

[0095] The SegFormer-B1 model includes 4 stages, each stage including 2 Transformer blocks.

[0096] The 4 stages are used to generate attention score matrices of different resolutions (e.g., 128x128, 64x64, 32x32, 16x16); each Transformer block generates an attention score matrix, and each stage includes 2 Transformer blocks, so each stage generates 2 attention score matrices, specifically: The 4 stages are used to obtain 2 128x128, 2 64x64, 2 32x32, and 2 16x16 attention score matrices, respectively.

[0097] Note that "128x128" means "128 pixels x 128 pixels", and the units of "64x64", "32x32", and "16x16" are all pixels.

[0098] ​The final attention score of each image block is related to the attention score matrix of each stage. Therefore, it is necessary to extract and fuse multi-scale attention information through Transformer blocks of different resolutions, as follows: Step 1: Average the two attention score matrices obtained at each stage to fuse the attention information of the two Transformer blocks at each stage, where: For the first three stages, the average values ​​are taken with windows of size 8×8, 4×4, and 2×2, respectively, so as to convert the attention score matrices of the first three stages into 16×16 averaged attention score matrices; For the fourth stage, the average of the two attention score matrices is directly taken to obtain the 16×16 averaged attention score matrix of the fourth stage; Step 2: Perform weighted summation on the 16×16 averaged attention score matrix obtained in the four stages to obtain a global attention matrix containing information at different scales; Step 3: Use an 8×8 window to average the global attention matrix containing information of different scales to obtain the final 2×2 attention matrix.

[0099] Finally, according to the final 2×2 attention matrix, the attention score of each image block is sorted from high to low, and the GPS location information corresponding to the image blocks in the front is sent to the following drone for high-precision semantic segmentation.

[0100] In the third embodiment, in the multi-source information fusion step, the multi-source information fusion is performed with the initial semantic segmentation result as follows: The initial semantic segmentation result is a coarse-grained segmentation result; The high-precision semantic segmentation result is a fine-grained segmentation result; The coarse-grained segmentation results are combined with the fine-grained segmentation results by direct coverage or probability comparison to generate the final semantic segmentation results.

[0101] In this embodiment, the initial semantic segmentation model is used to perform initial semantic segmentation on the low-definition large-scale image to obtain a coarse-grained segmentation result (i.e., the initial semantic segmentation result). The coarse-grained segmentation result contains global context information. High-precision semantic segmentation models are used to perform high-precision semantic segmentation on high-definition, small-scale images of areas requiring high-precision processing, obtaining fine-grained segmentation results (i.e., high-precision semantic segmentation results). The fine-grained segmentation results contain local detail information. Combining the global information modeling capability of the initial semantic segmentation model and the local detail capture capability of the high-precision semantic segmentation model, the advantages of heterogeneous models are fully utilized, significantly improving the accuracy and efficiency of semantic segmentation.

[0102] In this embodiment, through the collaboration of multiple UAVs, the fine-grained prediction (i.e., fine-grained segmentation result) of important image blocks (i.e., areas that require high-precision processing) is used to enhance the coarse-grained prediction (i.e., coarse-grained segmentation result).

[0103] In this embodiment, the initial semantic segmentation result (i.e., the coarse-grained segmentation result) includes the predicted category and probability information of each pixel; Similarly, the high-precision semantic segmentation result (i.e., fine-grained segmentation result) includes the predicted category and probability information of each pixel.

[0104] In addition, in one embodiment, a direct overlay method is used to combine the coarse-grained segmentation results with the fine-grained segmentation results to generate the final semantic segmentation results: The category predicted for each pixel in the high-precision semantic segmentation result is directly overlaid on the initial semantic segmentation result to obtain the final semantic segmentation result.

[0105] In addition, in one embodiment, the coarse-grained segmentation result is combined with the fine-grained segmentation result by using a probability comparison method to generate the final semantic segmentation result: Compare the probability information of the category predicted for each pixel in the high-precision semantic segmentation result and the category predicted for each pixel in the initial semantic segmentation result, select the category predicted for each pixel with high probability, and obtain the final semantic segmentation result.

[0106] In the fourth embodiment, the adaptation steps during the pilot drone test include: Layer normalization statistics calculation steps: Iteratively calculate the runtime mean and variance of new images collected by the Lingfei UAV in the channel and spatial dimensions; Statistics update step: Based on the runtime mean and variance of the new images collected by the Lingfei UAV in the channel and spatial dimensions, the exponential moving average method is used to update the statistics in real time.

[0107] In this implementation, for the initial semantic segmentation model of the pilot drone: The layer normalization (LN) adaptation method is used (i.e., the layer normalization statistics calculation step), and the statistics are updated by the exponential moving average (EMA) (i.e., the statistics updating step).

[0108] In addition, in one embodiment, the runtime mean and variance of the new image collected by the pilot drone in the channel and spatial dimensions are calculated iteratively as follows: Assume that the input feature map corresponding to the new image collected by the Lingfei UAV is , whose dimensions are ;in, is the number of channels, is the height of the input feature map, is the width of the input feature map; Then at iteration t, the runtime mean calculated over the channel and spatial dimensions is:

[0109] Then at the tth iteration, the variance calculated in the channel and spatial dimensions is: .

[0110] In addition, in one embodiment, based on the runtime mean and variance of the new image collected by the pilot drone in the channel and spatial dimensions, the exponential moving average method is used to update the statistics in real time as follows:

[0111] .

[0112] In this embodiment, the exponential moving average method is referred to as EMA.

[0113] In this embodiment, the statistics updated in real time by the leading UAV use the statistics of the training data of the leading UAV as the initialization statistics.

[0114] A fifth embodiment provides a method for on-demand dynamic collaboration of tasks for heterogeneous multi-agents, wherein the heterogeneous multi-agents include a leading drone and multiple following drones, and fly in a master-slave formation flight mode with the leading drone as the main body. The method includes the following steps: Steps for receiving coordinate information: any follower drone receives the coordinate information of the area requiring high-precision processing assigned by the leader drone; Steps for obtaining small-scale images: any following drone obtains high-definition small-scale images of the area that needs high-precision processing based on the coordinate information of the area that needs high-precision processing; High-precision semantic segmentation step: Any of the following drones uses a high-precision semantic segmentation model to perform high-precision semantic segmentation on the high-definition small-scale images of the area that requires high-precision processing, and obtains high-precision semantic segmentation results; Returning high-precision results: Any follower drone returns the high-precision semantic segmentation results to the leader drone. Adaptation step during follow-up drone testing: Any follow-up drone dynamically updates the high-precision semantic segmentation model by fusing the statistics of newly collected images from other follow-up drones in real time to adapt to cross-device testing.

[0115] In this embodiment, the high-definition small-area image and the low-definition large-area image mentioned above are not absolute numerical descriptions, that is, they do not limit the definition of high-definition or the range of large-area. They are relative expressions, namely: Compared with low-definition large-scale images, high-definition small-scale images are clearer in definition and have a smaller shooting range.

[0116] In this embodiment, a camera is also installed on the following drone for collecting high-definition small-scale images.

[0117] In this embodiment, the following drones are all equipped with low-cost cameras for collecting images at a relatively low cost.

[0118] In this embodiment, if Figure 1 As shown in Figure 1, a flow chart of the on-demand dynamic collaboration method for heterogeneous multi-agent tasks is shown, including the processing flow of the leading drone and the following drone, where: Drones are also called edge devices. In the figure, "edge side" refers to the computing modules deployed on the drone. For example, the "leading drone edge side" deploys a "high-performance computing module," while the "following drone edge side" deploys a "medium-performance computing module."

[0119] In the figure, "model adaptation" refers to "adaptation during cross-device testing."

[0120] The “high-altitude image capture” in the figure refers to the leader drone capturing images at high altitudes to collect low-definition and large-scale images.

[0121] The “attention-based image patch selection” in the figure refers to using the initial semantic segmentation model to perform initial semantic segmentation and identify areas that require high-precision processing.

[0122] The “initial segmentation result” in the figure refers to the initial semantic segmentation result.

[0123] The “output attention score” in the figure refers to extracting the attention score of each image block to identify areas that require high-precision processing.

[0124] The “information fusion” in the figure refers to multi-source information fusion.

[0125] The “final segmentation score” in the figure refers to the final semantic segmentation result.

[0126] The "GPS coordinate set" in the figure refers to the coordinate information (i.e., GPS location information) of the area that requires high-precision processing.

[0127] The “GPS-based low-altitude image capture” in the figure refers to obtaining high-definition small-scale images of the area that needs high-precision processing based on the coordinate information of the area that needs high-precision processing.

[0128] The “CNN-based semantic segmentation” in the figure refers to the use of a high-precision semantic segmentation model for high-precision semantic segmentation.

[0129] The "improved segmentation results" in the figure refer to high-precision semantic segmentation results.

[0130] In the sixth embodiment, the high-precision semantic segmentation model is a CNN-based semantic segmentation model.

[0131] In addition, in one embodiment, the CNN-based semantic segmentation model is a DeepLabv3+ model, which is used to process high-definition small-scale images and generate local detail information.

[0132] In the seventh embodiment, the adaptation steps during the following drone test include: Batch normalization statistics calculation: Iteratively calculate the runtime mean and variance of each channel in the spatial dimension for each new image collected by the following drone; Each following drone sends the runtime mean and variance of each channel of the collected new image obtained by its iterative calculation in the spatial dimension to other following drones; Statistics sharing and fusion: Each following drone receives the runtime mean and variance of each channel in the spatial dimension of the new images collected by other following drones and stores them in the memory pool; based on the data stored in the memory pool, the exponential sliding average method is used to update the statistics.

[0133] In this implementation, for the high-precision semantic segmentation model of the following drone: The batch normalization (BN) adaptation method (i.e., the batch normalization statistics calculation step) is used to achieve dynamic adjustment through statistics sharing and fusion among multiple devices (i.e., the statistics sharing and fusion step).

[0134] In this embodiment, the adaptability of the model is improved by transmitting and fusing statistics among multiple devices (ie, statistics sharing and fusion steps).

[0135] In addition, in one embodiment, the runtime mean and variance of each channel in the spatial dimension for each new image collected by the following drone are iteratively calculated as follows: Assume that the input feature map corresponding to each new image collected by the flying drone is , whose dimensions are ;in, is the number of channels, is the height of the input feature map, is the width of the input feature map; Assume that each following drone is A follow-up drone ; Then at the t-th iteration, for each channel The runtime mean computed over the spatial dimension is:

[0136] Then at the t-th iteration, for each channel The variance computed over the spatial dimension is: .

[0137] In addition, in an embodiment, each follower drone receives the runtime mean and variance of each channel over the spatial dimension of the collected new image sent by other follower drones, and stores them into a memory pool; based on the data stored in the memory pool, the statistics are updated using the exponential moving average method as follows:

[0138]

[0139] where the parameter controls the memory length.

[0140] In the embodiment, the parameter takes a small value to support long-term memory, and takes a large value to support short-term memory.

[0141] In the embodiment, the statistics updated by each follower drone in real time use the statistics of the training data of each follower drone as the initial statistics.

[0142] Embodiment eight provides a task on-demand dynamic cooperation device for heterogeneous multi-agent, the heterogeneous multi-agent includes a leader drone and a plurality of follower drones, and flies in a master-slave formation flight mode with the leader drone as the master; the device includes the following modules: A global task planning module: the leader drone formulates a global task plan according to task requirements and its own sensing capability; An acquisition of large-range image module: the leader drone acquires a low-quality large-range image collected according to the global task plan; A dynamic allocation module: the leader drone performs initial semantic segmentation on the low-quality large-range image using an initial semantic segmentation model, obtains an initial semantic segmentation result, and identifies a region requiring high-precision processing, and dynamically allocates coordinate information of the region requiring high-precision processing to a follower drone; A multi-source information fusion module: the leader drone receives a high-precision semantic segmentation result returned by any one follower drone, performs multi-source information fusion on the high-precision semantic segmentation result and the initial semantic segmentation result, and obtains a final semantic segmentation result; The leader UAV is adapted to test the module: the leader UAV dynamically updates the initial semantic segmentation model by updating the statistics of the new images collected in real time, to adapt to cross-device testing.

[0143] Embodiment nine provides a task on-demand dynamic cooperation device for a heterogeneous multi-agent, the heterogeneous multi-agent including a leader UAV and a plurality of follower UAVs, and flying in a master-slave formation flight mode with the leader UAV as the master; the device includes the following modules: A receiving coordinate information module: any one of the follower UAVs receives coordinate information of an area requiring high-precision processing distributed by the leader UAV; A small-range image acquisition module: any one of the follower UAVs acquires a high-definition small-range image of the area requiring high-precision processing according to the coordinate information of the area requiring high-precision processing; A high-precision semantic segmentation module: any one of the follower UAVs performs high-precision semantic segmentation on the high-definition small-range image of the area requiring high-precision processing using a high-precision semantic segmentation model to obtain a high-precision semantic segmentation result; A high-precision result returning module: any one of the follower UAVs returns the high-precision semantic segmentation result to the leader UAV; A follower UAV test adaptation module: any one of the follower UAVs dynamically updates the high-precision semantic segmentation model by fusing statistics of new images collected by other follower UAVs in real time, to adapt to cross-device testing.

[0144] Embodiment ten provides an electronic device, including: One or more processors; A storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the task on-demand dynamic cooperation method for a heterogeneous multi-agent as described in any one of embodiments one to four or the task on-demand dynamic cooperation method for a heterogeneous multi-agent as described in any one of embodiments five to seven.

[0145] Embodiment eleven provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, when the computer program is executed by a processor, implementing the task on-demand dynamic cooperation method for a heterogeneous multi-agent as described in any one of embodiments one to four or the task on-demand dynamic cooperation method for a heterogeneous multi-agent as described in any one of embodiments five to seven.

[0146] In a twelfth embodiment, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the on-demand dynamic collaboration method for heterogeneous multi-agent tasks as described in any one of embodiments one to four or the on-demand dynamic collaboration method for heterogeneous multi-agent tasks as described in any one of embodiments five to seven.

[0147] In a thirteenth embodiment, a drone system is provided, comprising a leading drone and a plurality of following drones, the drone system flying in a master-slave formation flight mode with the leading drone as the main member; The leading UAV is provided with a computing module for implementing the task-on-demand dynamic collaboration method for heterogeneous multi-agents described in any one of implementation modes one to four; The following drone is provided with a computing module for implementing the task-on-demand dynamic collaboration method for heterogeneous multi-agents as described in any one of implementation modes five to seven.

[0148] Fourteenth embodiment: A comparative experiment is conducted using a method based on SegFormer-B1, a method based on DeepLabv3+, and the present invention (referred to as this solution).

[0149] From the experimental results, we can see that: The framework constructed in this solution (a method for on-demand dynamic collaboration of heterogeneous multi-agent tasks) demonstrates excellent performance on both SDD (Table 1) and FloodNet (Table 2) datasets. This multi-UAV collaborative framework enables parallel execution of airborne semantic segmentation tasks, significantly reducing the inference latency and computational burden of a single UAV. Compared to single-UAV systems relying on high-cost sensors, this approach achieves higher segmentation accuracy using only low-cost sensors.

[0150] Table 1: Performance of this solution on the SDD dataset

[0151] Table 2: Performance of this solution on the FloodNet dataset

[0152] In terms of inference latency, a single drone can complete an onboard semantic segmentation task with a latency exceeding 500ms, and even exceeding 1800ms using the SegFormer model, hindering rapid decision-making for urgent tasks. This invention reduces single-device inference latency by approximately 3-5 times.

[0153] Due to the master-slave formation flight mode, the distance between drones is relatively close, the data transmission volume is small, and real-time missions are supported.

[0154] In terms of segmentation accuracy, our method shows significant improvement on uncorrupted images. Compared with the best baseline method: The segmentation accuracy on the SDD dataset was improved by 5.9%; It improves by 3.3% on the FloodNet dataset.

[0155] On the dataset with dynamic environment destruction, the performance improvement is even more significant: On the SDD-C dataset with a damage level of 5, the segmentation accuracy was improved by an average of 11.0%; in foggy weather, it even increased by 15.0%. On the FloodNet-C dataset with damage level 3, the segmentation accuracy was improved by an average of 6.5%; and by 9.4% in foggy weather.

[0156] In the above implementation manner, the abbreviated Chinese / English full names are explained as follows:

[0157] A computer device or system is provided in the above embodiment. The hardware device of this part is a general model and is not shown in the form of a diagram. The system includes a processor and a memory, wherein the processor and the memory can be connected through a bus or other means. The memory is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, and corresponding program instructions / modules. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory, so as to realize the data space entity resolution data quality enhancement method in the above method embodiment.

[0158] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, an intranet, a mobile communication network, and combinations thereof.

[0159] One or more modules are stored in the memory. When the processor executes, the method steps in the embodiment are executed. In this way, the purpose of the invention can be achieved through the method, device and process of the present invention. The specific details of the above-mentioned computer equipment can be understood by referring to the corresponding descriptions and effects in the embodiment, and will not be repeated here.

[0160] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD). The storage medium can also include a combination of the above-mentioned types of memory.

[0161] The above further describes the technical solution provided by the present invention in detail through several specific embodiments in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the several specific embodiments described above are not intended to limit the present invention. Any reasonable changes and improvements to the present invention, reasonable combinations of implementation methods and equivalent replacements based on the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A task-on-demand dynamic collaboration method for heterogeneous multi-agents, characterized by: The heterogeneous multi-agent comprises a leading UAV and a plurality of following UAVs, and flies in a master-slave formation flight mode with the leading UAV as the main body; the method comprises the following steps: Global mission planning step: The leading UAV formulates a global mission plan based on mission requirements and its own perception capabilities; The step of acquiring a large-scale image: the pilot drone acquires low-resolution large-scale images according to the global mission plan; Dynamic allocation step: The leading UAV uses the initial semantic segmentation model to perform initial semantic segmentation on the low-definition large-scale image, obtains the initial semantic segmentation result, identifies the area requiring high-precision processing, and dynamically allocates the coordinate information of the area requiring high-precision processing to the following UAV; Multi-source information fusion step: the leading UAV receives the high-precision semantic segmentation result returned by any following UAV, and performs multi-source information fusion on it with the initial semantic segmentation result to obtain the final semantic segmentation result; Adaptation step during the pilot drone test: The pilot drone dynamically updates the initial semantic segmentation model by updating the statistics of the new images it collects in real time to adapt during cross-device testing.

2. The task-on-demand dynamic collaboration method for heterogeneous multi-agents according to claim 1 is characterized in that: In the dynamic allocation step, the areas requiring high-precision processing are identified as follows: The initial semantic segmentation model divides the initial semantic segmentation model into multiple image blocks and extracts the attention score of each; Multiple image blocks are sorted from high to low according to the attention scores, and the image blocks that are ranked in front by a given number are regarded as areas that need high-precision processing.

3. The task-on-demand dynamic collaboration method for heterogeneous multi-agents according to claim 1 is characterized in that: In the multi-source information fusion step, the multi-source information fusion is performed with the initial semantic segmentation result as follows: The initial semantic segmentation result is a coarse-grained segmentation result; The high-precision semantic segmentation result is a fine-grained segmentation result; The coarse-grained segmentation results are combined with the fine-grained segmentation results by direct coverage or probability comparison to generate the final semantic segmentation results.

4. The task-on-demand dynamic collaboration method for heterogeneous multi-agents according to claim 1 is characterized in that: The adaptation steps during the pilot drone test include: Layer normalization statistics calculation steps: Iteratively calculate the runtime mean and variance of new images collected by the Lingfei UAV in the channel and spatial dimensions; Statistics update step: Based on the runtime mean and variance of the new images collected by the Lingfei UAV in the channel and spatial dimensions, the exponential moving average method is used to update the statistics in real time.

5. A task-on-demand dynamic collaboration method for heterogeneous multi-agents, characterized by: The heterogeneous multi-agent comprises a leading UAV and a plurality of following UAVs, and flies in a master-slave formation flight mode with the leading UAV as the main body; the method comprises the following steps: Steps for receiving coordinate information: any follower drone receives the coordinate information of the area requiring high-precision processing assigned by the leader drone; Steps for obtaining small-scale images: any following drone obtains high-definition small-scale images of the area that needs high-precision processing based on the coordinate information of the area that needs high-precision processing; High-precision semantic segmentation step: Any of the following drones uses a high-precision semantic segmentation model to perform high-precision semantic segmentation on the high-definition small-scale images of the area that requires high-precision processing, and obtains high-precision semantic segmentation results; Returning high-precision results: Any follower drone returns the high-precision semantic segmentation results to the leader drone. Adaptation step during follow-up drone testing: Any follow-up drone dynamically updates the high-precision semantic segmentation model by fusing the statistics of newly collected images from other follow-up drones in real time to adapt to cross-device testing.

6. The task-on-demand dynamic collaboration method for heterogeneous multi-agents according to claim 1 is characterized in that: The high-precision semantic segmentation model is a semantic segmentation model based on CNN.

7. The task-on-demand dynamic collaboration method for heterogeneous multi-agents according to claim 5 is characterized in that: The adaptation steps during the following drone test include: Batch normalization statistics calculation: Iteratively calculate the runtime mean and variance of each channel in the spatial dimension for each new image collected by the following drone; Each following drone sends the runtime mean and variance of each channel of the collected new image obtained by its iterative calculation in the spatial dimension to other following drones; Statistics sharing and fusion: Each following drone receives the runtime mean and variance of each channel in the spatial dimension of the new images collected by other following drones and stores them in the memory pool; based on the data stored in the memory pool, the exponential sliding average method is used to update the statistics.

8. A task-on-demand dynamic collaboration device for heterogeneous multi-agents, characterized by: The heterogeneous multi-agent comprises a leading UAV and multiple following UAVs, and flies in a master-slave formation flight mode with the leading UAV as the main one. The device comprises the following modules: Global mission planning module: The pilot drone formulates a global mission plan based on mission requirements and its own perception capabilities; Acquisition of large-scale image module: The pilot drone acquires low-resolution large-scale images according to the global mission planning; Dynamic allocation module: The leading UAV uses the initial semantic segmentation model to perform initial semantic segmentation on the low-definition large-scale image, obtains the initial semantic segmentation results, identifies the areas that require high-precision processing, and dynamically allocates the coordinate information of the areas that require high-precision processing to the following UAV; Multi-source information fusion module: The leading UAV receives the high-precision semantic segmentation result returned by any following UAV, and performs multi-source information fusion with the initial semantic segmentation result to obtain the final semantic segmentation result; The Lingfei UAV test-time adaptation module: The Lingfei UAV dynamically updates the initial semantic segmentation model by updating the statistics of the new images it collects in real time to adapt to cross-device testing.

9. A task-on-demand dynamic collaboration device for heterogeneous multi-agents, characterized by: The heterogeneous multi-agent comprises a leading UAV and multiple following UAVs, and flies in a master-slave formation flight mode with the leading UAV as the main one. The device comprises the following modules: Coordinate information receiving module: any follower drone receives the coordinate information of the area that requires high-precision processing assigned by the leader drone; Small-area image acquisition module: any following drone can acquire high-definition small-area images of the area that needs high-precision processing based on the coordinate information of the area that needs high-precision processing; High-precision semantic segmentation module: Any following drone uses a high-precision semantic segmentation model to perform high-precision semantic segmentation on high-definition small-scale images of areas that require high-precision processing, obtaining high-precision semantic segmentation results; High-precision result return module: Any follow-up drone returns high-precision semantic segmentation results to the leading drone; Adaptation module for following drones during testing: Any following drone dynamically updates the high-precision semantic segmentation model by integrating the statistics of newly collected images from other following drones in real time to adapt to cross-device testing.

10. An electronic device, characterized in that: include: one or more processors; A storage device configured to store one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the task-on-demand dynamic collaboration method for heterogeneous multi-agents as described in any one of claims 1 to 4 or the task-on-demand dynamic collaboration method for heterogeneous multi-agents as described in any one of claims 5 to 7.