Space cooperation mapping and multimodal instruction distribution system and method for air-ground heterogeneous multi-agent cluster

By using a spatial collaborative mapping and multimodal command distribution system for air-ground heterogeneous multi-agent clusters, the problems of collaborative mapping and control adaptation between air-ground heterogeneous agents are solved. This system enables unified expression of air viewpoint and ground feedback and accurate distribution of multimodal commands, thereby improving the accuracy and stability of collaboration in complex areas.

CN122431402APending Publication Date: 2026-07-21GLOBALTOUR GROUP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GLOBALTOUR GROUP LTD
Filing Date
2026-06-08
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing multi-agent cooperative control technologies lack a collaborative mapping mechanism between heterogeneous air and ground agents, and do not fully integrate the contact force, roll, yaw and travel obstruction information of ground agents, resulting in unstable judgment from the air perspective and difficulty in adapting control commands to different execution objects, affecting the accuracy and reliability of cooperative tasks.

Method used

A spatial collaborative mapping and multimodal command distribution system for air-ground heterogeneous multi-agent clusters is adopted. The probabilistic grid information of airborne agents is obtained through airborne environmental perception terminals, and the motion and contact feedback information of ground agents is obtained through ground-based state feedback terminals. These are mapped to the same generalized spatial coordinate system to generate heterogeneous collaborative spatial mapping information. The connection points are calculated by combining hovering constraints and passage constraints, and multimodal control commands are generated and distributed.

Benefits of technology

It improves the collaboration accuracy and execution reliability of heterogeneous air-ground multi-agent clusters in complex target operation areas, solves the problems of inconsistent spatial representation and control command adaptation among heterogeneous air-ground agents, and enhances the collaboration accuracy and stability of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431402A_ABST
    Figure CN122431402A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of multi-agent cooperative control, and provides a space cooperation mapping and multi-modal instruction distribution system and method for an air-ground heterogeneous multi-agent cluster, which comprises an air-based environment perception terminal, a ground-based state feedback terminal, a heterogeneous space mapping terminal, a cooperation connection calculation terminal, a multi-modal instruction generation terminal and a cluster instruction distribution terminal; the system is used for performing object matching, time sequence arrangement and distribution execution on the flight attitude control instruction, the ground motion control instruction and the operation action control instruction according to the equipment types, task roles, current positions and communication states of aerial agents and ground agents, so that the aerial agents and the ground agents complete a space cooperation task at a position corresponding to the target connection point information. The application has the effects of improving the cooperation accuracy, execution stability and task adaptability when the air-ground heterogeneous multi-agent cluster executes a cooperation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of multi-agent cooperative control, specifically to a spatial cooperative mapping and multimodal command distribution system and method for air-ground heterogeneous multi-agent clusters. Background Technology

[0002] With the development of drones, mobile robots, legged robots, automated guided vehicles (AGVs), and edge computing devices, multi-agent swarms have been increasingly applied in scenarios such as disaster relief, regional inspection, warehousing and transshipment, complex factory operations, mining area inspections, agricultural and forestry operations, and port yard scheduling. Compared to single-agent tasks, multi-agent swarms can improve task coverage and response efficiency through spatial division of labor, information sharing, and collaborative execution. Aerial agents typically possess wide field of view, large coverage area, and high environmental perception efficiency, making them suitable for overhead scanning of target work areas, terrain recognition, and obstacle distribution assessment. Ground agents typically have strong close-range operation capabilities, can support end-effectors, and can directly contact the ground and provide feedback on traffic status, making them suitable for tasks such as handling, grasping, detection, delivery, obstacle removal, or close-range operations. Therefore, in complex work areas, the collaborative control of aerial and ground agents has become an important development direction for multi-agent collaboration technology.

[0003] Existing technologies already include collaborative control schemes for robot swarms or drone swarms. For example, Chinese invention patent application CN117359639A discloses a collaborative control method and system for robot swarms. This scheme acquires the degree-of-freedom information of individual robots, determines the individual motion structure model, determines the connection point parameters between adjacent individual robots according to the process sequence of the target task, connects the individual motion structure models in a unified spatial coordinate system, and then reverse-calculates the stage check points and maps the collaborative control commands to different robots. This scheme can improve the accuracy of robot swarm motion control, but its focus is on constructing a swarm motion control model based on the degree-of-freedom information of individual robots and completing the motion structure connection between robots according to the connection point parameters. For task scenarios involving both aerial and ground-based intelligent agents, this scheme does not further consider the probabilistic grid information formed by the aerial intelligent agent's top-down perception, nor does it involve the ground-based feedback information such as contact force, tilt, yaw, and travel obstruction formed by the ground-based intelligent agent during actual passage. Therefore, it is difficult to directly solve the spatial collaborative mapping problem between heterogeneous aerial and ground-based intelligent agents caused by differences in motion mode, perception scale, and execution interface.

[0004] Therefore, while existing multi-agent cooperative control technologies cover robot swarm motion control, UAV swarm cooperative search, and UAV swarm scheduling, they still have the following shortcomings: First, existing solutions mostly focus on cooperation between homogeneous agents, or only perform task scheduling between UAVs and ground stations or UAV pods, lacking heterogeneous spatial cooperative mapping mechanisms for aerial agents and ground-based mobile agents; second, existing solutions typically use aerial perception results as the basis for search or localization, failing to fully incorporate contact forces, roll, yaw, and travel obstruction information generated by ground agents during actual passage, leading to inaccurate judgments of usable locations from an aerial perspective. Third, existing docking or collaborative control schemes typically focus on geometric location, degree of freedom connection, or task splitting, without incorporating air hovering constraints, ground access constraints, air-to-ground arrival time difference, end-effector execution difficulty, and communication link status into the target docking point determination process. Fourth, existing command distribution methods mostly issue control commands according to task objects or preset paths, making it difficult to convert the same target docking point information into separate flight attitude control commands for airborne intelligent agents, motion control commands for ground intelligent agents, and operational action control commands for end-effectors or mounted mechanisms, and further perform object matching, timing arrangement, and distribution confirmation.

[0005] Therefore, it is necessary to provide a spatial collaborative mapping and multimodal command distribution system and method for air-ground heterogeneous multi-agent clusters. This system can uniformly map the airborne probabilistic grid information acquired by airborne agents and the ground-based motion state information and ground-based contact feedback information acquired by ground agents to the same generalized spatial coordinate system. This forms heterogeneous collaborative spatial mapping information that can characterize the spatial reachability of airborne agents, the travel stability of ground agents, and the positional relationship of air-ground collaboration. Based on this, the system can solve the target docking point information, generate and distribute multimodal control commands adapted to different agents and actuators, thereby improving the accuracy and reliability of collaboration among air-ground heterogeneous multi-agent clusters in complex target operation areas. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of the above-mentioned solutions by proposing a spatial cooperative mapping and multimodal instruction distribution system and method for air-ground heterogeneous multi-agent clusters.

[0007] The present invention adopts the following technical solution:

[0008] In a first aspect, the present invention discloses a spatial cooperative mapping and multimodal command distribution system for an air-ground heterogeneous multi-agent cluster, including an air-based environmental perception terminal, a ground-based state feedback terminal, a heterogeneous spatial mapping terminal, a cooperative connection and calculation terminal, a multimodal command generation terminal, and a cluster command distribution terminal.

[0009] The airborne environmental perception terminal is used to acquire airborne environmental perception information obtained by the aerial intelligent agent through overhead perception of the target operation area, and to generate airborne probability grid information based on the airborne environmental perception information to characterize the terrain distribution, obstacle distribution and passability probability distribution of the target operation area.

[0010] The ground state feedback terminal is used to acquire ground motion state information and ground contact feedback information generated by the ground intelligent agent during its movement or operation within the target work area; the ground motion state information includes at least the pose information, tilt state information and yaw state information of the ground intelligent agent, and the ground contact feedback information includes at least the contact force information and travel resistance information generated during the contact between the ground intelligent agent and the ground;

[0011] The heterogeneous spatial mapping terminal is connected to the airborne environment sensing terminal and the ground-based state feedback terminal respectively, and is used to map the airborne probability grid information, the ground-based motion state information and the ground-based contact feedback information to the same generalized spatial coordinate system to generate heterogeneous cooperative spatial mapping information for characterizing the spatial reachability of airborne intelligent agents, the travel stability of ground intelligent agents and the positional relationship of air-ground cooperation.

[0012] The collaborative docking solution terminal is connected to the heterogeneous space mapping terminal and is used to determine candidate docking areas between the air agent and the ground agent that meet the cooperation conditions based on the heterogeneous collaborative space mapping information, combined with the hovering constraints of the air agent, the passage constraints of the ground agent, and the air-ground collaborative task constraints. The target docking point information is calculated based on the spatial reachability, passage stability, and collaborative execution cost corresponding to different locations in the candidate docking areas.

[0013] The multimodal instruction generation terminal is connected to the collaborative connection and resolution terminal, and is used to generate flight attitude control instructions adapted to the execution of the airborne intelligent agent and ground motion control instructions adapted to the execution of the ground intelligent agent according to the target connection point information. When the collaborative task involves end operation, it generates operation action control instructions adapted to the end execution mechanism of the ground intelligent agent or the mounted mechanism of the airborne intelligent agent.

[0014] The cluster instruction distribution terminal is connected to the multimodal instruction generation terminal and is used to perform object matching, timing arrangement, and distribution execution of the flight attitude control instructions, the ground motion control instructions, and the operation action control instructions according to the device type, task role, current location, and communication status of the airborne intelligent agent and the ground intelligent agent, so that the airborne intelligent agent and the ground intelligent agent can complete the spatial cooperation task at the location corresponding to the target docking point information.

[0015] Secondly, this invention also discloses a spatial cooperative mapping and multimodal command distribution method for air-to-ground heterogeneous multi-agent clusters, applied to the aforementioned spatial cooperative mapping and multimodal command distribution system for air-to-ground heterogeneous multi-agent clusters. The spatial cooperative mapping and multimodal command distribution method for air-to-ground heterogeneous multi-agent clusters includes:

[0016] S1, acquire airborne environmental perception information obtained by the aerial agent from the top-down perception of the target operation area, and generate airborne probability grid information representing the terrain distribution, obstacle distribution and passability probability distribution of the target operation area based on the airborne environmental perception information.

[0017] S2, acquire ground motion status information and ground contact feedback information generated by the ground intelligent agent during its movement or operation within the target work area;

[0018] S3, map the airborne probability grid information, the ground motion state information and the ground contact feedback information to the same generalized spatial coordinate system to generate heterogeneous cooperative spatial mapping information for characterizing the spatial reachability of airborne intelligent agents, the travel stability of ground intelligent agents and the positional relationship of air-ground cooperation;

[0019] S4. Based on the heterogeneous collaborative space mapping information, combined with the hovering constraints of the airborne intelligent agent, the passage constraints of the ground intelligent agent, and the air-ground collaborative task constraints, candidate docking areas that meet the collaboration conditions between the airborne intelligent agent and the ground intelligent agent are determined, and the target docking point information is calculated according to the spatial reachability, passage stability, and collaboration execution cost corresponding to different locations in the candidate docking areas.

[0020] S5. Based on the target docking point information, generate flight attitude control commands adapted to the execution of the airborne intelligent agent and ground motion control commands adapted to the execution of the ground intelligent agent, and generate operation action control commands adapted to the end-effector of the ground intelligent agent or the mounting mechanism of the airborne intelligent agent when the collaborative task involves end-effector operation.

[0021] S6. Based on the device type, task role, current location, and communication status of the airborne intelligent agent and the ground intelligent agent, perform object matching, timing arrangement, and distribution execution of the flight attitude control command, the ground motion control command, and the operation action control command, so that the airborne intelligent agent and the ground intelligent agent can complete the spatial cooperation task at the location corresponding to the target docking point information.

[0022] The beneficial effects achieved by this invention are:

[0023] This invention acquires airborne environmental perception information from an aerial agent's top-down perception of the target work area via an airborne environmental perception terminal. Based on this information, it generates airborne probability grid information representing the terrain distribution, obstacle distribution, and passability probability distribution of the target work area. This allows the system to utilize the aerial agent's high-altitude perspective to provide a global or local spatial representation of the target work area. Compared to methods relying solely on local detection by ground-based agents, this approach expands the environmental perception range and allows for early acquisition of obstacle distribution, terrain types, and passability probabilities within the target work area, providing airborne spatial data for subsequent candidate connection area selection.

[0024] This invention acquires ground motion and contact feedback information generated by a ground-based intelligent agent during its movement or operation within the target work area via a ground-based state feedback terminal. It incorporates pose, tilt, yaw, contact force, and travel resistance information into the collaborative judgment process, enabling the system to correct potential misjudgments of ground conditions from an aerial perspective using the actual travel feedback from the ground-based intelligent agent. For soft ground, gravel areas, hidden ditches, localized slopes, areas with insufficient adhesion, or micro-topographical changes that are difficult to accurately assess visually, this invention can reflect the true travel status through ground contact feedback and attitude disturbance information, thereby reducing the risk of selecting docking points in unstable ground areas.

[0025] This invention maps airborne probabilistic grid information, ground-based motion state information, and ground-based contact feedback information to the same generalized spatial coordinate system through a heterogeneous spatial mapping terminal, generating heterogeneous collaborative spatial mapping information. This allows the spatial accessibility of airborne agents, the travel stability of ground-based agents, and the positional relationships of air-ground collaboration to be uniformly calculated within the same spatial representation. Therefore, airborne agents and ground-based agents no longer plan independently based on their respective local information, but can instead make collaborative judgments based on a unified spatial coordinate system, improving the consistency and computability of the spatial relationship representation of heterogeneous multi-agent air-ground systems.

[0026] This invention utilizes a collaborative docking terminal to determine candidate docking areas and calculate target docking point information based on heterogeneous collaborative spatial mapping information, combined with hovering constraints of airborne agents, travel constraints of ground agents, and air-to-ground collaborative task constraints. This eliminates reliance on shortest distance, single reachability, or preset docking locations for target docking point determination; instead, it comprehensively considers spatial reachability, ground travel stability, air-to-ground arrival time difference, collaborative execution cost, and the stability advantages of adjacent candidate locations. This approach reduces the possibility of isolated, low-cost locations being mistakenly selected as docking points, improving the stability and practical feasibility of target docking points in complex environments.

[0027] This invention utilizes a multimodal command generation terminal to generate flight attitude control commands, ground motion control commands, and operational action control commands based on target docking point information. This allows the same target docking point information to be converted into control content adapted to different execution objects. For airborne agents, it can generate commands related to trajectory adjustment, speed adjustment, and hovering attitude control; for ground agents, it can generate commands related to travel path, speed adjustment, and base attitude stabilization; and for end effectors or airborne mounted mechanisms, it can generate commands related to action position, action attitude, and action triggering timing. Thus, the system expands air-to-ground collaborative tasks from simple location arrival to a multimodal collaborative process where position, attitude, action, and timing are all controlled.

[0028] This invention utilizes a cluster command distribution terminal to perform object matching, timing arrangement, and distribution of flight attitude control commands, ground motion control commands, and operational action control commands according to the device type, task role, current location, and communication status of airborne and ground-based intelligent agents. This ensures that different control commands are accurately matched to intelligent agents or actuators with corresponding execution capabilities. Through link status monitoring and distribution confirmation mechanisms, this invention can also trigger redistribution or target connection point recalculation when there are command reception anomalies, communication link quality degradation, or changes in the status of the execution object, thereby improving the continuity and reliability of air-to-ground collaborative missions in communication-unstable environments.

[0029] In summary, this invention combines the wide-range overhead perception capability of airborne intelligent agents, the near-ground passage feedback capability of ground intelligent agents, the unified mapping capability of heterogeneous spaces, the dynamic calculation capability of target docking points, and the multimodal command distribution capability. It solves the problems in the prior art such as inconsistent spatial representation among airborne and ground-based heterogeneous intelligent agents, difficulty in incorporating real ground passage status into docking judgment, single basis for target docking point selection, and difficulty in adapting control commands to different execution objects. This improves the collaboration accuracy, execution stability, and task adaptability of airborne and ground-based heterogeneous multi-agent clusters when performing collaborative tasks such as delivery, grasping, detection, loading, inspection, or rescue in complex target operation areas.

[0030] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the overall structure of the present invention;

[0032] Figure 2 This is a schematic diagram of the method flow for the spatial cooperative mapping and multimodal instruction distribution method of the air-ground heterogeneous multi-agent cluster in this invention;

[0033] Figure 3 This is a statistical chart showing the distribution of air-ground connection adaptability and neighborhood stability advantage values ​​of candidate connection locations in Embodiment 2 of the present invention.

[0034] Figure 4 This is a statistical diagram of the air-ground connection adaptability-joint cooperation cost screening relationship in Embodiment 2 of the present invention. Detailed Implementation

[0035] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated in advance. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.

[0036] Example 1: This example provides a spatial cooperative mapping and multimodal command distribution system for air-to-ground heterogeneous multi-agent clusters. Combined with... Figure 1 As shown, the spatial collaborative mapping and multimodal command distribution system of the air-ground heterogeneous multi-agent cluster includes an air-based environmental perception terminal, a ground-based state feedback terminal, a heterogeneous spatial mapping terminal, a collaborative connection and calculation terminal, a multimodal command generation terminal, and a cluster command distribution terminal.

[0037] The airborne environmental perception terminal is used to acquire airborne environmental perception information obtained by the aerial intelligent agent through overhead perception of the target operation area, and to generate airborne probability grid information based on the airborne environmental perception information to characterize the terrain distribution, obstacle distribution and passability probability distribution of the target operation area.

[0038] The ground state feedback terminal is used to acquire ground motion state information and ground contact feedback information generated by the ground intelligent agent during its movement or operation in the target work area; the ground motion state information includes at least the pose information, tilt state information and yaw state information of the ground intelligent agent, and the ground contact feedback information includes at least the contact force information and travel resistance information generated during the contact between the ground intelligent agent and the ground.

[0039] The heterogeneous spatial mapping terminal is connected to the airborne environment sensing terminal and the ground-based state feedback terminal respectively. It is used to map airborne probabilistic grid information, ground-based motion state information and ground-based contact feedback information to the same generalized spatial coordinate system, and generate heterogeneous cooperative spatial mapping information to characterize the spatial reachability of airborne intelligent agents, the travel stability of ground intelligent agents and the positional relationship of air-ground cooperation.

[0040] The collaborative docking solution terminal is connected to the heterogeneous space mapping terminal. It is used to determine the candidate docking areas between the air agent and the ground agent that meet the cooperation conditions based on the heterogeneous collaborative space mapping information, combined with the hovering constraints of the air agent, the passage constraints of the ground agent, and the air-ground collaborative task constraints. Based on the spatial reachability, passage stability and collaborative execution cost of different locations in the candidate docking areas, the target docking point information is calculated.

[0041] The multimodal command generation terminal is connected to the collaborative docking and calculation terminal. It is used to generate flight attitude control commands adapted to the execution of airborne intelligent agents and ground motion control commands adapted to the execution of ground intelligent agents, respectively, based on the target docking point information. When the collaborative task involves end operations, it generates operation action control commands adapted to the end execution mechanism of the ground intelligent agent or the mounted mechanism of the airborne intelligent agent.

[0042] The cluster command distribution terminal is connected to the multimodal command generation terminal. It is used to perform object matching, timing arrangement and distribution of flight attitude control commands, ground motion control commands and operation action control commands according to the device type, task role, current position and communication status of the airborne intelligent agent and the ground intelligent agent, so that the airborne intelligent agent and the ground intelligent agent can complete the spatial cooperation task at the location corresponding to the target docking point information.

[0043] Optionally, the space-based environmental perception terminal includes a regional image acquisition module, a space-based pose synchronization module, and a probabilistic grid generation module;

[0044] The area image acquisition module is used to acquire image information of the target operation area collected by the aerial intelligent agent at different flight altitudes and different viewing angles;

[0045] The airborne pose synchronization module is used to acquire the position information, flight altitude information and attitude angle information of the airborne intelligent agent corresponding to the image information of the target operation area, and to synchronize the position information, flight altitude information and attitude angle information of the airborne intelligent agent with the image information of the target operation area in time.

[0046] The probabilistic raster generation module is used to divide the target operation area into rasteres based on the image information of the target operation area after time synchronization, and generate terrain category information, obstacle occupancy probability information and passability probability information in each raster cell to obtain the empty base probabilistic raster information.

[0047] Optionally, the ground condition feedback terminal includes a ground pose acquisition module, an inertial disturbance acquisition module, a contact force acquisition module, and a sluggish state recognition module;

[0048] The ground pose acquisition module is used to obtain the current position, orientation angle, and travel speed of the ground intelligent agent within the target operation area;

[0049] The inertial disturbance acquisition module is used to acquire information on the roll angle change, yaw angle change, and base vibration of the ground-based intelligent agent during its movement.

[0050] The contact force acquisition module is used to acquire the supporting contact force, driving contact force, or foot contact force between the ground-based intelligent agent and the ground.

[0051] The stagnation state recognition module is used to identify the stagnation state of the ground agent at the corresponding location based on the travel speed, roll angle change information, yaw angle change information and contact force information, and generate travel stagnation information.

[0052] Optionally, the heterogeneous spatial mapping terminal includes a coordinate unification module, a space-ground feature alignment module, and a collaborative spatial field generation module;

[0053] The coordinate unification module is used to convert the flight coordinates of the airborne intelligent agent, the grid coordinates of the airborne probability grid information, and the ground motion coordinates of the ground intelligent agent to the generalized spatial coordinate system.

[0054] The air-ground feature alignment module is used to spatially align the passability probability information, obstacle occupancy probability information, ground pose information, contact force information, and travel resistance information at the same or adjacent spatial locations.

[0055] The collaborative spatial field generation module is used to generate heterogeneous collaborative spatial mapping information, including air hovering suitability, ground traffic stability, and air-to-ground connection suitability, based on the spatially aligned information.

[0056] Optionally, the collaborative connection solution terminal includes a candidate area screening module, a collaborative cost calculation module, and a target connection point determination module;

[0057] The candidate area filtering module is used to filter candidate connection areas that meet the hovering conditions of airborne intelligent agents and the arrival conditions of ground intelligent agents from heterogeneous collaborative spatial mapping information based on the spatial accessibility of airborne intelligent agents, the travel stability of ground intelligent agents, and the location relationship of air-ground cooperation.

[0058] The collaboration cost calculation module is used to calculate the air hovering cost, ground passage cost, time synchronization cost, and collaboration execution cost for each candidate location in the candidate connection area.

[0059] The target docking point determination module is used to determine the target docking point information based on the cost of hovering in the air, the cost of ground passage, the cost of time synchronization, and the cost of collaborative execution. The target docking point information includes the target docking location, the target docking time window, and the target docking attitude requirements.

[0060] Optionally, the cooperation cost calculation module is also used to calculate the joint cooperation cost of the corresponding candidate location based on the flight energy consumption required for the airborne agent to reach the candidate location, the travel energy consumption required for the ground agent to reach the candidate location, the degree of ground contact disturbance, and the air-to-ground arrival time difference.

[0061] The target connection point determination module is used to determine the candidate positions that meet the preset connection constraints in terms of joint cooperation cost and have a stable advantage relative to adjacent candidate positions as the positions corresponding to the target connection point information.

[0062] Optionally, the multimodal command generation terminal includes an air command mapping module, a ground command mapping module, and an operation command mapping module;

[0063] The airborne command mapping module is used to generate trajectory adjustment commands, speed adjustment commands, and hovering attitude control commands for the airborne intelligent agent based on the target docking location, target docking time window, and target docking attitude requirements in the target docking point information.

[0064] The ground command mapping module is used to generate the ground intelligent agent's travel path command, speed adjustment command, and base attitude stabilization command based on the target docking location and target docking time window in the target docking point information;

[0065] The operation instruction mapping module is used to generate motion position instructions, motion posture instructions, and motion triggering timing instructions for end effectors or mounting mechanisms during space collaborative tasks, including grasping, delivery, loading, or detection operations.

[0066] Optionally, the cluster command distribution terminal includes an object matching module, a timing orchestration module, a link status monitoring module, and a distribution confirmation module;

[0067] The object matching module is used to match flight attitude control commands, ground motion control commands, and operation action control commands to the corresponding execution objects based on the device type, task role, and current location of the airborne and ground intelligent agents.

[0068] The timing orchestration module is used to orchestrate the triggering order of control commands for different execution objects according to the target connection time window;

[0069] The link status monitoring module is used to obtain the communication link quality and instruction reception status of different execution objects;

[0070] The distribution confirmation module is used to send distribution confirmation information to the collaborative connection and resolution terminal after confirming that the corresponding execution object has completed the instruction reception, and to trigger redistribution or re-resolution of the target connection point when there is an instruction reception abnormality.

[0071] A spatial cooperative mapping and multimodal command distribution method for air-to-ground heterogeneous multi-agent clusters is proposed, applied to the aforementioned spatial cooperative mapping and multimodal command distribution system for air-to-ground heterogeneous multi-agent clusters, combined with... Figure 2 As shown, the spatial cooperative mapping and multimodal command distribution method for air-ground heterogeneous multi-agent clusters includes:

[0072] S1, acquire airborne environmental perception information obtained by the aerial agent from the top-down perception of the target operation area, and generate airborne probability grid information based on the airborne environmental perception information to characterize the terrain distribution, obstacle distribution and passability probability distribution of the target operation area.

[0073] S2, acquire ground motion status information and ground contact feedback information generated by the ground intelligent agent during its movement or operation within the target work area;

[0074] S3 maps airborne probabilistic grid information, ground-based motion state information, and ground-based contact feedback information to the same generalized spatial coordinate system, generating heterogeneous collaborative spatial mapping information to characterize the spatial reachability of airborne intelligent agents, the travel stability of ground-based intelligent agents, and the spatial relationship of air-ground cooperation.

[0075] S4. Based on heterogeneous collaborative spatial mapping information, combined with the hovering constraints of the airborne intelligent agent, the passage constraints of the ground intelligent agent, and the air-ground collaborative task constraints, candidate docking areas that meet the collaboration conditions between the airborne intelligent agent and the ground intelligent agent are determined. Based on the spatial accessibility, passage stability, and collaboration execution cost corresponding to different locations in the candidate docking areas, the target docking point information is calculated.

[0076] S5 generates flight attitude control commands adapted to the execution of airborne intelligent agents and ground motion control commands adapted to the execution of ground intelligent agents based on the target docking point information. When the collaborative task involves end operations, it generates operation action control commands adapted to the end execution mechanism of the ground intelligent agent or the mounted mechanism of the airborne intelligent agent.

[0077] S6 performs object matching, timing arrangement, and distribution of flight attitude control commands, ground motion control commands, and operation action control commands according to the device type, task role, current location, and communication status of the airborne and ground intelligent agents, so that the airborne and ground intelligent agents can complete spatial collaborative tasks at the locations corresponding to the target docking point information.

[0078] Furthermore, in this embodiment, the aerial intelligent agent can be a drone, a compound-wing aircraft, a tethered flight platform, or a mobile flight device equipped with airborne sensors, while the ground intelligent agent can be a wheeled robot, a tracked robot, a legged robot, a humanoid robot, an automated guided vehicle, or a ground-based mobile device equipped with ground-based sensors. The aerial and ground intelligent agents do not require the same kinematic model and control interface. The aerial intelligent agent primarily undertakes tasks such as high-angle environmental perception, rapid target area coverage, aerial assisted positioning, or aerial payload delivery, while the ground intelligent agent primarily undertakes tasks such as near-ground passage, contact feedback acquisition, end-effector operations, target handling, or on-site execution. By unifying the top-down perception advantage of the aerial intelligent agent with the contact feedback advantage of the ground intelligent agent, the problem of inaccurate judgment of ground micro-traffic status when relying solely on airborne images can be avoided, as can the problems of insufficient environmental coverage and lack of an overall perspective in path planning when relying solely on the local perception of ground robots can also be avoided.

[0079] Furthermore, the target operation area can be a disaster relief area, a warehousing and transshipment area, an outdoor inspection area, a complex factory area, a mining area, an agricultural and forestry operation area, a port storage yard, or other areas requiring air-ground collaborative tasks. The target operation area may contain environmental factors affecting air-ground collaboration, such as road breaks, undulating slopes, obstacles, soft ground, water accumulation, gravel, steps, ditches, or temporary stockpiles. Airborne probabilistic grid information is used to describe the spatial distribution of the target operation area on a larger scale, while ground motion state information and ground contact feedback information are used to describe the attitude disturbances and contact resistance experienced by the ground-based agent when actually passing through the target area at a near-ground scale. Since the spatial scale, sampling frequency, and physical meaning of these two types of information differ, directly using them for independent control of the airborne agent and the ground-based agent respectively can easily lead to situations where the docking area selected by the airborne agent is visually usable but not actually stably reachable on the ground, or the ground-based agent can reach the area but the airborne agent cannot hover stably or complete the load handover. Therefore, in this embodiment, the air-based probabilistic grid information and the ground-based feedback information are uniformly mapped to the generalized spatial coordinate system through the heterogeneous spatial mapping terminal, and then the connection point is calculated by the collaborative connection calculation terminal, so that the air-ground collaborative task no longer depends on the local judgment of a single device.

[0080] Furthermore, the airborne environmental perception information can include visible light images, infrared images, depth images, laser point clouds, aerial agent positioning information, flight altitude information, and attitude angle information of the target operation area. The area image acquisition module can acquire image information of the target operation area according to a preset cruise trajectory, a temporary replanning trajectory, or a mission-focused area triggering strategy. For the same target operation area, the area image acquisition module can obtain coarse-grained global images and fine-grained local images at different flight altitudes; it can also obtain multi-view images of the target area at different top-down angles to reduce raster misjudgment caused by single-view occlusion. When acquiring image information, the airborne pose synchronization module synchronously records the aerial agent position information, flight altitude information, roll angle, pitch angle, and yaw angle corresponding to the image frame, and binds the image frame with the corresponding pose data based on the timestamp. Through this processing method, the probabilistic raster generation module can convert the ground feature areas in the image to the actual spatial location of the target operation area when generating airborne probabilistic raster information, reducing raster offset caused by changes in UAV attitude, flight altitude, or camera tilt.

[0081] Furthermore, when dividing the target operation area into grids, the probabilistic grid generation module can divide the area according to a fixed grid scale or adaptively based on obstacle density, the importance of the task's focus area, and the air-ground connection requirements. For areas with low obstacle density, larger grid cells can be used to reduce computation; for locations that may serve as connection areas, path turning areas, or end-operation areas, smaller grid cells can be used to improve spatial resolution. Each grid cell can record terrain category information, obstacle occupancy probability information, passability probability information, and image confidence information. Terrain category information can be used to distinguish between hard surfaces, grass, mud, gravel, slopes, steps, or water; obstacle occupancy probability information characterizes the likelihood that the grid cell is occupied by fixed obstacles, moving obstacles, or temporary debris; passability probability information characterizes the feasibility of a ground agent passing through the grid cell; and image confidence information characterizes the visibility and recognition reliability of the corresponding grid cell. This information is not stored in isolation but is spatially aligned with ground-based feedback information during subsequent heterogeneous spatial mapping.

[0082] Furthermore, the ground motion state information collected by the ground state feedback terminal can come from the odometer, inertial measurement unit, positioning module, joint encoder, wheel speed sensor, visual odometer, or lidar positioning unit of the ground intelligent agent. Ground contact feedback information can come from foot force sensors, wheel torque sensors, chassis suspension sensors, drive current sensors, contact pressure sensors, or feedback data from the actuators of the ground intelligent agent. When the ground intelligent agent is a wheeled robot, the contact force information can include changes in traction between the drive wheel and the ground, slippage, wheel load changes, or wheel speed deviation; when the ground intelligent agent is a legged robot or humanoid robot, the contact force information can include foot support force, landing impact force, support phase duration, and base posture recovery; when the ground intelligent agent is a tracked robot, the contact force information can include changes in track tension, drive resistance, and ground adhesion status. By uniformly abstracting the contact feedback information of different types of ground intelligent agents, the system can obtain the real near-ground passage status within the target operating area without limiting the specific ground platform structure.

[0083] Furthermore, the obstruction state identification module can identify the obstruction state at a corresponding location based on the deviation between the expected and actual travel speeds of the ground agent within adjacent time windows, the magnitude of changes in roll angle and yaw angle, the amplitude of base vibration, and the degree of abrupt changes in contact force. Obstruction states can include normal passage, slight obstruction, attitude disturbance, insufficient adhesion, difficult passage, and impassable states. For the same spatial location, if the airborne probability grid information indicates a high probability of passage, but the ground-based state feedback terminal continuously reports high contact force, large roll angle changes, or significant speed reduction, the heterogeneous spatial mapping terminal can reduce the ground passage stability corresponding to that location. Thus, the system can utilize actual ground contact feedback to correct for issues that are difficult to identify from an aerial perspective, such as soft ground, gravel disturbance, hidden ditches, or localized slopes, making subsequent docking point selection more consistent with actual execution conditions.

[0084] Furthermore, the generalized spatial coordinate system can use the global positioning coordinates, local map coordinates, task reference point coordinates, or temporarily constructed operational coordinates of the target operation area as a reference. The coordinate unification module can transform the flight coordinates of aerial agents, the grid coordinates corresponding to the airborne probabilistic grid information, and the ground motion coordinates of ground agents into the same coordinate expression. For aerial agents, the coordinate unification module can calculate the image pixel positions in the airborne environmental perception information into spatial grid positions based on the aerial agent's positioning information, flight altitude information, and attitude angle information; for ground agents, the coordinate unification module can bind the ground-based feedback information to the spatial position it traverses or adjacent grid cells based on the ground agent's positioning information, trajectory, and attitude changes. Through the above transformation, the airborne perception results and the ground-based feedback results can be compared, fused, and interleaved within the same target operation area.

[0085] Furthermore, when performing spatial alignment, the air-to-ground feature alignment module can directly align the airborne probabilistic raster information and ground-based feedback information within the same raster cell. If there is a deviation between the actual location traversed by the ground agent and the center location of the raster, the ground-based feedback information can be assigned to the nearest raster cell or distributed to multiple adjacent raster cells according to spatial proximity. If there is a time difference between the airborne sensing coverage area and the actual ground trajectory, the historical sensing results and current feedback results at the same location can be time-series aligned by combining the sampling timestamps. The air-to-ground feature alignment module can also mark information sources. When a region only has airborne probabilistic raster information but lacks ground-based feedback information, this region can be designated as a region to be verified. When a region has both airborne probabilistic raster information and ground-based feedback information, this region can be designated as a high-confidence fusion region. When there is a significant conflict between the airborne probabilistic raster information and ground-based feedback information in a region, this region can be designated as a verification region or a low-stability region.

[0086] Furthermore, the collaborative spatial field generation module can generate hovering suitability, ground traffic stability, and air-to-ground docking suitability for various spatial locations within the target operation area in a generalized spatial coordinate system. Hovering suitability is related to obstacle occupancy probability, overhead clearance conditions, hovering radius of the aerial agent, airflow disturbance risk, positioning accuracy, and communication link quality. Ground traffic stability is related to trafficability probability, travel obstruction information, contact force changes, roll state, yaw state, and the ground agent's mobility. Air-to-ground docking suitability is related to the reachability of the aerial agent's location, the stability of the ground agent's location, the arrival time difference between the two parties, the task execution space, the range of motion of the end-effector, and the safety distance. By forming heterogeneous collaborative spatial mapping information containing the above content, the system can convert information originally scattered across different agents, sensors, and control interfaces into a unified spatial description that can be used for docking calculations.

[0087] Furthermore, when selecting candidate docking areas, the candidate region filtering module can first exclude grid cells with an obstacle occupancy probability higher than a preset occupancy threshold, then exclude grid cells with an aerial hovering suitability lower than a preset hovering threshold, and further exclude grid cells with a ground passage stability lower than a preset passage threshold. For the remaining grid cells, the candidate region filtering module can form candidate docking areas based on the relative distance between the aerial agent and the ground agent, the task execution radius, the reachability range of the end-effector, and the safety isolation distance. This candidate docking area is not a single location point, but consists of a set of spatial locations or multiple local areas that meet the basic cooperation conditions. By forming candidate docking areas first and then calculating the cooperation cost, the computational burden on the cooperative docking solution terminal can be reduced, and invalid solutions can be avoided in obviously unusable areas.

[0088] Furthermore, the collaboration cost calculation module can calculate the hovering cost, ground travel cost, time synchronization cost, and collaboration execution cost separately. The hovering cost characterizes the control cost required for an aerial agent to reach a candidate location and maintain its mission attitude, and can be related to flight distance, remaining battery power, hovering stability, wind disturbance risk, obstacle distance, and positioning accuracy. The ground travel cost characterizes the travel cost required for a ground agent to reach a candidate location, and can be related to path length, slope change, contact disturbance, travel obstruction state, attitude recovery difficulty, and remaining energy. The time synchronization cost characterizes the degree of time matching between the aerial and ground agents reaching the candidate location, and can be related to the expected arrival time difference, waiting time, and mission time window. The collaboration execution cost characterizes the execution difficulty for both parties to complete delivery, grasping, detection, loading, or other spatial collaborative actions at the candidate location, and can be related to end-effector accessibility, docking attitude requirements, relative position error, safety distance, and action triggering sequence. By calculating these costs separately and using them to determine the target docking point, it is possible to avoid using only the shortest path, closest distance, or single accessibility as the basis for docking point selection.

[0089] Furthermore, when determining target docking point information, the target docking point determination module can select locations from the candidate docking area where the joint cooperation cost satisfies preset docking constraints, and further determine whether this location has a stable advantage relative to adjacent candidate locations. A stable advantage can be understood as the candidate location not being a randomly occurring low-cost point in adjacent grids or adjacent local areas, but rather a location with continuous availability and low disturbance risk within the spatial neighborhood. By introducing a comparison process with adjacent candidate locations, the system can avoid misjudging isolated locally optimal locations as target docking points. For example, a location may appear as an open area in the airborne image, but its adjacent locations all exhibit high ground contact disturbances or large tilt changes; therefore, even if the single-point cost is low, this location may not be suitable as a target docking point. Target docking point information may include the target docking location, target docking time window, hovering altitude of the airborne agent target, hovering attitude of the airborne agent target, arrival attitude of the ground agent target, direction of action of the end effector, and allowable error range.

[0090] Furthermore, when generating flight attitude control commands, the multimodal command generation terminal can convert target docking point information into executable flight path sequence, velocity change sequence, hovering altitude command, attitude angle adjustment command, and hovering hold command for the aerial agent. For tasks requiring payload delivery or reception, flight attitude control commands can also include payload release altitude, payload orientation, hovering duration, and safe evacuation direction. When generating ground motion control commands, the multimodal command generation terminal can convert target docking point information into executable path point sequence, velocity control command, base attitude stabilization command, obstacle avoidance adjustment command, and position hold command for the ground agent. For legged robots or humanoid robots, ground motion control commands can also include landing area constraints, gait switching commands, and base attitude compensation commands. For wheeled or tracked robots, ground motion control commands can also include turning radius control, drive speed adjustment, and anti-slip control commands.

[0091] Furthermore, when space collaboration tasks involve grasping, delivery, loading, or detection operations, the operation command mapping module can generate motion position commands, motion attitude commands, and motion triggering sequence commands for the end effector or mounting mechanism based on the target docking attitude requirements and target docking time window in the target docking point information. The end effector can be a robotic arm, gripper, suction mechanism, detection probe, load carrier, delivery mechanism, or other mechanism capable of performing space operations. The operation command mapping module can convert the relative pose relationship between the airborne agent and the ground agent into local motion coordinates of the end effector and set the timing of motion waiting, motion triggering, and motion withdrawal based on the time difference between their arrival at the target docking point. Through this processing method, the system outputs commands that are not merely "reach a certain position" movement commands, but a complete multimodal command combination covering position arrival, attitude maintenance, end effector actions, and collaborative triggering for both airborne and ground agents.

[0092] Furthermore, when performing object matching, the cluster command distribution terminal can distribute different types of control commands to corresponding execution objects based on the equipment type, task role, current location, remaining energy, payload status, communication status, and execution capabilities of airborne and ground-based intelligent agents. For scenarios involving multiple airborne intelligent agents, the cluster command distribution terminal can assign the primary airborne intelligent agent to the target docking point and the auxiliary airborne intelligent agents to environmental monitoring or communication relay locations. For scenarios involving multiple ground-based intelligent agents, the cluster command distribution terminal can assign ground-based intelligent agents with end-effector capabilities as docking execution objects and ground-based intelligent agents with transport or path clearing capabilities as auxiliary execution objects. By combining task role and equipment capabilities for distribution, the object matching module can avoid sending end-effector commands to ground-based intelligent agents without end-effectors or sending hovering attitude commands to airborne platforms without corresponding flight capabilities.

[0093] Furthermore, the timing orchestration module can orchestrate the triggering order of flight attitude control commands, ground motion control commands, and operational action control commands based on the target docking time window and the expected arrival time of each execution object. For example, a trajectory adjustment command can be sent to the airborne agent first, causing it to enter a pre-hovering area near the target docking point; simultaneously, a travel path command can be sent to the ground agent, causing it to approach the target docking point; once both have entered the allowable error range corresponding to the target docking point, hovering attitude control commands, base attitude stabilization commands, and operational action control commands are triggered; after completing delivery, grasping, loading, or detection operations, a withdrawal command or a next task connection command is triggered. Through timing orchestration, energy consumption losses caused by prolonged ineffective hovering of the airborne agent can be reduced, as can the risk of attitude disturbances caused by the ground agent arriving early in complex terrain areas and waiting for extended periods can also be reduced.

[0094] Furthermore, the link status monitoring module can acquire the communication link quality, command reception status, acknowledgment feedback status, and execution feedback status of each execution object. Communication link quality can include signal strength, link latency, packet loss rate, or number of retransmissions. The command reception status indicates whether the corresponding execution object has successfully received the control command; the acknowledgment feedback status indicates whether the corresponding execution object has confirmed that the control command can be executed; and the execution feedback status indicates whether the corresponding execution object has started execution, interrupted execution, or completed execution. After confirming that the corresponding execution object has completed command reception, the distribution confirmation module can send distribution confirmation information back to the collaborative connection and resolution terminal. When there is an command reception anomaly, link latency exceeding the limit, a change in the execution object's status, or a task time window failure, the distribution confirmation module can trigger redistribution or re-resolution of the target connection point. By introducing a feedback loop in the command distribution phase, the resolution result of the connection point can be kept consistent with the actual communication and execution status, avoiding situations where air-to-ground agents arrive at different locations or perform collaborative actions at inconsistent times due to unilateral command distribution.

[0095] Furthermore, in a specific operational process, the system first uses an airborne environmental perception terminal to control one or more aerial agents to perform a top-down scan of the target operational area, acquiring image information of the target operational area as well as the corresponding aerial agent position information, flight altitude information, and attitude angle information. An airborne pose synchronization module binds the image information with the aerial agent's pose data, and a probabilistic grid generation module generates airborne probabilistic grid information based on the bound data. Simultaneously, a ground-based state feedback terminal acquires the ground agent's current position, orientation angle, travel speed, roll angle change information, yaw angle change information, base vibration information, contact force information, and travel drag information within the target operational area. A heterogeneous space mapping terminal unifies the aforementioned airborne and ground-based information into a generalized spatial coordinate system and generates heterogeneous collaborative space mapping information. A collaborative docking solution terminal filters candidate docking areas from the heterogeneous collaborative space mapping information and determines the target docking point information based on airborne hovering cost, ground passage cost, time synchronization cost, and collaborative execution cost. The multimodal instruction generation terminal converts the target docking point information into control instructions that are adapted to different intelligent agents. The cluster instruction distribution terminal distributes instructions to airborne and ground intelligent agents according to the object matching and timing arrangement results, enabling airborne and ground intelligent agents to complete spatial collaborative tasks at the locations corresponding to the target docking point.

[0096] Furthermore, in another specific operational process, when a ground agent reports significant travel obstruction in a certain area during its journey, the heterogeneous spatial mapping terminal can update the ground traffic stability of that area and send the updated heterogeneous cooperative spatial mapping information to the cooperative docking resolution terminal. The cooperative docking resolution terminal then re-selects candidate docking areas based on the updated heterogeneous cooperative spatial mapping information. If the ground traffic stability of the area where the original target docking point is located is lower than a preset traffic threshold, or the air-to-ground arrival time difference exceeds the target docking time window, the target docking point determination module can re-determine the target docking point information. The multimodal instruction generation terminal updates the flight attitude control instructions, ground motion control instructions, and operational action control instructions based on the re-determined target docking point information, and the cluster instruction distribution terminal distributes the updated instructions to the corresponding execution objects. Through this process, the system can dynamically correct docking points and cooperative instructions when there is a discrepancy between actual ground feedback and air-based prediction, thereby improving the continuity and reliability of air-to-ground cooperative missions in complex environments.

[0097] Furthermore, in another specific operational scenario, when the link status monitoring module detects a decline in the communication link quality of an airborne or ground-based agent, the cluster command distribution terminal can handle the situation based on the task role of the executing object and the current command execution stage. If the communication link quality decline occurs before the connection task begins, the corresponding command distribution can be paused and the object can be re-matched. If the communication link quality decline occurs when the airborne or ground-based agent is already approaching the target connection point, the system can determine whether to continue executing the current connection task based on the distributed commands and the target connection time window. If the communication link quality decline causes command confirmation timeouts or missing execution feedback, the distribution confirmation module can trigger the target connection point to recalculate or select a backup executing object. These processing methods enable the system to incorporate communication status into the spatial collaboration process, rather than mechanically issuing control commands after the connection point is determined, thereby improving the collaboration security of multi-agent clusters in unstable communication environments.

[0098] Furthermore, the heterogeneous collaborative spatial mapping information in this embodiment can be continuously updated as an intermediate data structure for air-ground collaborative tasks. The airborne environmental perception terminal can reacquire image information of the target operation area according to a preset period, task triggering conditions, or abnormal feedback triggering conditions; the ground-based state feedback terminal can output ground motion state information and ground contact feedback information in real time or periodically during the continuous movement of the ground agent; the heterogeneous spatial mapping terminal can superimpose or replace the newly acquired information to the corresponding position in the generalized spatial coordinate system, so that the heterogeneous collaborative spatial mapping information is continuously updated with the task execution process. The collaborative docking solution terminal can re-evaluate the candidate docking area and target docking point information when the heterogeneous collaborative spatial mapping information changes, the state of the execution object changes, the communication state changes, or the task constraints change. Thus, the system has the ability to continuously collaborate in dynamic environments, rather than only performing one-time path planning or one-time docking point selection before the task begins.

[0099] Furthermore, the spatial collaboration tasks in this embodiment can include airborne agents delivering supplies to ground agents, ground agents transferring samples to airborne agents, airborne agents assisting ground agents in locating targets, ground agents performing close-range detection under the guidance of airborne agents, and airborne and ground agents jointly completing tasks such as area inspection, rescue delivery, dangerous area exploration, or equipment maintenance. For delivery tasks, the target docking point information can primarily constrain the hovering height of the airborne agent, the orientation of the mounting mechanism, and the receiving posture of the ground agent; for grasping tasks, the target docking point information can primarily constrain the range of motion of the end effector, the position of the target object, and the relative posture of both parties; for detection tasks, the target docking point information can primarily constrain the detection distance, detection angle, equipment stabilization time, and safe isolation distance. By converting different task types into a unified docking point calculation and multimodal instruction generation process, the system can adapt to various heterogeneous air-ground collaboration scenarios.

[0100] Furthermore, the system in this embodiment can be deployed in a cluster control server, edge computing device, ground station, vehicle-mounted computing platform, or air-ground collaborative mission control platform. The airborne environmental perception terminal, ground-based status feedback terminal, heterogeneous space mapping terminal, collaborative connection and calculation terminal, multimodal instruction generation terminal, and cluster instruction distribution terminal can be implemented as software modules, hardware modules, or a combination of software and hardware modules. Data interaction between terminals can occur via wired communication, wireless communication, local area network, mobile communication network, ad hoc network, or dedicated communication link. For application scenarios with sufficient computing resources, the heterogeneous space mapping terminal and collaborative connection and calculation terminal can be centrally deployed at a ground station or edge server; for application scenarios with limited communication or requiring rapid response, some mapping calculations, candidate region screening, or instruction generation processes can also be deployed locally on the airborne agent or ground agent to reduce communication latency and the load on the central node.

[0101] Furthermore, this embodiment acquires global or local airborne probabilistic grid information of the target operation area through an airborne environmental perception terminal, acquires the motion state and contact feedback of ground-based intelligent agents during actual passage through a ground-based state feedback terminal, maps data from different sources, scales, and physical meanings to a generalized spatial coordinate system through a heterogeneous spatial mapping terminal, determines target docking point information from aspects such as spatial accessibility, passage stability, and collaborative execution cost through a collaborative docking solution terminal, and completes instruction generation and distribution execution adapted to different intelligent agents through a multimodal instruction generation terminal and a cluster instruction distribution terminal. Compared with schemes that only centrally schedule homogeneous UAVs or homogeneous ground robots, this embodiment can handle the differences between airborne flight platforms and ground mobile platforms in terms of motion mode, perception dimension, actuator, and communication status, enabling airborne and ground heterogeneous multi-agent clusters to complete tasks with spatial intersection and action coordination relationships within complex target operation areas.

[0102] Example 2: Building upon Example 1, this example further provides an algorithm implementation for spatial collaborative mapping, target connection point calculation, and multimodal command distribution in a heterogeneous air-to-ground multi-agent cluster. This example does not alter the connection relationships between terminals and modules in Example 1. Instead, it further expands the internal processing logic of the heterogeneous spatial mapping terminal, collaborative connection calculation terminal, multimodal command generation terminal, and cluster command distribution terminal, enabling the airborne probabilistic grid information, ground motion state information, and ground contact feedback information to form a computable, updatable, and distributable basis for collaborative control.

[0103] Combination Figure 3 and Figure 4 As shown, in this embodiment, the target work area is divided into multiple spatial grid cells, any spatial grid cell being denoted as... ,in This represents the grid number within the target operational area. The set of aerial agents is denoted as... The set of ground-based intelligent agents is denoted as Aerial intelligent agents At any moment For grid cells The resulting space-based observation information is denoted as Ground-based intelligent agents At any moment For grid cells The resulting foundation feedback information is denoted as The heterogeneous spatial mapping terminal transforms data from different sources into a generalized spatial coordinate system, and then forms a heterogeneous cooperative state vector for each grid cell:

[0104] ;

[0105] in, Represents grid cells At any moment Heterogeneous cooperative state vector; Indicates the probability of obstacle occupancy; Indicates the probability that the route is passable; Indicates the confidence level of the empty-base image; This represents the amount of contact force disturbance generated by a ground-based intelligent agent within the grid cell or its neighborhood. Indicates the amount of roll disturbance; Indicates the yaw disturbance amount; Indicates the amount of speed reduction; This indicates the communication link quality of objects near the grid cell. The communication link quality can be determined based on signal strength, link delay, packet loss rate, retransmission count, or command confirmation feedback status. Through the above state vector, information originally scattered across airborne sensing, ground motion feedback, and communication status is uniformly expressed as calculable parameters at the same spatial location, facilitating subsequent connection point selection and command generation.

[0106] Furthermore, when generating heterogeneous collaborative spatial mapping information, the heterogeneous spatial mapping terminal does not simply superimpose the space-based probabilistic grid information with the ground-based feedback information. Instead, it first calculates the space-ground fusion reliability of each grid cell. The space-ground fusion reliability characterizes the degree of consistency and usability between the space-based observation results and the ground-based feedback results within that grid cell. The calculation method is as follows:

[0107] ;

[0108] in, Represents grid cells At any moment Reliability of air-to-ground integration; This represents the passability probability given in the empty-base probability raster information; This represents the ground traffic status value obtained based on ground motion state information and ground contact feedback information. The range of the ground traffic status value can be normalized to 0 to 1. The larger the value, the more stable the actual ground traffic status. This represents a stability constant for the difference in the probability of passage, used to prevent the denominator from becoming too small; Indicates the confidence level of the empty-base image; This indicates the amount of ground contact force disturbance; A reference threshold representing ground contact force disturbance; This represents the sensitivity coefficient of reliability degradation caused by contact force disturbance. The higher the value, the more consistent the airborne observation and the ground-based feedback are, and the smaller the contact disturbance at that location; The lower the value, the more likely the location is to have misjudgment of the airborne image, hidden ground obstruction, or local contact anomalies, making it unsuitable for direct use as a high-confidence connection area.

[0109] Among them, ground traffic status value Ground-based intelligent agents can pass through grid cells The velocity decay, attitude disturbance, and contact force change in the region or its neighborhood are jointly determined, and the calculation method is as follows:

[0110] ;

[0111] in, Represents grid cells The ground traffic status value can be normalized to 0 to 1. This indicates the amount of speed reduction of the ground-based intelligent agent at that location; This represents the reference value for permissible speed reduction under normal traffic conditions; Indicates the amount of roll disturbance; This indicates the allowable roll disturbance reference value; Indicates the yaw disturbance amount; This indicates the permissible yaw disturbance reference value; Indicates the amount of contact force disturbance; This indicates the allowable contact force disturbance reference value; , , , These represent the stability constants of the corresponding parameters. The formula employs an exponential decay form, ensuring that a significant increase in any ground disturbance parameter leads to a rapid decrease in the ground traffic state value, thus avoiding the problem of relying solely on the average state to mask localized strong disturbances.

[0112] Furthermore, the collaborative spatial field generation module calculates the aerial hovering suitability, ground traffic stability, and air-to-ground connection adaptability of grid cells based on air-to-ground fusion reliability, aerial hovering conditions, and ground traffic status. Aerial hovering suitability describes the degree to which an aerial agent is suitable for hovering above a target location and performing collaborative actions; the calculation method is as follows:

[0113] ;

[0114] in, Represents grid cells The suitability for hovering in the air; Indicates the probability of obstacle occupancy; Indicates the confidence level of the empty-base image; This indicates the safe distance between the grid cell and any obstacle above or around it; Indicates reference values ​​for safe clearance from obstacles; This represents the estimated value of airflow disturbance or hovering disturbance at that location; This indicates the allowable airflow disturbance reference value; This represents an estimated energy consumption required for an aerial agent to reach this location and remain hovering. This indicates a reference value for hovering energy consumption; , , ... Used to demonstrate the nonlinear suppression effect on hovering suitability when the safety clearance is insufficient, when The smaller the size, the stronger the suppression of hover suitability, so that even if the image confidence is high, the location near the obstacle will not be easily selected as the connection area.

[0115] Ground traffic stability describes the stability of a ground agent when it arrives at and stays in a corresponding grid cell. The calculation method is as follows:

[0116] ;

[0117] in, Represents grid cells Ground traffic stability; Indicates the ground traffic status value; Indicates the reliability of air-to-ground integration; This indicates the slope or elevation change index corresponding to the grid cell. The slope or elevation change index can be determined by airborne elevation estimation, ground attitude change, local map elevation difference, or multi-frame point cloud fitting results. This indicates the allowable slope reference value; This indicates the degree of ground roughness or local undulation. Ground roughness or local undulation can be determined by point cloud undulation, ground vibration feedback, contact force fluctuation, or changes in load on wheel ends or feet. This indicates the allowable roughness reference value; , These represent the stability constants of the corresponding parameters. Through this calculation method, ground traffic stability is determined not only by the passability probability visible from the air, but also by the attitude perturbations, contact perturbations, and terrain undulations of the ground agent during actual passage.

[0118] The air-to-ground docking adaptability is used to describe the overall adaptability of a grid cell as a collaborative location between an airborne agent and a ground-based agent. The calculation method is as follows:

[0119] ;

[0120] in, Represents grid cells The adaptability of open-air connections; Indicates the suitability for hovering in the air; Indicates ground traffic stability; This indicates the time difference between the estimated arrival times of the aerial agent and the ground agent at this location. This indicates the allowable time difference reference value; This represents the estimated relative spatial distance error when both parties perform cooperative actions at this location; This indicates the allowable relative spatial distance error reference value; This represents the estimated relative posture deviation when both parties perform cooperative actions; This indicates the allowable relative attitude deviation reference value; , , These represent the stability constants of the corresponding parameters. This formula couples three factors: air suitability, ground stability, and spatiotemporal synchronization. A significant deficiency in any of these factors will reduce the air-to-ground connection suitability, thus preventing the connection point from being only accessible on one side and unable to complete collaborative actions.

[0121] Furthermore, the candidate area screening module performs preliminary screening of the target work area based on the adaptability of air-ground connections. For any given grid cell... If a device meets the following conditions, it will be included in the candidate connection area:

[0122] ;

[0123] in, This indicates the lower limit of hovering suitability. This indicates the lower limit of ground traffic stability; This indicates the lower limit of the adaptability of air-ground connection; This represents the upper limit of the probability of obstacle occupancy. Grid cells that meet the above conditions are included in the candidate connection area set. Multiple spatially adjacent candidate locations together constitute a candidate connection area. Through this screening method, the collaborative connection calculation terminal can eliminate locations that are obviously unsuitable as air-to-ground junctions before entering the joint collaborative cost calculation, reducing the amount of invalid calculations and reducing the risk of connection points falling on the edge of obstacles, ground disturbance areas, or areas with weak communication coverage.

[0124] Furthermore, the cooperation cost calculation module calculates the joint cooperation cost for each candidate location. The joint cooperation cost is not a simple linear sum of the various indicators, but rather reflects the interplay between airborne energy consumption, ground energy consumption, contact disturbance, time synchronization, and the difficulty of action execution through multiplicative coupling and nonlinear penalties. The calculation method is as follows:

[0125] ;

[0126] in, Indicates candidate position At any moment The cost of joint collaboration; This represents the flight energy consumption and hovering energy consumption required for an aerial intelligent agent to reach and hover at the candidate location; This indicates a reference value for in-flight energy consumption; This represents the energy consumption required for a ground-based intelligent agent to reach and remain at the candidate location. This represents a reference value for ground energy consumption. This indicates the complexity of the action to complete the end effector or mounting operation at the candidate location. The action complexity can be determined by one or more of the following: the range of motion of the end effector, the attitude adjustment angle, the duration of the grasping or delivery action, the number of obstacle avoidance steps in the action path, and the release accuracy requirements of the mounting mechanism. This represents a reference value for the complexity of the action. Indicates the compatibility between open and ground transportation; Indicates the quality of the communication link; Indicates the time difference of arrival at the open ground; This indicates the amount of ground contact force disturbance; , , These represent the nonlinear amplification exponents for airborne energy consumption, ground energy consumption, and motion complexity, respectively. , , , , , , These represent the stability constants of the corresponding parameters. Through the aforementioned joint cooperation cost, even candidate locations with good geometric distance conditions will be significantly penalized when the time difference is too large, the contact disturbance is too strong, the communication quality is low, or the action complexity is high, thus making the target docking point more in line with the actual cooperative execution requirements.

[0127] Furthermore, to avoid the target connection point falling into isolated, low-cost locations, the target connection point determination module calculates a neighborhood stability advantage value for each candidate location. Let the candidate locations be... The spatial neighborhood is The spatial neighborhood can include several grid cells adjacent to the candidate location, or it can include multiple grid cells within a preset radius centered on the candidate location. The neighborhood stability advantage value is calculated as follows:

[0128] ;

[0129] in, Indicates candidate position The neighborhood stability dominance value; Indicates candidate position The cost of joint collaboration; Indicates candidate position spatial neighborhood Any neighboring raster cell within; Represents neighborhood grid cells The cost of joint collaboration; Indicates the number of grid cells in the neighborhood; Represents neighborhood grid cells The adaptability of open-air connections; The variance represents the fit of open space access within the neighborhood; This represents the stability constant. The formula determines, on the one hand, whether the candidate location itself has a low joint cooperation cost, and on the other hand, whether the connectivity fitness within its neighborhood is stable. When the candidate location itself has a low cost and the neighborhood fitness fluctuation is small, The value is relatively large; when a certain location has a low cost at a single point, but the surrounding grid adaptability fluctuates greatly, its neighborhood stability advantage value will be suppressed and it will not be selected as the target connection point.

[0130] Furthermore, the target connection point determination module can jointly determine the target connection point based on the joint cooperation cost and the neighborhood stability advantage value. Target connection point The determination method is as follows:

[0131] ;

[0132] in, This represents the grid cell corresponding to the target docking point; Indicates time The set of candidate connection areas; Indicates candidate position The cost of joint collaboration; Indicates candidate position The neighborhood stability dominance value; This represents a stability constant used to prevent the denominator from becoming too small. Through this determination method, the target docking point not only needs to have a low joint cooperation cost, but also good stability within its spatial neighborhood, reducing the likelihood of frequent jumps or falling into locally accidental low-cost regions. The target docking point information is provided by... The corresponding spatial location, target rendezvous time window, hovering altitude of the aerial agent target, hovering attitude of the aerial agent target, arrival attitude of the ground agent target, direction of action of the end-effector, and allowable error range are all considered together.

[0133] Furthermore, after the target rendezvous point is determined, the multimodal command generation terminal performs inverse mapping of airborne control commands and ground control commands based on the target rendezvous point information. For the airborne agent, the target rendezvous point information is converted into target hovering position, target hovering altitude, target attitude angle, arrival time, and hovering duration. To avoid directly interpreting spatial deviation as the underlying attitude angle output, this embodiment first calculates the spatial deviation correction vector of the airborne agent in its own body coordinate system, and then the flight control interface of the airborne agent converts this spatial deviation correction vector into attitude angle correction commands, speed control commands, or hovering duration commands. The spatial deviation correction vector of the airborne agent can be calculated as follows:

[0134] ;

[0135] in, Indicates that the aerial intelligent agent is at any time Spatial deviation correction vector in its own body coordinate system; This represents the rotation matrix corresponding to the current attitude of the aerial agent; Represent its inverse matrix; Indicates the spatial location corresponding to the target connection point; Indicates the current spatial location of the aerial intelligent agent; This indicates the distance between the aerial agent and the target docking point; Indicates the reference distance for correcting spatial deviations in the air; This represents the stability constant. This formula is used to transform the spatial deviation of the target docking point relative to the current attitude of the airborne agent into the body coordinate system, and controls the correction magnitude of the long-range and short-range phases through a nonlinear attenuation term, so that the airborne agent can more smoothly enter the hovering attitude when approaching the target docking point.

[0136] For ground-based intelligent agents, the multimodal command generation terminal generates ground motion control commands based on the target docking point location, ground traffic stability, and travel obstruction information. The velocity correction of the ground-based intelligent agent can be calculated as follows:

[0137] ;

[0138] in, Indicates the ground-based intelligent agent at time... The speed correction amount can be used as an adjustment input for the reference speed of the ground agent, rather than directly replacing the underlying speed closed-loop control of the ground agent; This represents the reference travel speed of the ground-based intelligent agent; This indicates the path distance from the current location of the ground-based intelligent agent to the target docking point; Indicates the reference distance for ground speed adjustment; The stability constant representing the ground path distance term; This indicates the ground traffic stability of the grid cell corresponding to the target connection point; This represents the contact force disturbance amount of the grid cell corresponding to the target connection point; This represents the tilt disturbance of the grid cell corresponding to the target docking point; This indicates the allowable contact force disturbance reference value; This indicates the allowable roll disturbance reference value; , These represent the stability constants of the corresponding parameters. This formula enables the ground-based intelligent agent to automatically reduce or increase its speed based on the ground stability when approaching the target docking point. In particular, it can actively slow down its travel speed when the contact disturbance is strong or the risk of tilting is high, thereby improving attitude stability during the docking phase.

[0139] Furthermore, when space collaboration tasks involve actions of end-effectors or mounted mechanisms, the operation command mapping module needs to calculate the triggering sequence of the end-effector actions. Let the estimated time for the aerial agent to reach the target docking point be... The estimated time for the ground-based intelligent agent to reach the target docking point is The preparation time required for the end effector to perform the action is The duration of the action is The trigger time for the end effect can be determined as follows:

[0140] ;

[0141] in, Indicates the trigger time of the end effect; This indicates the estimated time for the aerial agent to arrive at the target docking point; This indicates the estimated time for the ground-based intelligent agent to reach the target docking point; Indicates the preparation time for the final action; This indicates the allowable time difference reference value; This represents the time synchronization sensitivity coefficient. When the expected arrival times of both parties are close, the action triggering time is primarily determined by the arrival and preparation time of the later-arriving party. When the expected arrival times differ significantly, the triggering time is subject to additional delay correction, thus preventing premature triggering of the terminal action before one party has reached a stable connection state. Action hold time. Used to determine the end effector in The duration during which the end-effector continues to grasp, deliver, detect, or mount can be determined as follows:

[0142] ;

[0143] in, Indicates the end time of the terminal action; Indicates the trigger time of the end effect; This indicates the duration of the end-effector action. By simultaneously determining the trigger time and end-effector action end time, the operation command mapping module can form a complete action sequence including preparation, triggering, holding, and termination, reducing collaboration failures caused by premature or insufficient action holding during the docking process between air and ground agents.

[0144] Furthermore, before distributing instructions, the cluster instruction distribution terminal can calculate the instruction distribution priority for each execution object. The instruction distribution priority is used to determine which devices in a multi-agent cluster should receive control instructions first, and which devices should serve as backup or auxiliary execution objects. It can be an aerial intelligent agent, a ground-based intelligent agent, a communication relay node, or an execution platform carrying an end effector. For the execution object... The priority of instruction dispatch can be calculated as follows:

[0145] ;

[0146] in, Indicates the execution object At any moment Instruction dispatch priority; Indicates the degree of matching between the execution object and the current task role; Indicates the degree of matching between the capabilities of the execution objects; Indicates the quality of the communication link of the executing object; Indicates the remaining energy of the executing object; This represents the minimum energy reference value required to perform the current task; This represents the inhibition coefficient of priority due to insufficient energy. This indicates the distance between the executing object and the target connection point or task area; A reference value indicating the distance from the execution object to the target connection point or task area; The stability constant representing the task distance term; Indicates the current communication link latency of the executing object; Indicates the allowable communication delay reference value; This represents the stability constant of the communication link delay term. Using this priority calculation method, the system can comprehensively consider task roles, equipment capabilities, communication status, remaining energy, and spatial distance, avoiding the selection of execution targets solely based on proximity or strongest communication.

[0147] Furthermore, the cluster command distribution terminal can perform object matching based on command distribution priority and target connection point information. For airborne commands, airborne agents with corresponding hovering capabilities, payload capacity, and communication link quality are prioritized; for ground movement commands, ground agents with arrival capabilities, passage capabilities, and attitude stabilization capabilities are prioritized; for operational action control commands, execution objects with end effectors or mounted mechanisms are prioritized. When the highest priority execution object experiences reception anomalies, link anomalies, or insufficient energy during the command confirmation phase, the cluster command distribution terminal can select the second highest priority execution object as a backup execution object and feed back the object change information to the cooperative connection resolution terminal to determine whether the target connection point is still applicable.

[0148] Furthermore, in a specific processing flow of this embodiment, the heterogeneous spatial mapping terminal first unifies the spatial coordinates of the airborne probabilistic grid information and the ground-based state feedback information to form a heterogeneous cooperative state vector for each grid cell; then it calculates the ground traffic status value and air-to-ground fusion reliability, and generates airborne hovering suitability, ground traffic stability, and air-to-ground connection suitability based on these; the cooperative connection calculation terminal filters candidate connection areas according to the above indicators, and calculates the joint cooperative cost and neighborhood stability advantage value for the candidate locations; the target connection point determination module determines the target connection point based on the joint cooperative cost and neighborhood stability advantage value; the multimodal instruction generation terminal converts the target connection point into spatial deviation correction instructions for the airborne intelligent agent, speed and attitude control instructions for the ground intelligent agent, and end-effector action triggering instructions; the cluster instruction distribution terminal completes object matching, timing arrangement, and distribution confirmation according to the instruction distribution priority of the execution object. The above process forms a closed-loop processing chain for the air-to-ground collaboration process from environment mapping, connection calculation, instruction generation to instruction distribution.

[0149] Furthermore, in a dynamic environment, if the travel obstruction information reported by the ground-based state feedback terminal changes abruptly, or if the communication link quality reported by the link status monitoring module is lower than a preset communication threshold, the heterogeneous spatial mapping terminal updates the heterogeneous cooperative state vector of the corresponding grid cell and recalculates the air-to-ground fusion reliability, ground traffic stability, and air-to-ground connection adaptability. The cooperative connection resolution terminal recalculates the joint cooperative cost of candidate locations based on the updated indicators. When the joint cooperative cost of the original target connection point exceeds the preset cost upper limit, or the neighborhood stability advantage value is lower than the preset stability advantage threshold, the target connection point determination module re-determines the target connection point and outputs the updated target connection point information to the multimodal command generation terminal. Through this dynamic update mechanism, the system can correct the connection point and control commands based on actual ground traffic feedback, air hovering conditions, and changes in communication status.

[0150] Furthermore, the reference values ​​and stability constants in this embodiment can be pre-configured according to different operational scenarios, equipment types, and task requirements, or updated based on historical task data. For disaster relief scenarios, the constraints on ground traffic stability and communication link quality in target connection point calculation can be increased; for warehousing or factory collaboration scenarios, the constraints on time synchronization cost and action execution cost can be increased; for outdoor inspection or complex terrain scenarios, the impact of contact force disturbance, lateral tilt disturbance, and slope change on ground traffic stability can be increased. By configuring different reference values ​​and thresholds for different application scenarios, the same set of air-ground heterogeneous collaborative algorithms can be adapted to various air-ground collaborative tasks without changing the system structure in Embodiment 1.

[0151] Furthermore, in this embodiment, the probability parameters in the aforementioned formulas can range from 0 to 1, with larger values ​​indicating a higher probability of the corresponding state occurring or a stronger corresponding capability; the suitability, stability, adaptability, and reliability parameters can be normalized to 0 to 1, with larger values ​​indicating a more suitable grid cell for air-to-ground collaboration; the larger the cost parameter value, the higher the difficulty of collaborative execution at the corresponding candidate location; energy consumption parameters can be estimated from the remaining battery power of the execution object, task execution distance, hovering time, drive load, or feedback from the underlying controller; disturbance parameters can be obtained from inertial measurement units, contact force sensors, drive current, foot force feedback, wheel end torque, or chassis vibration data; communication link quality can be determined from signal strength, link delay, packet loss rate, retransmission count, or command confirmation feedback status. Each reference value, threshold, and stability constant can be pre-configured according to the type of airborne intelligent agent, the type of ground intelligent agent, the complexity of the target operation area, and the type of spatial collaboration task, and the symbols should maintain consistent meaning within the same embodiment.

[0152] Furthermore, this embodiment addresses the potential inconsistency between airborne visual judgment and ground-based actual contact feedback through air-to-ground fusion reliability calculation. It resolves the coupling judgment problem between airborne reachability, ground reachability, and collaborative executability through hierarchical calculation of airborne hovering suitability, ground passage stability, and air-to-ground docking suitability. It addresses the problem of isolated low-cost misjudgments in target docking point selection by combining collaborative cost and neighborhood stability advantage values. It converts target docking points into multimodal instructions executable by different execution objects through spatial deviation correction, velocity correction, and action triggering time calculation. Finally, it ensures that instructions are matched to appropriate airborne and ground agents through instruction distribution priority calculation. Therefore, this embodiment can improve the collaborative mapping accuracy, docking point selection stability, and instruction distribution reliability of air-to-ground heterogeneous multi-agent clusters under complex terrain, dynamic tasks, and unstable communication conditions.

[0153] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of protection of the present invention. Furthermore, the elements therein can be updated as technology develops.

Claims

1. A spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters, characterized in that, It includes airborne environmental sensing terminals, ground-based state feedback terminals, heterogeneous space mapping terminals, collaborative connection and calculation terminals, multimodal command generation terminals, and cluster command distribution terminals; The airborne environmental perception terminal is used to acquire airborne environmental perception information obtained by the aerial intelligent agent through overhead perception of the target operation area, and to generate airborne probability grid information based on the airborne environmental perception information to characterize the terrain distribution, obstacle distribution and passability probability distribution of the target operation area. The ground status feedback terminal is used to acquire ground motion status information and ground contact feedback information generated by the ground intelligent agent during its movement or operation within the target work area. The ground motion state information includes at least the pose information, tilt state information, and yaw state information of the ground agent; the ground contact feedback information includes at least the contact force information and travel resistance information generated during the contact between the ground agent and the ground. The heterogeneous spatial mapping terminal is connected to the airborne environment sensing terminal and the ground-based state feedback terminal respectively, and is used to map the airborne probability grid information, the ground-based motion state information and the ground-based contact feedback information to the same generalized spatial coordinate system to generate heterogeneous cooperative spatial mapping information for characterizing the spatial reachability of airborne intelligent agents, the travel stability of ground intelligent agents and the positional relationship of air-ground cooperation. The collaborative docking solution terminal is connected to the heterogeneous space mapping terminal and is used to determine candidate docking areas between the air agent and the ground agent that meet the cooperation conditions based on the heterogeneous collaborative space mapping information, combined with the hovering constraints of the air agent, the passage constraints of the ground agent, and the air-ground collaborative task constraints. The target docking point information is calculated based on the spatial reachability, passage stability, and collaborative execution cost corresponding to different locations in the candidate docking areas. The multimodal instruction generation terminal is connected to the collaborative connection and resolution terminal, and is used to generate flight attitude control instructions adapted to the execution of the airborne intelligent agent and ground motion control instructions adapted to the execution of the ground intelligent agent according to the target connection point information. When the collaborative task involves end operation, it generates operation action control instructions adapted to the end execution mechanism of the ground intelligent agent or the mounted mechanism of the airborne intelligent agent. The cluster instruction distribution terminal is connected to the multimodal instruction generation terminal and is used to perform object matching, timing arrangement, and distribution execution of the flight attitude control instructions, the ground motion control instructions, and the operation action control instructions according to the device type, task role, current location, and communication status of the airborne intelligent agent and the ground intelligent agent, so that the airborne intelligent agent and the ground intelligent agent can complete the spatial cooperation task at the location corresponding to the target docking point information.

2. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 1, characterized in that, The space-based environment sensing terminal includes a regional image acquisition module, a space-based pose synchronization module, and a probability grid generation module. The area image acquisition module is used to acquire target operation area image information collected by the aerial intelligent agent at different flight altitudes and different viewing angles; The airborne pose synchronization module is used to acquire the position information, flight altitude information and attitude angle information of the airborne intelligent agent corresponding to the image information of the target operation area, and to synchronize the position information, flight altitude information and attitude angle information of the airborne intelligent agent with the image information of the target operation area in time. The probability grid generation module is used to divide the target operation area into grids based on the target operation area image information after time synchronization, and generate terrain category information, obstacle occupancy probability information and passability probability information in each grid cell to obtain the empty base probability grid information.

3. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 2, characterized in that, The foundation state feedback terminal includes a ground pose acquisition module, an inertial disturbance acquisition module, a contact force acquisition module, and a stagnation state recognition module. The ground pose acquisition module is used to obtain the current position, orientation angle and travel speed of the ground intelligent agent in the target operation area; The inertial disturbance acquisition module is used to acquire information on the side tilt angle change, yaw angle change, and base vibration of the ground intelligent agent during its movement. The contact force acquisition module is used to acquire the supporting contact force, driving contact force, or foot contact force between the ground intelligent agent and the ground. The obstruction state identification module is used to identify the obstruction state of the ground agent at the corresponding position based on the travel speed, the roll angle change information, the yaw angle change information and the contact force information, and generate the obstruction information.

4. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 3, characterized in that, The heterogeneous spatial mapping terminal includes a coordinate unification module, an air-ground feature alignment module, and a collaborative spatial field generation module. The coordinate unification module is used to convert the flight coordinates of the airborne intelligent agent, the grid coordinates of the airborne probability grid information, and the ground motion coordinates of the ground intelligent agent to the generalized spatial coordinate system. The air-ground feature alignment module is used to spatially align the passability probability information, obstacle occupancy probability information, ground pose information, contact force information, and travel obstruction information at the same or adjacent spatial locations. The collaborative spatial field generation module is used to generate heterogeneous collaborative spatial mapping information, including air hovering suitability, ground traffic stability, and air-to-ground connection suitability, based on the spatially aligned information.

5. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 4, characterized in that, The collaborative connection solution terminal includes a candidate region screening module, a collaborative cost calculation module, and a target connection point determination module; The candidate area filtering module is used to filter candidate connection areas that meet the hovering conditions of the airborne intelligent agent and the arrival conditions of the ground intelligent agent from the heterogeneous collaborative spatial mapping information based on the spatial accessibility of the airborne intelligent agent, the travel stability of the ground intelligent agent, and the air-ground cooperation positional relationship. The collaboration cost calculation module is used to calculate the air hovering cost, ground passage cost, time synchronization cost, and collaboration execution cost for each candidate location in the candidate connection area. The target docking point determination module is used to determine target docking point information based on the hovering cost, the ground passage cost, the time synchronization cost, and the cooperative execution cost. The target docking point information includes the target docking location, the target docking time window, and the target docking attitude requirements.

6. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 5, characterized in that, The cooperation cost calculation module is also used to calculate the joint cooperation cost of the corresponding candidate location based on the flight energy consumption required for the airborne agent to reach the candidate location, the travel energy consumption required for the ground agent to reach the candidate location, the degree of ground contact disturbance, and the air-to-ground arrival time difference. The target connection point determination module is used to determine the candidate position that meets the preset connection constraints in terms of joint cooperation cost and has a stable advantage relative to the adjacent candidate position as the position corresponding to the target connection point information.

7. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 6, characterized in that, The multimodal command generation terminal includes an air command mapping module, a ground command mapping module, and an operation command mapping module; The airborne command mapping module is used to generate trajectory adjustment commands, speed adjustment commands, and hovering attitude control commands for the airborne intelligent agent based on the target docking location, target docking time window, and target docking attitude requirements in the target docking point information. The ground command mapping module is used to generate the ground intelligent agent's travel path command, speed adjustment command, and base attitude stabilization command based on the target docking location and target docking time window in the target docking point information. The operation instruction mapping module is used to generate motion position instructions, motion posture instructions, and motion triggering timing instructions for the end effector or mounting mechanism when the space collaborative task includes grasping, delivery, loading, or detection operations.

8. The spatial cooperative mapping and multimodal instruction distribution system for air-ground heterogeneous multi-agent clusters according to claim 7, characterized in that, The cluster instruction distribution terminal includes an object matching module, a timing orchestration module, a link status monitoring module, and a distribution confirmation module; The object matching module is used to match the flight attitude control command, the ground motion control command, and the operation action control command to the corresponding execution objects according to the device type, task role, and current position of the airborne intelligent agent and the ground intelligent agent; The timing arrangement module is used to arrange the triggering order of control instructions for different execution objects according to the target connection time window; The link status monitoring module is used to obtain the communication link quality and instruction reception status of different execution objects; The distribution confirmation module is used to send distribution confirmation information to the collaborative connection calculation terminal after confirming that the corresponding execution object has completed the instruction reception, and to trigger redistribution or recalculation of the target connection point when there is an instruction reception abnormality.

9. A spatial cooperative mapping and multimodal command distribution method for air-to-ground heterogeneous multi-agent clusters, applied to the spatial cooperative mapping and multimodal command distribution system for air-to-ground heterogeneous multi-agent clusters as described in claim 8, characterized in that, Spatial cooperative mapping and multimodal command distribution methods for heterogeneous air-ground multi-agent clusters include: S1, acquire airborne environmental perception information obtained by the aerial agent from the top-down perception of the target operation area, and generate airborne probability grid information representing the terrain distribution, obstacle distribution and passability probability distribution of the target operation area based on the airborne environmental perception information. S2, acquire ground motion status information and ground contact feedback information generated by the ground intelligent agent during its movement or operation within the target work area; S3, map the airborne probability grid information, the ground motion state information and the ground contact feedback information to the same generalized spatial coordinate system to generate heterogeneous cooperative spatial mapping information for characterizing the spatial reachability of airborne intelligent agents, the travel stability of ground intelligent agents and the positional relationship of air-ground cooperation; S4. Based on the heterogeneous collaborative space mapping information, combined with the hovering constraints of the airborne intelligent agent, the passage constraints of the ground intelligent agent, and the air-ground collaborative task constraints, candidate docking areas that meet the collaboration conditions between the airborne intelligent agent and the ground intelligent agent are determined, and the target docking point information is calculated according to the spatial reachability, passage stability, and collaboration execution cost corresponding to different locations in the candidate docking areas. S5. Based on the target docking point information, generate flight attitude control commands adapted to the execution of the airborne intelligent agent and ground motion control commands adapted to the execution of the ground intelligent agent, and generate operation action control commands adapted to the end-effector of the ground intelligent agent or the mounting mechanism of the airborne intelligent agent when the collaborative task involves end-effector operation. S6. Based on the device type, task role, current location, and communication status of the airborne intelligent agent and the ground intelligent agent, perform object matching, timing arrangement, and distribution execution of the flight attitude control command, the ground motion control command, and the operation action control command, so that the airborne intelligent agent and the ground intelligent agent can complete the spatial cooperation task at the location corresponding to the target docking point information.