Air-ground cooperation system and air-ground cooperation method based on own intelligence

Through the Vision-Language-Action (VLA) framework, the multimodal data fusion and intelligent decision-making of drones and unmanned vehicles are realized, solving the problem of insufficient perception and interaction capabilities of existing systems in complex environments, and improving the task execution efficiency and robustness of collaborative systems.

CN120560232APending Publication Date: 2025-08-29NANJING UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510680159.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing UAV and UAV collaboration systems have insufficient multimodal perception capabilities in complex environments, limited natural language interaction capabilities, low dynamic collaboration efficiency, and poor algorithm generalization, making it difficult to achieve cross-platform adaptation.

Method used

Using the Vision-Language-Action (VLA) framework, the natural language instructions and visual information are analyzed through the embodied intelligent unit, and multi-modal data is integrated to realize cross-modal data fusion and intelligent decision-making, dynamic task allocation and execution.

Benefits of technology

It improves the mission execution capabilities of drones and unmanned vehicles in complex environments, enhances multimodal perception and natural language interaction capabilities, and ensures the system's collaboration and robustness in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560232A_ABST
    Figure CN120560232A_ABST
Patent Text Reader

Abstract

The invention discloses an air-ground cooperation system and an air-ground cooperation method based on intelligence, and relates to the field of intelligent unmanned system control. The system comprises an intelligent unit, a multi-element sensing unit, a communication unit and a task execution unit. Based on a Vision-Language-Action (VLA) framework, an intelligent unit with a body realizes end-to-end analysis of a natural language instruction and a visual instruction, fuses sensing data of an unmanned aerial vehicle and an unmanned vehicle, generates a joint feature code with time-space alignment, and dynamically allocates a cooperative task. The communication unit supports low-delay end-to-end communication and real-time data synchronization between clusters, and ensures efficient interaction between agents. The multi-element sensing unit is integrated with a visual camera, a laser radar and high-precision positioning equipment, and air-ground environment multi-mode sensing is achieved. And the task execution unit supports the unmanned aerial vehicle to complete coordination actions in a complex scene. According to the scheme, the problems of insufficient multi-mode perception, limited natural language interaction and low dynamic cooperation efficiency of the existing system are solved, and the cooperation capability and system robustness of the unmanned aerial vehicle and the unmanned aerial vehicle in a complex environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent unmanned system control, and specifically to an air-ground collaborative system and an air-ground collaborative method based on embodied intelligence, which are particularly suitable for cross-modal interaction, dynamic task allocation and collaborative control of unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) in complex environments. Background Art

[0002] Amid the rapid development of artificial intelligence and robotics, air-ground collaborative systems, which combine drones and unmanned vehicles (UAVs) and their complementary capabilities, have shown significant potential for applications in disaster relief, logistics, military reconnaissance, and other fields. However, current mainstream systems rely on pre-set scripts or structured instructions, making it difficult to achieve natural interaction with human operators. Furthermore, their adaptability in complex and dynamic environments is limited, severely hindering their large-scale application.

[0003] The above-mentioned prior art has the following major defects:

[0004] 1. Insufficient multimodal perception capabilities: Existing systems fail to effectively integrate language commands with visual and lidar data, resulting in a disconnect between task understanding and execution.

[0005] 2. Limitations of natural language interaction: Reliance on structured command input and lack of end-to-end parsing and dynamic response capabilities for natural language commands.

[0006] 3. Low efficiency of dynamic collaboration: Task allocation and path planning of heterogeneous platforms (UAV and UGV) rely on human intervention and are difficult to adapt to real-time changes in complex environments.

[0007] Poor algorithm generalization: Traditional control methods are complex to tune and difficult to adapt across platforms, which limits the scalability of the system.

[0008] The Vision-Language-Action (VLA) model is an advanced multimodal machine learning framework that integrates vision, language, and action capabilities. Leveraging large-scale vision-language data and agent demonstrations, VLA enables agents to more efficiently learn new skills by fine-tuning pre-trained models, eliminating the need for starting from scratch. The framework aims to achieve a complete closed-loop capability, directly mapping perceptual input to robotic control actions. Summary of the Invention

[0009] To solve the above problems, this application provides an air-ground collaborative system and collaborative method based on embodied intelligence, which realizes natural language command parsing, cross-modal data fusion and intelligent decision-making through the Vision-Language-Action (VLA) framework, significantly improving the task execution capabilities of unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) in complex environments.

[0010] According to the first aspect of the present application, the present application provides an air-ground collaborative system based on embodied intelligence, comprising: an embodied intelligence unit, a multi-sensor unit, a communication unit, and a task execution unit. The multi-sensor unit is used to collect environmental data, target information, and intelligent body position data of the air-ground collaborative space to generate sensor data; the communication unit is signal-connected to the multi-sensor unit to realize cluster communication between each intelligent body and low-latency communication between each intelligent body and the terminal controller, ensuring real-time data interaction and cross-domain data synchronization of data transmission; the intelligent body includes a drone and an unmanned vehicle; the embodied intelligence unit is used to parse natural language instructions and visual instructions, fuse the sensor data, and issue task instructions based on dynamic task allocation strategies; the data transmitted by the communication unit includes the sensor data and the task instructions; the task execution unit is signal-connected to the communication unit to drive the drone and the unmanned vehicle to perform collaborative actions according to the task instructions to complete the preset action goals.

[0011] Furthermore, the embodied intelligence unit adopts the Vision-Language-Action (VLA) framework, including an input parsing module, a cross-modal fusion module, and a behavior generation module.

[0012] The input parsing module, comprising a language encoding submodule, a visual encoding submodule, and an action encoding submodule, is designed to accept and parse multimodal signal inputs containing natural language, visual images, and agent positions and actions, encoding them into vectors. The language encoding submodule uses a pre-trained language model to parse natural language instructions; the visual encoding submodule extracts image features; and the action encoding submodule encodes the position, trajectory, and motion behavior of drones and unmanned vehicles into vector representations.

[0013] The cross-modal fusion module is connected to the input parsing module signal and is used to process visual, language, and action coding information in the same time and space, so that the model can process the input of multimodal data in a common space.

[0014] The behavior generation module, signal-connected to the cross-modal fusion module, includes a behavior decision submodule and an action decoding submodule, which are responsible for generating behavior decisions, generating coordinated actions for the agents, and decoding and issuing them. The behavior decision submodule generates and assigns subtasks to each agent based on target instructions and environmental constraints; the action decoding submodule decodes the decision results into specific control instructions for the drone and unmanned vehicle.

[0015] Furthermore, the multimodal sensing unit includes a drone sensing module and an unmanned vehicle sensing module.

[0016] The drone sensor module, comprising a visual camera, a laser radar, and an RTK locator, is used to collect environmental data and location information from the drone. The visual camera captures RGB image data of the drone's surroundings. The laser radar emits short-pulsed laser beams to measure the distance, shape, and surface features of objects around the aircraft, generating high-precision three-dimensional point cloud data around the drone and enabling dynamic obstacle detection. The RTK locator supports centimeter-level precision positioning of the drone.

[0017] The unmanned vehicle sensor module includes a visual camera, a lidar, an RTK locator, and an inertial navigation system (INS), which are used to collect environmental data and position information on the unmanned vehicle. The visual camera of the unmanned vehicle sensor module is used to collect RGB image data around the unmanned vehicle; the lidar of the unmanned vehicle sensor module is used to scan the ground terrain, construct a ground grid map near the unmanned vehicle, and identify passable areas and obstacles; the RTK locator of the unmanned vehicle sensor module is used to support centimeter-level precision positioning of the unmanned vehicle; and the inertial navigation system of the unmanned vehicle sensor module uses inertial sensors (accelerometers and gyroscopes) to measure the acceleration and angular velocity of an object, and calculates the object's position, velocity, and attitude through integration operations. This is used to assist the unmanned vehicle in positioning when the RTK positioning signal is interfered with or in a denied environment.

[0018] Furthermore, the communication unit includes an end-to-end communication module and an inter-cluster communication module.

[0019] The end-to-end communication module handles communication between the terminal controller and each agent, enabling end-to-end information exchange between each agent and the decision-making core. The decision-making core transmits decoded task instructions to each agent, and the agent uploads its own location information to the terminal in real time.

[0020] The inter-cluster communication module is used for data communication between the intelligent agent clusters, supporting the real-time sharing of location information and behavior action information among the intelligent agents to achieve information interaction.

[0021] Furthermore, the task execution unit includes a drone execution module and an unmanned vehicle execution module.

[0022] The UAV execution module includes a robotic arm and a quadcopter power system. The robotic arm of the UAV execution module includes a robotic arm end gripper and a multi-degree-of-freedom joint, which achieves precise grasping through force feedback control. The quadcopter power system of the UAV execution module includes a battery, a galvanometer, an electronic speed controller, a brushless motor, and propellers, which are used to power the UAV and support flight and hovering functions.

[0023] The unmanned vehicle execution module includes a robotic arm and an omnidirectional drive power system. The robotic arm is used to carry heavy payloads and grasp objects. The omnidirectional drive power system includes a battery, ammeter, electronic speed controller, drive motor, and Mecanum wheels, which power the unmanned vehicle and enable all-terrain adaptive mobility.

[0024] According to a second aspect of the present application, an air-ground collaboration method based on embodied intelligence is provided, comprising the following steps:

[0025] Check the communication connection status between drones, unmanned vehicles and terminal controllers, and update the number and status of intelligent agents;

[0026] The embodied intelligence unit parses natural language instructions and visual instructions to generate task semantic vectors;

[0027] The multi-sensor unit collects RGB image data and three-dimensional point cloud data of the drone in the air and records its own position information to generate corresponding sensor data, and the multi-sensor unit collects a ground grid map and its own position information of the unmanned vehicle on the ground to generate corresponding sensor data, and the multi-sensor unit shares the sensor data with the embodied intelligent unit via the communication unit;

[0028] fusing the task semantic vector and the sensor data through the embodied intelligence unit to generate a spatiotemporally aligned joint feature encoding;

[0029] The embodied intelligent unit obtains the joint feature code and issues task instructions to the UAV and the unmanned vehicle based on a dynamic task allocation strategy, wherein the task instructions include the three-dimensional route of the UAV and the ground path of the unmanned vehicle, as well as the behavioral actions of the UAV and the unmanned vehicle;

[0030] The task execution unit receives the task instruction, drives the UAV and the unmanned vehicle to perform coordinated actions, and completes the preset action goals.

[0031] Furthermore, the embodied intelligent unit includes an input parsing module, which parses natural language instructions and visual instructions through the embodied intelligent unit to generate a task semantic vector, including: the input parsing module performs semantic segmentation and intent recognition on text instructions through a preset natural language processing model, extracts task keywords and action constraints, and generates a text semantic vector; the input parsing module performs feature extraction on the input visual image through a preset visual encoder to generate a visual semantic vector; the input parsing module performs spatiotemporal alignment on the text semantic vector and the visual semantic vector, and fuses them into a task semantic vector, wherein the spatiotemporal alignment is achieved through timestamp matching and spatial coordinate system conversion.

[0032] Furthermore, the embodied intelligent unit also includes a cross-modal fusion module, and the generation of spatiotemporal aligned joint feature coding by the embodied intelligent unit includes: the cross-modal fusion module uses a multi-head attention mechanism to perform cross-modal feature interaction on the task semantic vector, the three-dimensional point cloud data of the drone, and the raster map data of the unmanned vehicle; and uses a preset spatiotemporal encoder to embed the temporal information and spatial coordinates of the multimodal data into a unified feature space to generate a joint feature coding.

[0033] Furthermore, the embodied intelligent unit also includes a behavior generation module, and the embodied intelligent unit obtains the joint feature code and issues task instructions to the drone and the unmanned vehicle based on the dynamic task allocation strategy, including: the behavior generation module constructs an obstacle probability map based on the three-dimensional point cloud data of the drone, and generates a global path planning in combination with the unmanned vehicle grid map; the behavior generation module uses a dynamic window algorithm to detect moving obstacles in real time, and optimizes the local path through speed space sampling; the behavior generation module adjusts the motion trajectory of the drone and the unmanned vehicle according to the task priority and the path conflict level, and ensures collaborative obstacle avoidance; the behavior generation module issues task instructions to the drone and the unmanned vehicle based on the global path planning, the local path, and the motion trajectory.

[0034] Furthermore, the embodied intelligent unit implements a dynamic task migration mechanism. Specifically, during task execution, the task status is monitored in real time. If an anomaly such as communication interruption or low battery is detected, the dynamic task migration mechanism is triggered. Tasks are redistributed and paths are replanned based on the dynamic task migration mechanism. Tasks of abnormal vehicles are migrated to adjacent nodes, and the formation topology is adjusted and replanned. The paths of the remaining agents are updated based on current environmental data to ensure task continuity.

[0035] The beneficial effects of this application are:

[0036] The aforementioned embodied intelligence-based air-ground collaborative system and collaborative method utilizes an embodied intelligence unit based on the VLA framework to multimodally fuse sensor data from drones and unmanned vehicles with natural language and image input commands, enhancing the system's multimodal perception and natural language interaction capabilities. This framework boasts strong generalization capabilities, enabling real-time adjustments to mission strategies and improving the system's collaborative capabilities in complex environments. Furthermore, a dynamic task migration mechanism ensures stable operation even in the event of partial node failures, enhancing the system's robustness and security. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a system structure diagram of an air-ground collaborative system based on embodied intelligence for this application;

[0038] Figure 2This is a diagram showing the composition of an embodied intelligence unit in an air-ground collaborative system based on embodied intelligence for this application;

[0039] Figure 3 This is a diagram showing the composition of a multi-sensor unit for an air-ground collaborative system based on embodied intelligence;

[0040] Figure 4 A flow chart of a dynamic task migration mechanism is provided for this application;

[0041] Figure 5 The present application provides a flow chart of an air-ground collaboration method based on embodied intelligence. DETAILED DESCRIPTION

[0042] In the description of the present invention, it should be understood that the described embodiments are only part of the embodiments of the present invention, not all the embodiments.

[0043] Please refer to Figure 1 This example discloses an air-ground collaborative system based on embodied intelligence, which mainly includes an embodied intelligence unit 11, a multi-sensor unit 12, a communication unit 13, and a task execution unit 14. Each unit is introduced in detail below.

[0044] The embodied intelligence unit 11 is used to parse natural language instructions and visual instructions, integrate sensor data, and issue task instructions based on dynamic task allocation strategies.

[0045] The multi-sensor unit 12 is connected to the communication unit 13 by signal, and is used to collect environmental data, target information and intelligent body position data of the air-ground collaborative space to generate sensor data.

[0046] The signal connection between the communication unit 13 and the embodied intelligent unit 11 is used to realize cluster communication between each intelligent agent and low-latency communication between each intelligent agent and the terminal controller, ensuring real-time data interaction and cross-domain data synchronization of data transmission; the intelligent agents include drones and unmanned vehicles, and the data transmitted by the communication unit 13 includes sensor data and task instructions.

[0047] The task execution unit 14 is connected to the communication unit 13 by signal, and is used to drive the UAV and the unmanned vehicle to perform the collaborative task and complete the preset action target according to the task instruction.

[0048] In one embodiment, see Figure 1The embodied intelligence unit 11 includes an input parsing module 111, a cross-modal fusion module 112, and a behavior generation module 113. The input parsing module 111 receives and parses multimodal signals containing natural language, visual images, agent position information, and actions. The cross-modal fusion module 112 is connected to the input parsing module 111 by signal, receives the output of the input parsing unit 111, and processes the language, vision, and action encoding information in the same time and space. The behavior generation module 113 is connected to the cross-modal fusion module 112 by signal, generates behavior decisions, generates coordinated actions for each agent, and decodes and sends them.

[0049] In one embodiment, see Figure 1 The multi-sensor unit 12 includes a drone sensor module 121 and an unmanned vehicle sensor module 122. The drone sensor module 121 is mounted on the drone and is used to collect environmental data and location information. The unmanned vehicle sensor module 122 is mounted on the unmanned vehicle and is used to collect environmental data and location information. It will be appreciated that the drone sensor module 121 and the unmanned vehicle sensor module 122 contain high-precision sensors for acquiring information about the agent's surroundings and its own location data.

[0050] In one embodiment, see Figure 1 The communication unit 13 includes an end-to-end communication module 131 and an inter-cluster communication module 132. The end-to-end communication module 131 handles data communication between the terminal controller and each agent, enabling end-to-end information exchange between each agent and the decision-making core. It can be understood that the decision-making core sends task instructions to each agent via the end-to-end communication module 131, and the agents upload sensor data to the terminal in real time via the end-to-end communication module 131. The inter-cluster communication module 132 uses data transmission to exchange location and behavior information between agent clusters.

[0051] In one embodiment, see Figure 1 The task execution unit 14 includes a drone execution module 141 and an unmanned vehicle execution module 142. The drone execution module 141 and the unmanned vehicle execution module 142 are mechanical devices installed on the drone and the unmanned vehicle, respectively. For example, the drone execution module 141 may include a four-rotor power system and a robotic arm. The four-rotor power system provides the drone with the ability to fly and hover, and the robotic arm supports the drone to perform some grasping tasks. The unmanned vehicle execution module 142 may include an omnidirectional drive power system and a robotic arm. The omnidirectional drive power system provides the unmanned vehicle with all-terrain adaptive mobility capabilities, and the robotic arm supports the unmanned vehicle to complete the tasks of grasping and large load handling.

[0052] In one embodiment, see Figure 2The embodied intelligence unit 11 includes an input parsing module 111, a cross-modal fusion module 112, and a behavior generation module 113. The input parsing module 111 includes a language encoding submodule 21, a visual encoding submodule 22, and an action encoding submodule 23, while the behavior generation module 113 includes a behavior decision submodule 24 and an action decoding submodule 25.

[0053] In one embodiment, see Figure 2 , the language encoding submodule 21 processes the natural language input received by the system, parses the natural language instructions, and encodes the natural language instructions using a large language model based on the Transformer architecture; the visual encoding submodule 22 processes the image input, extracts image features, and encodes the visual image using a convolutional neural network (CNN) or Transformers technology; the action encoding submodule 23 converts the position information and actions of the drone and the unmanned vehicle into a low-dimensional vector space representation, and designs a set of encoding rules for the position information and the actions that the drone and the unmanned vehicle can perform. It can be understood that for the air-ground collaborative system based on embodied intelligence, natural language instructions and visual instructions can be input when the system is working. Each intelligent agent involves interaction with the environment in the execution of the task and needs to record the position of the drone and the unmanned vehicle itself and the execution action in real time. The three inputs are encoded using encoders respectively to understand the multimodal semantic features.

[0054] In one embodiment, see Figure 2 The behavior decision submodule 24 uses the multi-agent deep reinforcement learning (MADDPG) algorithm, combined with environmental constraints, to assign subtasks to each agent. The action decoding submodule 25 generates specific control instructions for the drone and unmanned vehicle based on the decision results, decodes them into a format understandable to the drone and unmanned vehicle, and issues the control instructions, allowing the drone and unmanned vehicle to perform coordinated actions.

[0055] In one embodiment, see Figure 3 The multi-sensor unit 12 includes a drone sensor unit 121 and an unmanned vehicle sensor unit 122. The drone sensor unit 121 includes a visual camera 31, a laser radar 32, and an RTK positioner 33, while the unmanned vehicle sensor unit 122 includes a visual camera 34, a laser radar 35, an RTK positioner 36, and an inertial navigation system 37.

[0056] In one embodiment, see Figure 3The visual camera 31 is installed on the UAV to collect RGB image data of the environment around the UAV; the laser radar 32 emits a short-pulse laser beam to measure the features of objects around the aircraft, which is used to generate high-precision three-dimensional point cloud data around the UAV and dynamically detect obstacles; the RTK locator 33 is used to support centimeter-level precise positioning of the UAV; the visual camera 34 is installed on the unmanned vehicle to collect RGB image data around the unmanned vehicle; the laser radar 35 is used to generate a ground grid map around the unmanned vehicle and detect obstacles; the RTK locator 36 is used to support centimeter-level precise positioning of the unmanned vehicle; and the inertial navigation system 37 is used for auxiliary positioning of the unmanned vehicle when the RTK positioning signal is interfered with or in a denied environment.

[0057] In one embodiment, based on the aforementioned embodied intelligence-based air-ground collaboration system, the present application further provides a corresponding air-ground collaboration method, which includes the following steps:

[0058] (1) Check the communication connection status between the UAV, the unmanned vehicle and the terminal controller, and update the number and status of the intelligent agents.

[0059] (2) The embodied intelligent unit parses natural language instructions and visual instructions to generate task semantic vectors. For example, the input parsing module uses a preset natural language processing model to perform semantic segmentation and intent recognition on text instructions, extracts task keywords and action constraints, and generates text semantic vectors; the input parsing module uses a preset visual encoder to extract features from the input visual image and generates visual semantic vectors; the input parsing module performs spatiotemporal alignment on the text semantic vector and the visual semantic vector to fuse them into a task semantic vector, where spatiotemporal alignment is achieved through timestamp matching and spatial coordinate system conversion.

[0060] (3) The multi-sensor unit collects RGB image data and three-dimensional point cloud data of the UAV in the air and records its own position information to generate corresponding sensor data, and the multi-sensor unit collects the ground grid map and its own position information of the UAV on the ground to generate corresponding sensor data. The multi-sensor unit shares the sensor data with the embodied intelligent unit through the communication unit;

[0061] (4) The embodied intelligence unit fuses the task semantic vector and sensor data to generate a spatiotemporally aligned joint feature code. For example, the embodied intelligence unit also includes a cross-modal fusion module, which uses a multi-head attention mechanism to perform cross-modal feature interaction on the task semantic vector, the three-dimensional point cloud data of the drone, and the raster map data of the unmanned vehicle; and uses a preset spatiotemporal encoder to embed the temporal information and spatial coordinates of the multimodal data into a unified feature space to generate a joint feature code.

[0062] (5) The embodied intelligence unit obtains joint feature coding and issues task instructions to the UAV and the unmanned vehicle based on a dynamic task allocation strategy. The task instructions include the UAV's three-dimensional route and the unmanned vehicle's ground path and are used to avoid dynamic obstacles in real time. For example, the embodied intelligence unit also includes a behavior generation module. The behavior generation module constructs an obstacle probability map based on the UAV's three-dimensional point cloud data and generates a global path plan based on the unmanned vehicle's grid map. The behavior generation module uses a dynamic window algorithm to detect moving obstacles in real time and optimizes the local path through speed space sampling. The behavior generation module adjusts the motion trajectory of the UAV and the unmanned vehicle according to the task priority and path conflict level and ensures collaborative obstacle avoidance. The behavior generation module issues task instructions to the UAV and the unmanned vehicle based on the global path plan, local path, and motion trajectory.

[0063] (5) Receive task instructions through the task execution unit, drive the UAV and the unmanned vehicle to perform coordinated actions and complete the preset action goals.

[0064] In this application, the air-ground collaboration method also includes executing a dynamic task migration mechanism through an embodied intelligent unit, specifically including: during the task execution process, real-time monitoring of the task status, and triggering the dynamic task migration mechanism if an anomaly is detected; performing task redistribution and path replanning based on the dynamic task migration mechanism, migrating the tasks of the abnormal intelligent agent to the adjacent node, and adjusting the formation topology path replanning, updating the paths of the remaining intelligent agents based on the current environmental data, and ensuring task continuity.

[0065] In a specific embodiment, see Figure 4 In response to unexpected situations, this system adopts a dynamic task migration mechanism. The dynamic migration mechanism includes:

[0066] Abnormality Detection S41: Real-time monitoring of each agent's status data. If a drone or unmanned vehicle is deemed abnormal, a dynamic migration mechanism is initiated. Understandably, battery charge levels below a threshold or continuous packet loss are considered abnormal.

[0067] Task Reassignment S42: The decision core reassigns the abnormal vehicle task. For example, if a drone exits a reconnaissance mission due to low battery, the task is transferred to the drone closest to the target area based on the real-time location and task priority of the remaining drones, and the task queue is updated.

[0068] Formation adjustment S43: Re-adjust the drone swarm or unmanned vehicle swarm formation according to the tasks of the remaining agents.

[0069] Path replanning S44: After the remaining agents update the task queue, the decision core generates a new action path for each agent and sends it to the drones and unmanned vehicles.

[0070] In a specific embodiment, please refer to Figure 5 .

[0071] System initialization S501. This step is the beginning of the work of the air-ground collaborative system. The UAV and unmanned vehicle are powered on, communicate with the terminal controller, establish a communication link between the intelligent agents, and the terminal controller (decision-making core) enters the ready state.

[0072] System self-check S502, before starting work, check whether the system has any abnormalities, and construct the unit state vector S = [s1, s2, ..., s N ](s i =1 means normal, s i =0 indicates an abnormality). For example, check whether the UAV and the unmanned vehicle have sufficient battery power, whether the RTK locator fails to locate correctly, and whether the communication link is normal.

[0073] Error troubleshooting S503 is performed after the system self-check S502 fails. It is used to troubleshoot the error reported in the system self-check S502. For example, if a drone shows that the GPS is not positioned, check the drone's RTK positioner.

[0074] Command parsing (S504), executed after the system passes self-check (S502), parses the input natural language and visual commands to generate a task semantic vector. For example, the command "drone searches area A for item B, autonomous vehicle goes to grab and bring back" is encoded into a vector.

[0075] Environmental Perception S505 is executed after the system self-check S502 passes. The drone and the unmanned vehicle perceive their surroundings through the multi-sensor unit 12 and generate sensor data. For example, the drone uses the visual camera 31 to collect RGB image data of the drone's surroundings, and the lidar 32 to collect real-time 3D point cloud data of the drone's surroundings. The unmanned vehicle uses the lidar 35 to construct a ground grid map of the unmanned vehicle's surroundings. This sensor data is shared with the decision-making core through the communication unit 13.

[0076] Cross-modal fusion S506: The cross-modal fusion unit 112 fuses the task semantic vector and the sensor data to generate a spatiotemporally aligned joint coding feature.

[0077] During task assignment S507, the behavior decision submodule 24 acquires the joint feature code and generates tasks for each drone and vehicle based on the dynamic task assignment strategy. For example, based on the instructions mentioned in the instruction parsing S504, the drone swarm searches area A. After detecting object B via the multi-sensor unit 12, it transmits its coordinates to the vehicle via the communication unit 13. The vehicle then travels to the coordinates, identifies object B, and uses its robotic arm to grab it before returning to its starting coordinates.

[0078] Path planning S508: After task assignment S507, each agent is assigned a specific path. For example, based on the tasks mentioned in task assignment S507, flight paths are planned for each drone swarm, generating a set of waypoint coordinates {[x1, y1, z1], ...}.

[0079] Action decoding S509 : decoding the decision core output content into task instructions through the action decoding submodule 25 .

[0080] Instruction distribution S510, the decoded task control instruction is distributed to each agent through the communication unit 13.

[0081] In the task execution S511, each agent that receives the task instruction executes the preset task through the task execution unit 14. After the task execution is completed, the system returns to the system self-check step S502 to check the system.

[0082] The dynamic task migration mechanism S513 is always enabled during task execution to deal with unexpected situations during task execution. Figure 4 .

[0083] The above content is a further detailed description of the present application in conjunction with specific implementation methods, and the specific implementation of the present application cannot be considered to be limited to these descriptions. For ordinary technicians in the technical field to which the present application belongs, several simple deductions or substitutions can be made without departing from the inventive concept of the present application.

Claims

1. An air-ground collaborative system based on embodied intelligence, characterized by: include: Embodied intelligence unit, multi-sensing unit, communication unit, task execution unit; The multi-sensor unit is used to collect environmental data, target information and agent position data of the air-ground collaborative space to generate sensor data; The communication unit is connected to the multi-sensor unit signal, and is used to realize cluster communication between each intelligent agent and low-latency communication between each intelligent agent and the terminal controller, ensuring real-time data interaction and cross-domain data synchronization of data transmission; the intelligent agent includes drones and unmanned vehicles; The embodied intelligence unit is used to parse natural language instructions and visual instructions, integrate the sensor data, and issue task instructions based on dynamic task allocation strategies; The data transmitted by the communication unit includes the sensing data and the task instructions; The task execution unit is connected to the communication unit by signal, and is used to drive the UAV and the unmanned vehicle to perform coordinated actions according to the task instructions to complete the preset action goals.

2. The air-ground collaborative system based on embodied intelligence according to claim 1, characterized in that: The embodied intelligence unit includes an input parsing module, a cross-modal fusion module, and a behavior generation module; The input parsing module is used to receive and parse multimodal signal inputs containing natural language, visual images, agent position information and movements, and encode them into vectors; The cross-modal fusion module is signal-connected to the input parsing module to process language, vision, and motion encoding information in the same space and time, so that the model can process multimodal data input in a common space; The behavior generation module is connected to the cross-modal fusion module signal to generate behavior decisions, generate intelligent body collaborative actions and decode and send them.

3. The air-ground collaborative system based on embodied intelligence according to claim 1, characterized in that: The multi-sensor unit includes a drone sensing module and an unmanned vehicle sensing module; The drone sensor module includes a visual camera, a laser radar, and an RTK locator, which are used to collect environmental data and location information on the drone. The visual camera of the drone sensor module is used to collect RGB image data around the drone. The laser radar of the drone sensor module is used to collect three-dimensional point cloud data around the drone. The RTK locator of the drone sensor module is used to support centimeter-level precise positioning of the drone. The unmanned vehicle sensing module includes a visual camera, a lidar, an RTK locator, and an inertial navigation system, which are used to collect environmental data and location information on the unmanned vehicle; the visual camera of the unmanned vehicle sensing module is used to collect RGB image data around the unmanned vehicle; the lidar of the unmanned vehicle sensing module is used to construct a ground grid map; the RTK locator of the unmanned vehicle sensing module is used to support centimeter-level precise positioning of the unmanned vehicle; the inertial navigation system of the unmanned vehicle sensing module is used for auxiliary positioning of the unmanned vehicle when the RTK positioning signal is interfered with or in a denied environment.

4. The air-ground collaborative system based on embodied intelligence according to claim 1, characterized in that: The communication unit includes an end-to-end communication module and an inter-cluster communication module; The end-to-end communication module is used to handle the communication between the terminal controller and each intelligent agent, and realize the end-to-end information interaction between each intelligent agent and the decision-making core; The inter-cluster communication module is used for data communication between the agent clusters, and supports the real-time sharing of location information and behavior action information among the agents.

5. The air-ground collaborative system based on embodied intelligence according to claim 1, wherein the task execution unit includes a drone execution module and an unmanned vehicle execution module; The UAV execution module includes a robotic arm and a quad-rotor power system; the robotic arm of the UAV execution module is used to achieve a precise grasping function; the quad-rotor power system of the UAV execution module includes a battery, an ammeter, an electronic speed controller, a brushless motor, and a propeller, which are used to power the UAV and support flight and hovering functions; The unmanned vehicle execution module includes a robotic arm and an omnidirectional drive power system; the robotic arm of the unmanned vehicle execution module is used to perform large-load handling and object grasping tasks; the omnidirectional drive power system of the unmanned vehicle execution module includes a battery, an ammeter, an electronic regulator, a drive motor, and a Mecanum wheel, which are used to power the unmanned vehicle and support all-terrain adaptive movement.

6. An air-ground collaboration method based on embodied intelligence, characterized in that: Applied to the air-ground collaboration system according to any one of claims 1 to 5, the air-ground collaboration method comprises: Check the communication connection status between the UAV, unmanned vehicle and terminal controller, and update the number and status of intelligent agents; The embodied intelligence unit parses natural language instructions and visual instructions to generate task semantic vectors; The multi-sensor unit collects RGB image data and three-dimensional point cloud data of the drone in the air and records its own position information to generate corresponding sensor data, and the multi-sensor unit collects a ground grid map and its own position information of the unmanned vehicle on the ground to generate corresponding sensor data, and the multi-sensor unit shares the sensor data with the embodied intelligent unit via the communication unit; fusing the task semantic vector and the sensor data through the embodied intelligence unit to generate a spatiotemporally aligned joint feature encoding; The embodied intelligent unit obtains the joint feature code and issues task instructions to the UAV and the unmanned vehicle based on a dynamic task allocation strategy, wherein the task instructions include the three-dimensional route of the UAV and the ground path of the unmanned vehicle, as well as the behavioral actions of the UAV and the unmanned vehicle; The task execution unit receives the task instruction and drives the UAV and the unmanned vehicle to perform coordinated actions to complete the preset action goals.

7. The air-ground collaboration method according to claim 6, characterized in that: The embodied intelligence unit includes an input parsing module, and the embodied intelligence unit parses natural language instructions and visual instructions to generate a task semantic vector, including: The input parsing module performs semantic segmentation and intent recognition on the text instructions through a preset natural language processing model, extracts task keywords and action constraints, and generates a text semantic vector; The input parsing module extracts features from the input visual image through a preset visual encoder to generate a visual semantic vector; The input parsing module performs spatiotemporal alignment on the text semantic vector and the visual semantic vector and fuses them into a task semantic vector, wherein the spatiotemporal alignment is achieved through timestamp matching and spatial coordinate system conversion.

8. The air-ground collaboration method according to claim 7, characterized in that: The embodied intelligence unit further includes a cross-modal fusion module, and the embodied intelligence unit generates a spatiotemporal aligned joint feature code, including: The cross-modal fusion module uses a multi-head attention mechanism to perform cross-modal feature interaction on the task semantic vector, the three-dimensional point cloud data of the drone, and the raster map data of the unmanned vehicle; and uses a preset spatiotemporal encoder to embed the temporal information and spatial coordinates of the multimodal data into a unified feature space to generate a joint feature code.

9. The air-ground collaboration method according to claim 6, characterized in that: The embodied intelligent unit further includes a behavior generation module, and obtaining the joint feature code through the embodied intelligent unit and issuing task instructions to the UAV and the unmanned vehicle based on the dynamic task allocation strategy includes: The behavior generation module constructs an obstacle probability map based on the three-dimensional point cloud data of the UAV, and generates a global path plan by combining it with the grid map of the unmanned vehicle; The behavior generation module uses a dynamic window algorithm to detect moving obstacles in real time and optimizes the local path through speed space sampling; The behavior generation module adjusts the motion trajectories of the UAV and the unmanned vehicle and ensures collaborative obstacle avoidance based on task priorities and path conflict levels; The behavior generation module issues task instructions to the UAV and the unmanned vehicle based on the global path planning, the local path, and the motion trajectory.

10. The air-ground collaboration method according to claim 6, characterized in that: The embodiment also includes executing a dynamic task migration mechanism through the embodied intelligent unit, specifically including: During task execution, the task status is monitored in real time, and if an anomaly is detected, the dynamic task migration mechanism is triggered; Based on the dynamic task migration mechanism, task redistribution and path replanning are performed to migrate the tasks of abnormal agents to adjacent nodes, and the formation topology path replanning is adjusted. The paths of the remaining agents are updated based on the current environmental data to ensure task continuity.

Citation Information

Cited By

  • Air-ground collaborative unified semantic interoperation method and system

    CN120725028A

  • Unmanned aerial vehicle body cognition alignment method based on man-machine cooperation

    CN120803003A

  • Vision-based cross-network interaction method and system

    CN121209719A

  • Autonomous control method and system for disconnection of unmanned aerial vehicle

    CN121900452A

  • An unmanned aerial vehicle air-ground cooperative pathfinding method and system based on a federal large model

    CN122488803A