Ground-air robot storage environment intelligent sensing and anomaly detection system based on AI vision

By constructing an intelligent perception and anomaly detection system for the air-to-ground robotic warehouse environment, the problem of isolated perception perspectives between the aerial platform and the ground platform was solved, achieving high-precision three-dimensional environmental mapping and anomaly detection, thereby improving the intelligence and automation level of warehouse management.

CN121599587APending Publication Date: 2026-03-03SHANGHAI TEKU INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511738924.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In existing warehousing robot systems, the perception perspectives of the aerial platform and the ground platform are isolated, and the data fusion level is shallow, resulting in incomplete anomaly detection and low collaborative efficiency.

Method used

Construct an AI vision-based intelligent perception and anomaly detection system for ground-to-air robotic warehousing environments, including multimodal data acquisition, ground-to-air collaborative localization and mapping, unified semantic cognitive model, multi-scale anomaly detection and reasoning, and task dynamic planning and execution unit, to achieve deep collaborative perception and collaborative operation between air and ground platforms.

Benefits of technology

It achieves globally unified and high-precision 3D environment mapping, improves the completeness and accuracy of environmental perception, can deeply understand the semantic information of the environment, significantly improves the depth and breadth of anomaly detection, and realizes a complete closed loop from perception to execution, thereby improving the automation and intelligence level of warehouse management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121599587A_ABST
    Figure CN121599587A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of crossing of artificial intelligence and intelligent robots, and particularly relates to a ground-air robot storage environment intelligent sensing and anomaly detection system based on AI vision. The invention discloses an AI vision-based intelligent sensing and anomaly detection system for a storage environment of a ground-air robot, and aims to solve the problem of incomplete anomaly detection caused by isolated sensing and shallow data fusion of a ground-air platform. The system comprises a multi-modal acquisition unit, a ground-air collaborative mapping unit, a unified semantic cognition unit, a multi-scale anomaly detection unit and a dynamic task planning unit. Through fusion of air wide-area perception and ground fine observation data, a high-precision three-dimensional semantic map is constructed, comprehensive detection of macroscopic layout, microscopic state and time sequence logic abnormity is realized, and the ground-air robot is driven to cooperatively respond, so that the perception integrity and intelligent decision-making capability of a storage environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology of artificial intelligence and intelligent robots, specifically relating to an AI vision-based intelligent perception and anomaly detection system for ground and air robot warehouse environments. Background Technology

[0002] Modern warehousing and logistics systems are rapidly developing towards automation and intelligence, with robotics technology serving as a core driving force, playing a crucial role in improving warehousing efficiency, reducing operating costs, and ensuring operational safety. Environmental perception is the foundation for achieving intelligent warehousing. Deploying various sensors and monitoring equipment to monitor the status of goods, equipment operation, and personnel activities in the warehousing environment in real time is a core element in ensuring smooth and safe warehousing processes.

[0003] Among these, utilizing mobile robots for environmental inspection and perception has become a key technological branch. By endowing robots with mobility and visual perception capabilities, the aim is to overcome the problems of blind spots and limited coverage of traditional fixed monitoring equipment, thereby obtaining more comprehensive and dynamic environmental information and providing data support for warehouse management decisions.

[0004] Current technologies for applying robots to perceive warehouse environments still face numerous challenges. Single-type robot systems have inherent limitations. For example, ground robots' field of vision is easily obstructed by shelves, and their movement paths are limited by ground aisles. While aerial robots have a wide field of vision, they struggle to acquire detailed ground information and have limited endurance. Furthermore, visual data collected by different robot platforms varies significantly in perspective, scale, and resolution, making the fusion of multi-source heterogeneous information extremely difficult and hindering the formation of a unified and accurate environmental perception. In addition, traditional anomaly detection algorithms often suffer from insufficient detection accuracy and high false alarm rates when dealing with complex and changing warehouse backgrounds due to factors such as lighting variations, item stacking, and dynamic interference, failing to meet the demands of high-reliability warehouse management. Therefore, how to construct an intelligent system that can integrate multi-dimensional ground and aerial perspectives, achieving efficient collaboration and accurate anomaly detection, has become a pressing technical challenge. Summary of the Invention

[0005] The purpose of this invention is to address the technical problems in existing warehouse robot systems, such as isolated perception perspectives between aerial and ground platforms, shallow data fusion levels, and difficulty in forming a unified environmental understanding, which leads to incomplete anomaly detection and low collaborative efficiency. This invention provides an AI vision-based intelligent perception and anomaly detection system for ground-to-air robot warehouse environments.

[0006] To achieve the above objectives, the technical solution provided by this invention is: an AI vision-based intelligent perception and anomaly detection system for ground-to-air robotic warehouse environments, which includes a multimodal data acquisition unit, a ground-to-air collaborative positioning and mapping unit, a unified semantic cognitive model unit, a multi-scale anomaly detection and reasoning unit, and a task dynamic planning and execution unit.

[0007] The multimodal data acquisition unit is configured to simultaneously acquire raw perception data of the warehouse environment from both aerial and ground perspectives. This unit comprises an aerial perception subunit and a ground perception subunit. The aerial perception subunit, mounted on an aerial robotic platform, integrates a high-resolution wide-angle camera, LiDAR, and inertial measurement module. It is responsible for capturing color image information, 3D point cloud data, and its own motion attitude data of a large-scale scene from a high-altitude perspective. The ground perception subunit, mounted on a ground robotic platform, integrates a high-magnification zoom camera, a structured light scanner, and a thermal imaging camera. It is responsible for performing detailed observation tasks at close range, acquiring high-resolution texture details, microscopic 3D morphological data, and temperature distribution information of specific targets.

[0008] The ground-to-air cooperative localization and mapping unit is connected to the data output of the multimodal data acquisition unit. Its function is to fuse heterogeneous data from both aerial and ground perspectives to construct a globally unified, high-precision 3D environmental map in real time, and to perform joint localization of the ground-to-air robotic platform within this map. This unit first utilizes wide-area data collected by the aerial perception subunit to generate a global sparse feature map through a vision-liquid radar fusion algorithm. This sparse feature map is transmitted to the ground robotic platform as a global pose prior for visual odometry calculations, effectively constraining the accumulation of its localization errors. Simultaneously, the ground perception subunit, during close-range observation, performs high-precision 3D measurements of stable, unchanging static features in the environment, such as shelf pillars and ground markings, and feeds back the precise world coordinates of these feature points to the aerial robotic platform. These precise coordinate points serve as global anchor points to correct trajectory drift generated by the aerial robotic platform during long-term flight, ultimately forming a globally consistent, dense 3D point cloud map that integrates macroscopic structure and microscopic details.

[0009] Furthermore, the unified semantic cognition model unit receives the fused 3D map and associated image data generated by the ground-air cooperative positioning and mapping unit, aiming to transform the raw geometric and visual information into machine-understandable structured semantic knowledge. This unit deploys a deep learning-driven scene understanding model. This model first performs instance segmentation on the fused point cloud map, identifying and separating individual physical entities in the warehouse environment, such as shelves, pallets, cargo boxes, robots, and workers. Subsequently, the model combines corresponding image texture information to perform category recognition, state assessment, and attribute extraction for each entity. State assessment includes determining whether the cargo packaging is intact, damaged, or leaking. Attribute extraction covers reading barcodes, QR codes, or text labels on cargo boxes. Finally, the unit outputs a dynamically updated digital twin scene map of the warehouse environment. In this scene map, nodes represent physical entities and their semantic attributes, and edges represent spatial relationships between entities, such as support, stacking, and proximity, as well as logical relationships, such as ownership and association.

[0010] As one embodiment of the present invention, the multi-scale anomaly detection and reasoning unit performs comprehensive anomaly state detection and analysis based on the dynamic scene graph output by the unified semantic cognitive model unit. The unit's working logic is divided into three levels. The first level is macro-layout anomaly detection, which detects large-scale spatial structural anomalies such as collapsed shelves, illegally occupied fire exits, and misplaced bulk goods by comparing the real-time scene graph with a preset standard warehouse layout model. The second level is micro-state anomaly detection, which identifies fine-grained entity state anomalies such as damaged packaging, liquid leakage, abnormal equipment surface temperature, and unreadable label information by analyzing the attributes of individual entities in the scene graph. The third level is temporal logic anomaly reasoning, where the unit records and analyzes the sequence of entity states evolving over time in the scene graph. By comparing this actual state sequence with the standard operating procedure model defined in the warehouse management system, the unit can infer process-related anomalies, such as goods exceeding the time limit for temporary storage, equipment movement deviating from the predetermined path, or unauthorized personnel activity.

[0011] Furthermore, the task dynamic planning and execution unit, serving as the decision-making and control center of the entire system, is connected to the output of the multi-scale anomaly detection and inference unit. Upon receiving any type of anomaly alarm signal, this unit first classifies the severity and type of the anomaly event according to a predefined rule base. Based on the classification results, the unit automatically generates a ground-air collaborative verification or handling task. This task is decomposed into a series of specific action commands for the aerial robot platform and the ground robot platform. For example, for a suspected ground leak anomaly detected from an aerial perspective, the task planning and execution unit first instructs the aerial robot to adjust its hovering position and angle for secondary confirmation from multiple perspectives, while simultaneously instructing the ground robot to plan a collision-free path to the target area and use its high-magnification camera and thermal imaging camera for close-up detailed investigation. The task commands are converted into robot-executable motion control commands and sent to the respective underlying control systems of the ground and aerial robots for execution, forming a complete closed loop from perception, cognition, detection, to action.

[0012] In another embodiment of the present invention, the unified semantic cognition model unit employs a deep learning model trained on a hybrid dataset consisting of synthetic and real data. The synthetic data is obtained through physical rendering and procedural generation of the warehouse's 3D design model, covering various extreme lighting conditions and rare anomaly samples. The real data is continuously collected and manually labeled in the actual system deployment environment, used for the model's continuous adaptation and optimization to specific environmental features. This hybrid training strategy ensures that the model possesses generalization capabilities while also exhibiting high specificity and recognition accuracy for specific application scenarios.

[0013] In another embodiment of the present invention, a two-way information feedback channel is established between the ground-air cooperative positioning and mapping unit and the task dynamic planning and execution unit. During task execution, if the ground robot's path planning is hindered by incomplete local map information, the task dynamic planning and execution unit will proactively request the ground-air cooperative positioning and mapping unit to instruct the aerial robot to fly to that area for supplementary surveying and to update the local environmental map in real time, thereby providing decision support for the ground robot. This dynamic, on-demand perception mechanism significantly improves the robot's autonomous navigation and task execution success rate in complex and dynamically changing environments.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: 1. This invention constructs a ground-air collaborative positioning and mapping unit, deeply integrating the wide-area perspective of an aerial platform with the detailed perspective of a ground platform to generate a globally unified and highly detailed 3D environmental model. This solution fundamentally resolves the contradiction between the perception range and accuracy of a single platform, achieving comprehensive and seamless digital mapping of the warehousing environment. It provides an unprecedented high-quality data foundation for upper-level intelligent applications, significantly improving the completeness and accuracy of environmental perception.

[0015] 2. The unified semantic cognition model unit proposed in this invention transforms raw, unstructured perceptual data into structured, dynamic scene graphs containing rich semantic information. This transformation surpasses traditional geometric reconstruction, endowing the system with a profound understanding of the environment, enabling it to recognize the categories, states, and interrelationships of objects. This semantic understanding-based perceptual mode allows anomaly detection to move beyond simple geometric deviations and delve into functional and logical levels, significantly improving the depth and breadth of anomaly detection.

[0016] 3. The multi-scale anomaly detection and reasoning unit designed in this invention integrates detection capabilities across three dimensions: macroscopic layout, microscopic state, and temporal logic. In particular, the introduction of temporal logic anomaly reasoning enables the system to discover hidden problems that appear normal in a single time slice but are abnormal in the overall business process. This marks the evolution of the system from a static environmental monitor to an intelligent analysis engine capable of understanding the efficiency and compliance of dynamic production processes, achieving a leap from "discovering problems" to "predicting risks."

[0017] 4. This invention constructs a complete closed loop from perception to execution. The task dynamic planning and execution unit can autonomously generate and issue ground-air collaborative tasks based on detected anomalies. This tightly integrated perception-decision-action design ensures rapid response and effective handling of abnormal events, integrating ground and air robots from two independent data acquisition tools into an organically collaborative, autonomously problem-solving intelligent operating team, significantly improving the automation and intelligence level of warehouse management. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the unified semantic cognition model unit in this invention. Detailed Implementation

[0019] Example 1 Please refer to Figure 1 and Figure 2To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.

[0020] This invention provides an AI vision-based intelligent perception and anomaly detection system for ground-to-air robotic warehouse environments. It aims to construct a dynamic, high-precision digital twin of the warehouse environment with semantic understanding capabilities through deep collaboration between aerial and ground robotic platforms. Based on this, it achieves multi-scale, multi-dimensional anomaly detection and autonomous response, thereby overcoming the technical bottlenecks of limited perception, superficial understanding, and inefficient collaboration in existing automated warehouse systems. In this embodiment, the system is deployed in a large-scale automated warehouse environment and consists of an aerial robotic platform, a ground robotic platform, and a central computing cluster deployed on a local edge computing server. The various components interact with each other in real time via a high-bandwidth, low-latency dedicated wireless network.

[0021] The system's overall technical process begins with the multimodal data acquisition unit synchronously capturing multi-view data of the warehouse environment. The captured data stream is then fed into the ground-air collaborative positioning and mapping unit, which is responsible for fusing heterogeneous sensor information to generate a globally unified, precisely measured 3D environment map, and continuously tracking the precise poses of the two robot platforms within this map. Subsequently, the unified semantic cognition model unit performs deep analysis of this 3D map and associated visual data, elevating the raw geometric information into structured semantic knowledge and constructing a dynamically updated scene map. Based on this scene map, the multi-scale anomaly detection and reasoning unit performs comprehensive anomaly detection and root cause analysis from three levels: macroscopic layout, microscopic state, and temporal logic. Once an anomaly is detected, the task dynamic planning and execution unit immediately intervenes, classifying and grading the anomaly, and autonomously planning, decomposing, and issuing ground-air collaborative tasks, instructing the robot platforms to perform verification or handling actions, forming a complete intelligent operation process from environmental perception, scene cognition, anomaly detection, and closed-loop action.

[0022] The core function of the multimodal data acquisition unit is to act as the system's sensing front end, simultaneously acquiring raw, high-dimensional sensing data of the warehouse environment from two complementary dimensions: air and ground. Physically, this unit consists of two independent sensing sub-units mounted on air and ground robot platforms respectively, ensuring data stream consistency in timestamps through a precise clock synchronization protocol.

[0023] The aerial perception subunit of this system is specifically mounted on a highly maneuverable quadcopter drone. This subunit integrates three core sensors. The first is a high-resolution wide-angle camera; in this embodiment, a camera module with 48 megapixels and a 120-degree field of view is used to capture true-color images of a wide-area scene. Its output video stream resolution is set to 4K, with a frame rate of 30 frames per second, providing rich texture and color details for subsequent scene understanding. The second is a 32-line mechanical rotating LiDAR, which performs a 360-degree horizontal scan at a rate of 10 revolutions per second, generating approximately 1.2 million three-dimensional spatial points per second, forming a wide-coverage, medium-density environmental point cloud data for rapid three-dimensional reconstruction of macroscopic structures. The third is an industrial-grade six-axis inertial measurement unit (IMU), which outputs the drone's three-axis acceleration and three-axis angular velocity data at a frequency of 200 Hz, providing high-frequency dynamic information for real-time estimation of motion attitude. All sensor data is tagged with a precise timestamp synchronized by the Global Navigation Satellite System during acquisition, and is initially packaged by the onboard processor and sent to the central computing cluster via a wireless link in a custom data frame format.

[0024] The ground perception subunit of this module is mounted on an autonomous mobile robot capable of navigating narrow aisles between shelves. This subunit is specifically configured for close-range, high-precision observation tasks and integrates three core sensors. The first is a high-magnification gimbal camera with 30x optical zoom, capable of clearly capturing tiny labels and damage details on cargo packaging from several meters away; its output image data is used for fine identification of microscopic conditions. The second is a structured light scanner based on the principle of striped grating projection; by projecting coded light patterns onto the surface of a target object and analyzing its deformation, it can reconstruct the local three-dimensional shape of the target object with sub-millimeter precision, used for high-precision modeling of specific goods or equipment. The third is a long-wave infrared thermal imaging camera, operating in the 8-14 micrometer band with a temperature resolution of 0.05 degrees Celsius, used for non-contact acquisition of temperature distribution maps of object surfaces to detect potential thermal anomalies such as equipment overheating or liquid leaks. Data collected by the ground platform is also timestamped, formatted, and uploaded in real time.

[0025] The ground-to-air collaborative positioning and mapping unit is tightly coupled with the data output of the multimodal data acquisition unit. Its core responsibility is to solve the problem of high-precision and robust positioning of the two mobile platforms in a unified coordinate system, and to simultaneously construct a global 3D environment map that integrates the macroscopic global structure and microscopic local details. The internal workflow of this unit is decomposed into four closely linked processing stages: global map construction, global prior-assisted ground positioning, global anchor-constrained aerial positioning, and final map fusion and optimization.

[0026] First, this unit utilizes wide-area color images and LiDAR point cloud data collected by the aerial perception subunit to calculate the aerial robot's trajectory in real time using a tightly coupled vision-LiDAR fusion odometry method, and initially generates a global sparse feature map. This sparse map consists of a series of points with stable geometric features in three-dimensional space. Each feature point is accompanied by its three-dimensional world coordinates and descriptors extracted from the image, such as fast robust features or orientation-fast and rotation-invariant feature descriptors, which have good invariance to changes in viewpoint and illumination.

[0027] Subsequently, this global sparse feature map is transmitted in real time to the ground robot platform as a global pose prior for visual odometry calculation. When the ground robot is locating itself in its local environment, it actively matches the visual features in its camera's field of view with the received global sparse feature map. Once a match is successful, the known world coordinates of these global feature points act like a lighthouse in navigation, directly correcting the ground robot's current pose estimation. This effectively constrains the drift caused by the accumulation of positioning errors over time, ensuring that the ground robot's activities are always anchored within a globally unified coordinate system.

[0028] Meanwhile, when performing close-range observation tasks, the ground perception subunit utilizes its structured light scanner to perform high-precision 3D measurements of static landmarks in the environment that remain stable over long periods, such as load-bearing columns of shelves, fixed fire hydrants in corners, or permanent marking lines sprayed on the ground. The 3D coordinates of these landmarks have extremely high accuracy in the local coordinate system of the ground robot itself. Using the known global pose of the ground robot, these high-precision local coordinates are transformed into the global world coordinate system and fed back as global anchor points to the central optimizer of the ground-air cooperative localization and mapping unit.

[0029] Finally, the central optimizer of this unit periodically executes a global pose graph optimization process. This pose graph is a complex network structure where nodes represent the poses of the ground and air robots at different times, and the 3D positions of all feature points on the map. The edges in the graph represent spatial constraints imposed by various sensor observations. These constraints include pre-integration constraints from the aerial robot's inertial measurement module, LiDAR point cloud registration constraints, visual reprojection error constraints, visual odometry constraints from the ground robot, and crucially, ground-to-air pose prior constraints provided by the sparse map and ground-to-air global anchor point constraints provided by precise measurements. This unit seeks the optimal solution that minimizes the sum of all constraint errors—that is, the best estimate of all robot poses and map point positions—by solving a nonlinear least squares problem. The objective function of this optimization problem can be formally expressed as:

[0030] In the above formula, This represents the set of pose states of all ground-to-air robots to be optimized. No. Robot pose at any given moment This represents the set of coordinates of all map points to be optimized. No. A map feature point. It is the pre-integration error of the inertial measurement module between adjacent poses of the aerial robot. and These are the reprojection errors of the point cloud from the LiDAR of the aerial robot and the map points from visual observation. It is the visual odometry error between adjacent poses of a ground robot. This refers to the global anchor point between the aerial robot's pose and the ground platform. The error term between them is a key constraint for achieving air-ground coordination. The covariance matrix, representing the errors of each item, is used to weight constraints from different sources. It is a robust kernel function, such as the Hubel function, used to reduce the negative impact of outliers, i.e., the negative impact of outliers on the optimization results. , , and By minimizing this joint cost function, different types of constraint sets (IMU, sensing, VO, anchor points) achieve a unified global consistency and high accuracy between the robot trajectory calculated by the system and the 3D map. The optimized map is ultimately output as a dense 3D color point cloud that integrates macroscopic structure and microscopic details, and is continuously updated in real time.

[0031] Furthermore, the core task of the unified semantic cognition model unit is to receive the high-precision fused 3D map output by the ground-air cooperative positioning and mapping unit, along with various image data associated with it and bearing precise pose labels. It then transforms this raw, purely geometric and visual information into structured semantic knowledge that can be directly understood and reasoned about by machines. The implementation of this unit relies on a deep learning-driven, multi-task, multi-modal scene understanding model deployed on the graphics processing unit array of a central computing cluster.

[0032] The workflow of this unit begins with instance segmentation of the fused point cloud map. Unlike simply classifying point clouds into different categories, instance segmentation aims to identify and accurately separate each individual physical entity in the warehouse environment. In this embodiment, a model architecture based on sparse convolutional neural networks is employed, which can efficiently process large-scale, non-uniformly distributed point cloud data. The model is pre-trained on massive amounts of warehouse scene point cloud data containing point-by-point instance labels, thereby learning how to independently segment point cloud clusters representing shelves, pallets, boxes, forklifts, robots, and even workers based on clues such as geometry and spatial proximity, and assigning a unique instance identifier to each entity.

[0033] After instance segmentation, the model enters the attribute extraction and state assessment phase. For each segmented entity instance, the system retrieves all image slices of different viewpoints and resolutions observed from the image database of the multimodal data acquisition unit, based on its 3D spatial location and recorded camera pose. These image slices, along with the instance's point cloud data, are fed into a multimodal feature fusion module. This module combines the entity's 3D geometric features and 2D texture features to perform multiple parallel recognition tasks. A classification subnetwork is responsible for identifying the entity's category, such as determining whether it is a "standard wooden pallet" or a "paper packaging box." A state assessment subnetwork is specifically used to determine the entity's current state, for example, by analyzing whether the texture of the packaging box surface is continuous and whether there are abnormal stains, to determine whether its state is "intact," "damaged," or "suspected of leaking." For specific types of equipment, such as motors or refrigeration units, the system extracts the corresponding area in the thermal imaging image, analyzes its average and maximum temperatures, and assesses whether its operating status is "normal" or "temperature abnormal." In addition, a dedicated optical character recognition and barcode decoding subnetwork operates on high-definition image slices to automatically extract and identify barcodes, QR codes, or text label information on cargo boxes, and associates and verifies them with inventory data in the warehouse management system.

[0034] Finally, this unit integrates all the parsed semantic information and outputs a dynamically updated digital twin scene diagram of the warehouse environment. This is a highly structured, object-oriented data structure. In this diagram, each node represents an identified physical entity, and the node's attributes fully record all the information of that entity: a unique instance identifier, category label, 3D bounding box, accurate 3D model or point cloud data, current state assessment result, and all extracted attribute information, such as barcode content and temperature value. The edges in the diagram are used to represent the rich spatial and logical relationships between entities. Spatial relationships are calculated by analyzing the relative positions of the 3D bounding boxes of each entity. For example, a "support" relationship is established when the bottom of the bounding box of a cargo box is in close contact with the top of the bounding box of a pallet; a "stack" relationship is established between multiple cargo boxes; and a "proximity" relationship is determined based on a distance threshold. Logical relationships are more complex. For example, by associating barcode information with order data, a "belonging" relationship can be established between a pallet and multiple cargo boxes. This scene is not static, but dynamically updated at a near real-time frequency as the ground-to-air robot platform continuously senses it, fully reflecting the instantaneous state of the warehousing environment at both the physical and informational levels.

[0035] The multi-scale anomaly detection and reasoning unit is the core of the system's intelligent analysis. Its function is to perform comprehensive and in-depth anomaly detection and root cause analysis based on the dynamic scene graph output by the unified semantic cognitive model unit. This unit abandons the traditional simple alarm mode based on a single threshold, and instead constructs a hierarchical and progressive anomaly detection logic from three interrelated dimensions: macroscopic, microscopic, and temporal.

[0036] The first layer of the detection logic is macro-layout anomaly detection. This unit pre-stores a standard warehouse layout model, also in the form of a scene graph. This model defines the standard positions of all fixed facilities, such as shelves and workbenches, as well as the precise three-dimensional spatial range of all functional areas, such as main aisles, fire exits, and temporary storage areas. The macro-layout anomaly detection module performs periodic graph matching and spatial geometry comparison between the real-time generated scene graph and this standard layout model. During the comparison, if a significant positional or orientation deviation of a shelf node is detected in the real-time scene, it may indicate that the shelf is tilting or collapsing. If any entity node representing a "pallet" or "goods" is detected whose three-dimensional bounding box overlaps with the spatial area representing a "fire exit," the system will immediately determine it as an "illegally occupied fire exit" anomaly. Similarly, if a large quantity of goods that should be stored in area A is found to have its node appearing in the scene graph of area B, an "official goods misplacement" alarm will be triggered.

[0037] The second layer of the detection logic is micro-level anomaly detection. This layer focuses on the internal attributes of individual entity nodes in the scene graph. This is a rule-based, fine-grained inspection process. The inference engine traverses every node in the scene graph and applies corresponding detection rules based on its category. For example, for all nodes categorized as "cargo boxes," the system checks their "status" attribute. If the attribute value is "damaged" or "leaking," a "cargo packaging damage" anomaly is triggered. For nodes categorized as "liquid containers," the system also checks for newly appearing unidentified geometric or color areas on the surrounding ground to cross-verify the possibility of leakage. For nodes categorized as "electrical equipment," the system continuously monitors their "temperature" attribute; once this value exceeds a preset safety threshold, such as 85 degrees Celsius, an "abnormal equipment surface temperature" alarm is generated. Furthermore, for all cargo nodes that should have labels, if their "label information" attribute is empty or fails to resolve, it is determined as a "label information cannot be recognized" anomaly.

[0038] The third layer of the detection logic is temporal logic anomaly reasoning, which is the most in-depth function of this unit. This unit maintains a historical state database, recording the complete sequence of state evolution of all key entities in the scene graph over a past period. Simultaneously, the system imports all standard operational process models from the upper-level warehouse management system, which are formalized as finite state machines or time-constrained Petri nets. The core task of the temporal logic anomaly reasoning module is to perform pattern matching and logical consistency checks on the actual state evolution sequences of entities observed in the real world against predefined standard operational process models. For example, a standard process might stipulate that goods entering the temporary storage area must be transferred to their permanent storage location within two hours. This module continuously tracks all goods nodes in the temporary storage area and starts a timer for them. If a goods node remains in the area for more than two hours, even if its own state is intact, the system will infer that this is a "goods overdue" process anomaly. Similarly, by analyzing the historical trajectory node sequence of a ground robot, if its path significantly deviates from the predetermined path specified by the task planning system, a "device movement trajectory deviation" anomaly will be triggered. Furthermore, if a node categorized as "staff member" is detected in a scene map outside of working hours or in an unauthorized area, the system can infer the potential security risk of "unauthorized personnel activity".

[0039] The task dynamic planning and execution unit, serving as the decision-making and action center of the entire system, receives and responds to various anomaly alarms generated by the multi-scale anomaly detection and reasoning unit, and autonomously transforms these alarms into a sequence of executable, coordinated, and specific task instructions for the ground-to-air robot. This unit's operational mechanism ensures a rapid, closed-loop response from anomaly detection to actual intervention.

[0040] Upon receiving any type of abnormal alarm signal, the unit's internal anomaly classification and priority assessment engine is activated first. This engine, based on a predefined, administrator-configurable rule base, quickly classifies the severity level and root cause type of the abnormal event. For example, "fire lane obstruction" is classified as "Level 1, Serious Safety Issue"; "damaged cargo packaging" is classified as "Level 2, Asset Loss"; and "cargo delays" is classified as "Level 3, Inefficiency." This classification result directly determines the urgency of subsequent response tasks and resource allocation strategies.

[0041] Based on the classification results, the collaborative task generation module of this unit automatically selects and instantiates the most suitable ground-air collaborative verification or handling task from the task template library. The task template is a pre-defined behavioral strategy framework for different anomaly types. Taking a "suspected ground liquid leak" anomaly initially detected from an aerial perspective as an example, the system will select a task template named "Ground-Air Collaborative Detailed Investigation." After instantiating this template, the system will automatically generate a high-level task containing multiple sub-objectives: confirming the leak source, assessing the leak range and substances, and recording on-site image evidence.

[0042] Subsequently, the task decomposition module further breaks down this high-level collaborative task into a series of specific, quantifiable action instructions for both the aerial and ground robotic platforms. For the aforementioned leak detection task, the decomposition result might be: for the aerial robotic platform, the generated instruction sequence includes "fly to a position 5 meters above the target area and hover," "adjust the gimbal camera's tilt angle to 75 degrees," and "perform 360-degree video recording around the target point," with the aim of secondary confirmation and monitoring of the surrounding environment from multiple perspectives. For the ground robotic platform, the generated instruction sequence is more complex, including "planning a collision-free path to the target point, avoiding potentially slippery areas," "upon arrival, using a high-magnification zoom camera to capture close-up images of the leak source," "switching to a thermal imaging camera to scan the temperature difference between the leaking liquid and the surrounding ground," and "packaging and uploading all detection data and marking the task as complete."

[0043] Ultimately, these decomposed atomic-level task instructions are encoded by the task instruction conversion and distribution module into standardized motion control commands that can be understood by the respective underlying control systems of the two robots, such as messages conforming to the communication protocols of specific robot operating systems. These commands are precisely distributed to the corresponding robot platforms for execution via wireless networks. During execution, the robots continuously report their status and task progress, thus forming a complete monitoring and feedback loop within the task dynamic planning and execution unit, ensuring seamless integration of the entire system from passive perception, intelligent cognition, and proactive detection to active closed-loop action.

[0044] Example 2 In this embodiment, the training strategy of the deep learning model adopted by the unified semantic cognition model unit is specified and described in detail. The core of this strategy is that the model is not only trained on real-world data, but also jointly trained on a hybrid dataset consisting of large-scale, highly diverse synthetic data and continuously collected, finely labeled real data, in order to simultaneously obtain excellent generalization ability and high adaptability to specific scenarios.

[0045] The generation process for the synthetic data portion is a systematic workflow based on physically based rendering and procedural content creation. First, using the warehouse's building information model or 3D design model, a basic 3D digital twin environment identical to the real warehouse is constructed in professional digital content creation tools such as Blender or 3ds Max. Then, this environment is imported into a real-time rendering engine with high-fidelity physically based rendering capabilities, such as Unreal Engine 5. Within this engine, a procedural content generation script is developed and executed. This script is responsible for automatically creating an endless stream of training samples that are difficult to obtain in batches in the real world. Specifically, the script can: randomize lighting conditions, simulating various lighting changes from sunrise to sunset, sunny to cloudy days, and emergency lighting; programmatically stack and arrange different types and packages of goods on shelves, simulating their natural drooping and tilting postures; and most importantly, the script can actively generate various rare anomaly samples based on a preset probability model. For example, by performing vertex displacement and texture map replacement on the 3D model of the cargo box, it can simulate different degrees of compression, damage, water stains, and oil stains; and by generating a particle system that simulates fluid physics on the ground, it can create liquid leakage scenarios of various shapes. For each generated scene, the rendering engine can simultaneously output multiple types of data: simulated color images and simulated depth images that conform to the characteristics of ground and air robot sensors, and most importantly, perfect pixel-by-pixel semantic labels and instance segmentation masks. These automatically generated, pixel-level accurate label data greatly reduce the cost of data annotation and enable the model to learn robust knowledge to various anomaly patterns at an early stage.

[0046] The real-world data portion is acquired and utilized through a closed-loop process of continuous learning after the system is actually deployed in the target warehouse. All raw data collected by the ground-to-air robots during daily inspections is stored. The system uses the current version of the Unified Semantic Cognitive Model Unit to perform initial automatic annotation on this data. Then, this pre-annotated data is submitted to a human review and annotation platform. Human operators on this platform do not annotate from scratch but check and correct the machine's annotation results, which greatly improves annotation efficiency. High-quality annotated data, after fine-tuning by humans, is added to the real-world database. This database is used for periodic fine-tuning and optimization of the model, ensuring that the model can continuously learn and adapt to the unique characteristics of specific warehouse environments, such as special cargo types, unique lighting conditions, or newly introduced equipment.

[0047] The hybrid training strategy is specifically manifested in the model training phase as a specially designed loss function. This loss function is a weighted sum of the synthetic data loss and the real data loss, and its form is as follows:

[0048] In this formula, It is the total loss that needs to be minimized. This represents all trainable parameters of a deep learning model. and These represent synthetic data batches and real data batches, respectively. and It is the loss value calculated by the model on the corresponding batch of data. This loss value is itself a composite function, which usually includes focus loss or Dice loss for instance segmentation, and cross-entropy loss for class recognition and state evaluation. It is a key hyperparameter, a weighting factor between 0 and 1, used to balance the contributions of synthetic and real data during training. In the early stages of training... The value is set relatively high, such as 0.8, so that the model can quickly learn general, robust feature representations from large-scale synthetic data. As training progresses, The value of is gradually reduced through a preset annealing strategy. This means that the model will increasingly focus on learning the fine features of specific scenes from high-quality real data, thus completing the transition from generalization to specialization. This hybrid training strategy ensures that the final model will neither overfit to the "virtual" feel of synthetic data nor suffer from insufficient generalization ability due to the scarcity and limitations of real data. Thus, while ensuring recognition accuracy, it also possesses a strong ability to cope with complex and ever-changing real-world environments.

[0049] Example 3 In this embodiment, a bidirectional information feedback channel is constructed between the ground-air cooperative localization and mapping unit and the task dynamic planning and execution unit. This channel realizes a dynamic, on-demand, and intelligent collaborative mechanism, significantly improving the robot's autonomous navigation and task execution success rate in unknown or dynamically changing environments.

[0050] The trigger condition for this feedback channel stems from the autonomous judgment made by the ground robot during task execution. When the task dynamic planning and execution unit issues a path planning instruction to the ground robot, the ground robot's local navigation system calculates an optimal path based on the global map it currently holds, provided by the ground-air cooperative positioning and mapping unit. If, during its calculation process, its path planning algorithm, such as the A algorithm or the RRT algorithm, repeatedly fails in a certain area and cannot find a collision-free path to the target, or if, during its movement, its forward-mounted obstacle avoidance LiDAR or depth camera detects an obstacle not marked on the global map, causing the path to be blocked, then the ground robot will determine that it is facing a "path obstruction due to incomplete or outdated local map information" event.

[0051] Once the event is triggered, the ground robot will immediately initiate a "map update request" to the task dynamic planning and execution unit through this two-way information feedback channel. This request is encapsulated in a standardized data packet, which includes at least: the unique identifier of the requesting ground robot, its current precise global pose, and a clearly defined 3D bounding box of the target area requiring supplementary surveying. The extent of this bounding box is typically determined based on the area where its path planning failed or the surrounding area of ​​unknown obstacles detected by local sensors.

[0052] The task scheduler within the task dynamic planning and execution unit treats this "map update request" as a high-priority event. It immediately assesses the aerial robot's current task status. If the aerial robot is idle, a new "on-demand survey" task is generated directly for it. If the aerial robot is performing a low-priority task, such as a routine full-coverage inspection, the scheduler will interrupt its current task and insert the "on-demand survey" task at the front of its task queue. The specific content of the new survey task is to instruct the aerial robot to fly over the target area requested by the ground robot.

[0053] After receiving instructions, the aerial robot autonomously plans its path to the target area and conducts a low-speed, multi-angle, detailed scan of the area, using its onboard high-resolution camera and LiDAR to capture the latest, high-density data. This newly acquired data is transmitted back to the central computing cluster in real time with high-priority tags.

[0054] Upon receiving these incremental data with high-priority labels, the ground-air cooperative positioning and mapping unit immediately initiates a real-time local map update process. Unlike the time-consuming global optimization, this process only performs rapid map updates on the requested survey area and its neighboring areas. It registers and merges the new point cloud data with the old map data, removes parts of the old map that conflict with the new observations (e.g., a moved cargo container), and adds newly appearing elements (e.g., a temporarily placed maintenance device). This updated local map patch is then seamlessly integrated back into the global map.

[0055] Once the local map update is complete, the ground-air cooperative localization and mapping unit sends a "map update complete" confirmation signal to the task dynamic planning and execution unit via an information channel. The task dynamic planning and execution unit then instructs the previously obstructed ground robot to re-acquire the updated map information from the ground-air cooperative localization and mapping unit and replan its path based on the new, more accurate environmental understanding. Through this dynamic closed-loop collaboration of "obstruction-request-survey-update-replanning," the system can gracefully handle dynamic environmental changes, ensuring the continuity and success rate of the ground robot's mission execution in complex environments. This elevates the relationship between the ground and air platforms from a simple "master-slave" relationship to a symbiotic relationship of intelligent collaboration and dynamic complementarity.

Claims

1. A ground-to-air robot intelligent perception and anomaly detection system for warehouse environment based on AI vision, characterized in that, include: A multimodal data acquisition unit is configured to simultaneously acquire raw perception data of the warehouse environment from both aerial and ground dimensions. The multimodal data acquisition unit includes an aerial perception subunit mounted on an aerial robot platform and a ground perception subunit mounted on a ground robot platform. The air-ground cooperative positioning and mapping unit is connected to the multimodal data acquisition unit and is used to fuse the wide-area data collected by the air perception subunit and the fine data collected by the ground perception subunit to construct a globally unified three-dimensional environment map and perform joint positioning of the air and ground robot platforms. The unified semantic cognition model unit is used to receive the three-dimensional environment map and associated image data generated by the ground-air cooperative positioning and mapping unit, and to transform the raw geometric and visual information into structured semantic knowledge through a deep learning-driven scene understanding model. The multi-scale anomaly detection and reasoning unit is used to perform macro-layout anomaly detection, micro-state anomaly detection and temporal logic anomaly reasoning based on the dynamic scene graph output by the unified semantic cognition model unit, so as to identify spatial structure anomalies, entity state anomalies and process anomalies. The task dynamic planning and execution unit is used to respond to the abnormal alarm signal generated by the multi-scale abnormal detection and reasoning unit, classify abnormal events, and autonomously generate a ground-air collaborative task that is decomposed into specific action instructions for the air and ground robot platforms, so as to form a closed loop of perception, cognition, detection and action.

2. The system according to claim 1, characterized in that, The airborne sensing subunit integrates a high-resolution wide-angle camera, a lidar, and an inertial measurement module to capture color image information, 3D point cloud data, and its own motion attitude data of a large-scale scene; the ground sensing subunit integrates a high-magnification zoom camera, a structured light scanner, and a thermal imaging camera to acquire high-definition texture details, microscopic 3D morphology data, and temperature distribution information of specific targets.

3. The system according to claim 1, characterized in that, A global sparse feature map is generated using the wide-area data. This global sparse feature map is transmitted to the ground robot platform as a global pose prior to constrain its positioning error. The precise world coordinates fed back by the ground perception subunit after performing high-precision 3D measurements of static features in the environment are used as global anchor points to correct the trajectory drift of the aerial robot platform. The ground-air cooperative positioning and mapping unit calculates the motion trajectory of the aerial robot platform in real time through a tightly coupled vision-lidar fusion odometry method and generates the global sparse feature map. When performing visual odometry calculation, the ground robot platform matches the visual features in its camera field of view with the global sparse feature map to use the global pose prior to correct its current pose estimate.

4. The system according to claim 3, characterized in that, The ground perception subunit uses the structured light scanner to perform high-precision three-dimensional measurements on static landmarks that remain stable in the environment over a long period of time; the ground-air cooperative positioning and mapping unit uses the global anchor points converted from the static landmarks to constrain the pose of the aerial robot platform during the global pose map optimization process.

5. The system according to claim 4, characterized in that, The ground-air cooperative localization and mapping unit performs global positioning map optimization by solving a nonlinear least squares problem. The goal of this problem is to minimize a joint cost function, which integrates multiple spatial constraints, including the pre-integration error of the inertial measurement module between adjacent poses of the aerial robot, the reprojection error of the aerial robot's lidar point cloud and visual observation to map points, the visual odometry error between adjacent poses of the ground robot, and the error term between the aerial robot's pose and the global anchor point provided by the ground platform.

6. The system according to claim 1, characterized in that, The 3D environment map is segmented to separate independent physical entities. Image texture information is then used to perform category recognition, state assessment, and attribute extraction for each entity. The final output is a dynamically updated digital twin scene map of the warehouse environment, where nodes represent physical entities and their semantic attributes, and edges represent the spatial and logical relationships between entities. The instance segmentation performed by the unified semantic cognitive model unit is used to independently segment point cloud clusters representing shelves, pallets, cargo boxes, robots, and workers based on geometric shapes and spatial proximity cues, and assign each cluster a unique instance identifier. The category recognition, state assessment, and attribute extraction are used to determine entity categories, assess whether entity packaging is intact or damaged, evaluate equipment operating temperature, and identify barcodes, QR codes, or text labels on cargo boxes.

7. The system according to claim 6, characterized in that, The unified semantic cognition model unit uses a deep learning model that is trained on a hybrid dataset consisting of synthetic and real data. The synthetic data is obtained by physically rendering and procedurally generating the 3D design model of the warehouse. The real data is continuously collected in the actual deployment environment of the system and manually corrected and labeled. The training process is accomplished by minimizing a total loss function consisting of a weighted sum of the synthetic data loss and the real data loss.

8. The system according to claim 1, characterized in that, The specific working logic of the multi-scale anomaly detection and reasoning unit is as follows: the macro-layout anomaly detection compares the real-time scene map with the preset standard warehouse layout model to detect collapsed shelves or blocked fire exits; the micro-state anomaly detection analyzes the attributes of individual entity nodes in the scene map to identify damaged goods packaging or abnormal equipment surface temperature; the temporal logic anomaly reasoning compares the actual sequence of entity state evolution over time with the predefined standard operating procedure model to infer that goods have been delayed for too long or that the equipment movement trajectory has deviated.

9. The system according to claim 1, characterized in that, The specific workflow of the task dynamic planning and execution unit is as follows: classify the received abnormal alarm signals by severity level and type based on a predefined rule base; select and instantiate a ground-air collaborative verification or handling task from the task template library based on the classification results; decompose the ground-air collaborative task into a hovering observation command sequence for the aerial robot platform and a path planning and close-range reconnaissance command sequence for the ground robot platform; and convert the command sequence into motion control commands that the robot can execute and issue them.

10. The system according to claim 1, characterized in that, A two-way information feedback channel is also established between the ground-air cooperative positioning and mapping unit and the task dynamic planning and execution unit. When the ground robot's path is blocked due to incomplete local map information, it initiates a map update request to the task dynamic planning and execution unit through the two-way information feedback channel. In response to the request, the task dynamic planning and execution unit instructs the aerial robot to fly to the target area for supplementary surveying, thereby triggering the ground-air cooperative positioning and mapping unit to update the local environmental map in real time.