Multi-sensor data fusion robot real environment perception warning system

By using a multi-sensor data fusion system to generate a dynamic semantic map of the environment, the shortcomings of multi-source heterogeneous data processing in robot environmental perception systems are solved, enabling efficient and accurate environmental perception and autonomous early warning decision-making.

CN121315951BActive Publication Date: 2026-05-01SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)
Filing Date
2025-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing robot environmental perception systems lack a unified and efficient layered fusion architecture, making it impossible to systematically process multi-source heterogeneous data, resulting in insufficient perception accuracy and robustness.

Method used

A multi-sensor data fusion system is adopted, including a multi-sensor feedback module, a cross-validation module, a data processing definition module, a heterogeneous feature fusion module, and an early warning decision module. Through time series and coordinate unification, cross-validation, feature extraction, and deep fusion, a dynamic semantic map of the environment is generated for hierarchical early warning decision-making.

Benefits of technology

It significantly improves the comprehensiveness and accuracy of environmental perception, reduces false alarms and missed alarms, provides reliable assurance for autonomous robot operation, adapts to a variety of complex scenarios, and realizes intelligent decision-making from raw perception to hierarchical early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121315951B_ABST
    Figure CN121315951B_ABST
Patent Text Reader

Abstract

The present application relates to the field of robot environment perception, and is used for solving the problem that the existing robot environment perception system lacks a unified and efficient hierarchical fusion architecture to systematically process multi-source heterogeneous data, in particular to a robot real environment perception early warning system for multi-sensor data fusion; the present application significantly improves the comprehensiveness and accuracy of environment perception through the cooperative work of laser radar, visual camera, millimeter wave radar and ultrasonic sensor; through the progressive processing flow of space-time registration, cross-checking, feature definition to deep fusion, the multi-source information is sorted by confidence and features are extracted, and the heterogeneous feature fusion module generates an environment dynamic semantic map rich in texture, structure, motion and semantic information by means of convolutional neural network, which provides accurate and comprehensive information basis for early warning decision, and at the same time, the dynamic hierarchical early warning mechanism provides reliable guarantee for robot safety and autonomous operation.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-sensor data fusion robot real-world perception and early warning system Technical Field

[0001] This invention relates to the field of robot environmental perception, specifically to a robot real-world environment perception and early warning system based on multi-sensor data fusion. Background Technology

[0002] With the widespread application of robotics technology in industrial automation, smart warehousing, autonomous driving, and public services, accurate and robust perception of the real-world environment has become a core prerequisite for ensuring the robot's autonomous and safe operation. The performance of the environmental perception system directly determines the robot's decision-making level and operational reliability.

[0003] Early systems typically relied on a single type of sensor, such as a vision camera, LiDAR, or millimeter-wave radar. While these solutions were simple in structure, they had inherent functional limitations. Therefore, to compensate for the shortcomings of a single sensor, subsequent technologies adopted solutions that combined multiple sensors. For example, cameras and LiDAR were installed on a robot simultaneously. However, most systems at this stage employed simple data switching or decision-level fusion strategies, i.e., selecting from the outputs of different sensors based on the scene, or only performing high-level comparisons of results after each sensor's independent perception, failing to achieve deep information complementarity.

[0004] In existing technologies, more advanced techniques are beginning to attempt data fusion, such as fusing point clouds from LiDAR with images from cameras to improve the accuracy of target detection. However, these fusion methods often remain at a simple combination of data or feature layers, lacking a unified and efficient layered fusion architecture to systematically process multi-source heterogeneous data.

[0005] To address the aforementioned technical problems, this application proposes a solution. Summary of the Invention

[0006] This invention significantly improves the comprehensiveness, accuracy, and robustness of environmental perception through the complementary fusion of multiple sensor sources. It achieves layer-by-layer abstraction and decision optimization from raw data to high-level semantics, effectively reducing false alarms and missed alarms. Its dynamic hierarchical early warning mechanism can adapt to various complex scenarios, from structured factories to unstructured environments, providing reliable protection for robot safety and autonomous operation. It solves the problem that existing robot environmental perception systems lack a unified and efficient hierarchical fusion architecture to systematically process multi-source heterogeneous data, and proposes a robot real-world environmental perception and early warning system based on multi-sensor data fusion.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] A robot real-world environment perception and early warning system based on multi-sensor data fusion includes a multi-sensor feedback module. The multi-sensor feedback module is used to receive signals fed back from multiple sensors, and to perform digital processing on the feedback signals, unify the time sequence and coordinates, and obtain multi-source heterogeneous data.

[0009] The cross-validation module can acquire multi-source heterogeneous data through the multi-sensor feedback module, and perform preliminary model generation on each individual multi-source heterogeneous data to obtain multiple single semantic maps. The cross-validation module selects different semantic maps for cross-validation and assigns a weight vector to each semantic map.

[0010] The data processing definition module acquires multi-source heterogeneous data through the multi-sensor feedback module, extracts data features from the multi-source heterogeneous data to generate various heterogeneous features, and obtains weight vectors through the cross-validation module. Based on the weight vectors, the various heterogeneous features are corrected to obtain unified feature layer feature data.

[0011] The heterogeneous feature fusion module obtains feature data through the data processing definition module, and fuses the feature data with depth, contour, speed and semantic features to obtain a dynamic semantic map of the environment.

[0012] The early warning decision module divides regions based on the dynamic semantic map of the environment and classifies the early warning level of each region in combination with its own movement characteristics, thereby generating an early warning region decision.

[0013] In a preferred embodiment of the present invention, the multi-sensor feedback module is connected to the binocular vision unit, the laser scanning unit, the radar detection unit, and the inertial measurement unit, and receives the feedback depth vision image, laser scanning image, 3D radar point cloud image, and motion inertial results. It performs time synchronization and spatial registration of the multi-sensor results through hardware synchronization and software synchronization. After time synchronization and spatial registration, the data from multiple sensors are recorded as multi-source heterogeneous data.

[0014] In a preferred embodiment of the present invention, the single semantic map obtained by the cross-validation module includes a visual model, a laser model, a 3D model, and a motion model. Specifically, a visual model is generated from a depth visual image, a laser model is generated from a laser scan image, a 3D model is generated from a radar echo image, and a motion model is generated from motion inertia results.

[0015] In a preferred embodiment of the present invention, the cross-validation module selects a laser model and a 3D model in the same time sequence. By mapping radar points in the 3D model to the laser model, and obtaining clusters of differences through cluster analysis and calculating the volume of the clusters of differences, the module then removes differences with excessively small volumes through volume filtering. Finally, the module obtains the difference regions in the laser model and the 3D model, and performs analysis based on the spatial distribution of the differences to achieve the detection of specific types of targets.

[0016] The cross-validation module records the regions shared by the laser model and the 3D model as high-weight regions, and the regions unique to the laser model and the 3D model as normal-weight regions.

[0017] In a preferred embodiment of the present invention, the cross-validation module compares the visual model with the laser model and the 3D model respectively, and verifies the unique regions. If a unique region in the laser model exists in the visual model, it is upgraded to a high-weight region; if a unique region in the laser model does not exist in the visual model, it is downgraded to a low-weight region. If a unique region in the 3D model exists in the visual model, it is upgraded to a high-weight region; if a unique region in the 3D model does not exist in the visual model, it is downgraded to a low-weight region.

[0018] In a preferred embodiment of the present invention, if an object exists in both the laser model and the 3D model but there are differences, it is recorded as an abnormal region. The laser model and the 3D model are compared with the visual model respectively to obtain the overlap degree of the abnormal regions in the visual model and the laser model, as well as the overlap degree of the abnormal regions in the visual model and the 3D model. The two sets of overlap degrees are compared, and the set with a higher overlap degree of the abnormal regions in the 3D model and the laser model is used as a high weight factor, and the set with a lower overlap degree is used as a low weight factor.

[0019] In a preferred embodiment of the present invention, after acquiring multi-source heterogeneous data, the data processing definition module obtains accurate surface texture and structural model through laser model, obtains moving objects in the environment through 3D model, and obtains whether the objects are in motion and their structural model.

[0020] The data processing definition module then uses a deep learning model to perform object recognition on the visual model to obtain the object semantics in the visual model.

[0021] The data processing definition module sorts the confidence levels of the high-weight and low-weight regions defined by the cross-validation module, generates surface textures, structural models, motion states, and object semantics with different weights, and records them as feature data.

[0022] After acquiring feature data, the heterogeneous feature fusion module performs deep fusion of the feature data through a convolutional neural network, ultimately generating a set of dynamic semantic maps of the environment that include texture information, motion information, structural information, and semantic information.

[0023] In a preferred embodiment of the present invention, the early warning decision module divides the robot body into a near-field area and a far-field area layer by layer outward from the center point. When dividing, the early warning decision module first obtains the set path for the robot's movement, then obtains the motion inertia result through the multi-sensor feedback module, and uses the path forward direction and the motion direction of the current motion inertia result as weight vectors to offset the near-field area graphic, forming the deformed near-field area and far-field area.

[0024] After dividing the near-field and far-field areas, the early warning decision module uses a dynamic semantic map of the environment to conduct risk warnings, identify objects in the near-field and far-field areas that pose a collision risk with the robot's set path, and determine the loss result after the collision based on semantics. The collision risk level is calculated by multiplying the collision risk and the result at any time, and the collision risk level is used as the warning level. The warning is then issued through the network.

[0025] Compared with the prior art, the beneficial effects of the present invention are:

[0026] 1. In this invention, through the collaborative work of LiDAR, visual camera, millimeter-wave radar and ultrasonic sensor, the invention achieves the complementarity of multi-physical dimension information. The visual sensor provides rich texture and semantics, the LiDAR provides precise geometric structure, the millimeter-wave radar provides accurate motion information, and the ultrasonic sensor is responsible for fine detection at close range, fundamentally eliminating the perception blind spots of a single sensor.

[0027] 2. In this invention, a progressive processing flow from spatiotemporal registration, cross-validation, feature definition to deep fusion is used to sort and extract features from multi-source information based on confidence. The heterogeneous feature fusion module uses a convolutional neural network to generate a dynamic semantic map of the environment rich in texture, structure, motion, and semantic information, providing an accurate and comprehensive information foundation for early warning decisions. Based on this map, combined with intelligent region division and dynamic risk assessment of robot motion vectors, the decision-making process from primitive perception to hierarchical early warning and autonomous obstacle avoidance is realized, improving the accuracy of early warning and the intelligence level of robot actions.

[0028] 3. In this invention, by using a weighted partitioning and dynamic region partitioning mechanism, attention is prioritized to the high-weight region after cross-validation, and vector deformation based on the motion direction is applied to the near-field region. This allows perception and computing resources to be focused on key threats on the robot's motion path, which not only improves real-time performance but also makes early warning and decision-making more in line with the dynamic needs of actual operation scenarios. Attached Figure Description

[0029] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0030] Figure 1 is a system block diagram of the present invention;

[0031] Figure 2 is a system flowchart of the present invention. Detailed Implementation

[0032] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0033] Example 1:

[0034] Please refer to Figures 1 and 2. The robot real-world perception and early warning system based on multi-sensor data fusion includes a multi-sensor feedback module, a cross-validation module, a data processing definition module, a heterogeneous feature fusion module, and an early warning decision module.

[0035] The multi-sensor feedback module is connected to the binocular vision unit, laser scanning unit, radar detection unit and inertial measurement unit, and receives the feedback depth vision image, laser scanning image, 3D radar point cloud image and motion inertial results.

[0036] After acquiring the results from multiple sensors, the multi-sensor feedback module performs time-series synchronization of the results through hardware and software synchronization. Specifically:

[0037] The hardware synchronization method involves sending a synchronization trigger signal to the sensors that can be synchronized through the hardware clock source equipped on the robot, and using all sensor data collected based on the synchronization trigger signal as unified time data for time synchronization.

[0038] When the sensor cannot access the hardware clock source or the synchronization trigger signal cannot be sent, the software inserts a high-precision timestamp into the data fed back by the sensor and aligns the data with the same high-precision timestamp.

[0039] After hardware and software synchronization, the multi-sensor feedback module verifies the data types corresponding to the selected time. If it simultaneously includes depth vision images, laser scan images, radar echo images, and motion inertial results, it determines that the sensor data for the selected time is complete. If it does not simultaneously include depth vision images, laser scan images, radar echo images, and motion inertial results, it determines that the sensor data for the selected time is missing. From the missing data types, it selects data with similar timestamps, which are two sets of data from the frame before and after the selected time. The state corresponding to the selected time is inferred through the motion model and used as supplementary data. The supplementary data is then added to complete the sensor data types for the selected time.

[0040] After acquiring data from different sensors, the multi-sensor feedback module performs coordinate transformation on the feedback data, unifies the coordinates of the various data, and completes spatial registration.

[0041] After performing time synchronization and spatial registration, the multi-sensor feedback module records the data from multiple sensors as multi-source heterogeneous data.

[0042] After acquiring multi-source heterogeneous data through the multi-sensor feedback module, the cross-validation module selects depth vision images to generate a visual model, selects laser scanning images to generate a laser model, selects radar echo images to generate a 3D model, and selects motion inertial results to generate a motion model. The visual model, laser model, motion model, and 3D model are all recorded as a single semantic map.

[0043] The cross-validation module selects laser models and 3D models from the same time sequence and compares them. By mapping radar points in the 3D model to the laser model, and performing cluster analysis using Euclidean clustering algorithm, clusters of differences are obtained. The volume of the difference clusters is calculated using the algorithm, and then small-volume difference points are removed by volume filtering. Finally, the difference regions in the laser model and 3D model are obtained, and analysis is performed based on the spatial distribution of the difference points to achieve the detection of specific types of targets or sensor anomalies. For example, if a region is unique to the 3D model but not to the laser model, it indicates that the object is not obvious to lasers and may be a light-transmitting or radar-absorbing material. If a region is unique to the laser model but not to the 3D model, it indicates that the material is not sensitive to radar waves. The cross-validation module records the regions shared by the laser model and 3D model as high-weight regions and the regions unique to the laser model and 3D model as normal-weight regions.

[0044] The cross-validation module then selects a visual model and compares it with the laser model and the 3D model. The visual model verifies the unique regions in the laser model and the 3D model. If a unique region in the laser model exists in the visual model, it is upgraded to a high-weight region; if a unique region in the laser model does not exist in the visual model, it is downgraded to a low-weight region. Similarly, if a unique region in the 3D model exists in the visual model, it is upgraded to a high-weight region; if a unique region in the 3D model does not exist in the visual model, it is downgraded to a low-weight region.

[0045] If an abnormal region exists in both the laser model and the 3D model but differs from it, it is recorded as an abnormal region. The laser model and the 3D model are then compared with the visual model to obtain the overlap degree of the abnormal regions in the visual model and the laser model, as well as the overlap degree of the abnormal regions in the visual model and the 3D model. The two sets of overlap degrees are compared, and the set with a higher overlap degree in the 3D model and the laser model is taken as a high-weight factor, while the set with a lower overlap degree is taken as a low-weight factor.

[0046] After acquiring multi-source heterogeneous data, the data processing definition module obtains accurate surface texture and structural model through laser model, and obtains moving objects in the environment through 3D model, and obtains whether the objects are in motion and their structural model.

[0047] The data processing definition module then uses a deep learning model to perform object recognition on the visual model, obtaining the object semantics in the visual model, namely object type, material, etc.

[0048] The data processing definition module sorts the confidence scores of the high-weight and low-weight regions defined by the cross-validation module, thereby generating surface textures, structural models, motion states, and object semantics with different weights and recording them as feature data.

[0049] After acquiring feature data, the heterogeneous feature fusion module performs deep fusion of the feature data through a convolutional neural network, and finally generates a set of dynamic semantic maps of the environment containing texture information, motion information, structural information and semantic information. The heterogeneous feature fusion module can display the dynamic semantic map of the environment through augmented reality interaction technology according to the preset settings, so that managers can perform human-computer interaction.

[0050] Example 2:

[0051] Please refer to Figures 1 and 2.

[0052] In a robot real-world environment perception and early warning system based on multi-sensor data fusion, the early warning decision module, after acquiring a dynamic semantic map of the environment, divides the area according to the dynamic semantic map. That is, it divides the area into near-field and far-field regions layer by layer outward from the robot body as the center point. When dividing the area, the early warning decision module first acquires the set path for the robot's movement, then obtains the motion inertia result through the multi-sensor feedback module, and uses the path forward direction and the current motion inertia result as weight vectors to offset the near-field region graphic, forming deformed near-field and far-field regions. For example, the near-field region is expanded from a circle to an ellipse. The major axis of the ellipse is determined by the path forward direction and the current motion direction. The robot body is located at a focus of the ellipse away from the motion direction, so that the division of the near-field region can fit the robot's own motion vector, expanding the near-field range in the forward direction and shrinking the near-field range in the far-field direction, making the division of the area more in line with the moving robot.

[0053] After dividing the near-field and far-field areas, the early warning decision module uses a dynamic semantic map of the environment to provide risk warnings. It identifies objects in the near-field and far-field areas that pose a collision risk with the robot's set path and determines the loss result after a collision based on semantics. The collision risk level is calculated by multiplying the collision risk and the result at any time. The collision risk level is then used as the warning level, which is issued via the network. At the same time, the robot's movement is automatically adjusted based on deep learning to perform autonomous obstacle avoidance, thereby preventing collisions during robot movement.

[0054] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A robot real-world environment perception and early warning system based on multi-sensor data fusion, characterized in that, It includes a multi-sensor feedback module, which is used to receive signals fed back from multiple sensors, and to perform digital processing on the fed-back signals, unify the time sequence and coordinates, and obtain multi-source heterogeneous data. The cross-validation module acquires multi-source heterogeneous data through the multi-sensor feedback module, performs preliminary model generation on each type of individual multi-source heterogeneous data, and obtains multiple single semantic maps. The cross-validation module selects different semantic maps for cross-validation and assigns a weight vector to each semantic map. The data processing definition module acquires multi-source heterogeneous data through the multi-sensor feedback module, extracts data features from the multi-source heterogeneous data, generates multiple heterogeneous features, and obtains weight vectors through the cross-validation module. Based on the weight vectors, it corrects the multiple heterogeneous features to obtain unified feature layer feature data. Heterogeneous The feature fusion module, a heterogeneous feature fusion module, acquires feature data through the data processing definition module, and fuses the feature data with depth, contour, speed and semantic features to obtain a dynamic semantic map of the environment. The early warning decision module divides regions based on a dynamic semantic map of the environment and classifies early warning levels based on its own movement characteristics, generating early warning region decisions. The cross-validation module selects laser and 3D models from the same time sequence. It maps radar points from the 3D model to the laser model, obtains difference point clusters through clustering analysis, calculates the volume of these clusters, and then filters out excessively small difference points to obtain the difference regions in the laser and 3D models. It then analyzes the spatial distribution of these difference points. The cross-validation module records regions shared by the laser and 3D models as high-weight regions and regions unique to both models as normal-weight regions. Finally, the cross-validation module compares the visual model with both the laser and 3D models, further analyzing the unique regions. The verification process is as follows: if a unique region in the laser model also exists in the visual model, it is upgraded to a high-weight region; if a unique region in the laser model does not exist in the visual model, it is downgraded to a low-weight region. Similarly, if a unique region in the 3D model exists in the visual model, it is upgraded to a high-weight region; if a unique region in the 3D model does not exist in the visual model, it is downgraded to a low-weight region. If an object exists in both the laser and 3D models but there are differences, it is recorded as an abnormal region. The laser and 3D models are then compared with the visual model to obtain the overlap degree of abnormal regions in both the visual and laser models, as well as the overlap degree of abnormal regions in both the visual and 3D models. The two sets of overlap degrees are compared, and the set with the higher overlap degree between the 3D and laser models is designated as a high-weight factor, while the set with the lower overlap degree is designated as a low-weight factor.

2. The robot real-world environment perception and early warning system based on multi-sensor data fusion according to claim 1, characterized in that, The multi-sensor feedback module is connected to the binocular vision unit, laser scanning unit, radar detection unit, and inertial measurement unit, and receives the feedback depth vision image, laser scanning image, 3D radar point cloud image, and motion inertial results. It performs time synchronization and spatial registration of the multi-sensor results through hardware synchronization and software synchronization. After time synchronization and spatial registration, the data from multiple sensors are recorded as multi-source heterogeneous data.

3. The robot real-world environment perception and early warning system based on multi-sensor data fusion according to claim 1, characterized in that, The single semantic map obtained by the cross-validation module includes a visual model, a laser model, a 3D model, and a motion model. Specifically, a visual model is generated from depth visual images, a laser model is generated from laser scan images, a 3D model is generated from radar echo images, and a motion model is generated from motion inertia results.

4. The robot real-world environment perception and early warning system based on multi-sensor data fusion according to claim 1, characterized in that, After acquiring multi-source heterogeneous data, the data processing definition module obtains accurate surface texture and structural models through a laser model, and identifies moving objects in the environment through a 3D model, determining whether the objects are in motion and their structural models. The data processing definition module then uses a deep learning model to perform object recognition on the visual model, obtaining object semantics from the visual model. Based on the high-weight and low-weight regions defined by the cross-validation module, the data processing definition module sorts the confidence levels and generates surface textures, structural models, motion states, and object semantics with different weights, recording them as feature data. After acquiring the feature data, the heterogeneous feature fusion module performs deep fusion of the feature data through a convolutional neural network, ultimately generating a dynamic semantic map of the environment containing texture, motion, structural, and semantic information.

5. The robot real-world environment perception and early warning system based on multi-sensor data fusion according to claim 1, characterized in that, The early warning decision module divides the robot body into a near-field area and a far-field area layer by layer outward. When dividing the robot body, the early warning decision module first obtains the set path for the robot's movement, then obtains the motion inertia result through the multi-sensor feedback module, and uses the path forward direction and the motion direction of the current motion inertia result as weight vectors to offset the near-field area graphic, forming the deformed near-field area and far-field area. After dividing the near-field and far-field areas, the early warning decision module uses a dynamic semantic map of the environment to conduct risk warnings, identify objects in the near-field and far-field areas that pose a collision risk with the robot's set path, and determine the loss result after the collision based on semantics. The collision risk level is calculated by multiplying the collision risk and the result at any time, and the collision risk level is used as the warning level. The warning is then issued through the network.

Citation Information

Patent Citations

  • Target-level multi-source sensor asynchronous fusion method based on mining area scene

    CN116304971A

  • Intelligent robot control method based on multi-sensor cooperation

    CN119990181A