Autonomous obstacle avoidance mobile robot control method and system with fusion perception
By generating a real-time environmental semantic map through multi-source fusion perception technology and performing dynamic obstacle avoidance learning, the problem of real-time obstacle avoidance and adaptability of mobile robots in complex environments is solved, and efficient autonomous obstacle avoidance control is achieved.
Patent Information
- Application Number
- CN202610870574.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-08-25
AI Technical Summary
In existing technologies, mobile robots lack closed-loop feedback in environmental perception and dynamic obstacle avoidance decision-making in complex and dynamic environments, resulting in insufficient real-time obstacle avoidance and adaptability, making it difficult to cope with sudden dynamic obstacles and unstructured terrain.
A multimodal environmental perception dataset is generated by collecting data from multiple heterogeneous sensors. Spatiotemporal alignment and multi-source fusion feature map generation are performed to construct a real-time environmental semantic map. Dynamic obstacle avoidance learning is then carried out to construct an obstacle avoidance action sequence and update the environmental semantic map, forming an autonomous navigation closed-loop control strategy.
It improves the accuracy of environmental perception and the real-time performance of obstacle avoidance decisions in dynamic environments, enhances autonomous adaptability, and enables reliable obstacle avoidance in complex environments.
Smart Images

Figure CN122632841A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics technology, specifically to a control method and system for an autonomous obstacle avoidance mobile robot based on fused perception. Background Technology
[0002] Mobile robots are widely used in industrial manufacturing, logistics and warehousing, service companionship, emergency rescue, and autonomous driving. In real-world applications, robots need to navigate autonomously in complex, dynamic, and unstructured environments, such as human-robot collaborative workshops, crowded shopping malls, and outdoor areas with temporary obstacles. Uncertainties in the environment, such as the sudden appearance of pedestrians, object displacement, and changes in lighting, pose significant challenges to the safe and efficient operation of mobile robots. Traditional autonomous obstacle avoidance relies on environmental information acquired by LiDAR or vision cameras, combined with pre-set behavioral rules or simple geometric models for obstacle detection and avoidance. For example, obstacle avoidance based on ultrasonic or infrared sensors can only detect nearby obstacles, lacking understanding of distant obstacles and environmental semantics. While vision-based obstacle avoidance methods are information-rich, they are easily affected by lighting, shadows, and texture features. When facing dynamic, complex, and ever-changing real-world environments, they often exhibit significant shortcomings such as insufficient perception accuracy, weak environmental understanding, and limited obstacle avoidance strategies, making it difficult to cope with sudden dynamic obstacles and unstructured terrain. Furthermore, most current multi-source fusion methods still remain at the level of simple superposition of the perception layer or feature layer, lacking the ability to deeply align and understand the semantics of multimodal data in a spatiotemporal manner. This makes it difficult to generate real-time environmental representations that can be directly used for obstacle avoidance decisions. Moreover, when faced with dynamic obstacles, obstacle avoidance control strategies often lack the ability to predict the movement trend of obstacles, resulting in unnatural obstacle avoidance paths, untimely obstacle avoidance, or getting stuck in local optima, thus failing to achieve the synergistic evolution of obstacle avoidance strategies and environmental cognition.
[0003] Therefore, current technologies suffer from several technical problems, including a lack of closed-loop feedback between environmental perception and dynamic obstacle avoidance decision-making, and insufficient real-time performance and adaptability of mobile robots in complex dynamic environments. Summary of the Invention
[0004] This application provides a control method and system for autonomous obstacle avoidance mobile robots based on fused perception, which solves the technical problems in the prior art of lacking closed-loop feedback between environmental perception and dynamic obstacle avoidance decision-making, as well as the insufficient real-time performance and adaptability of mobile robots in complex dynamic environments. It achieves the technical effect of improving the accuracy of environmental perception, the real-time performance of obstacle avoidance decision-making, and the autonomous adaptive control capability of mobile robots in dynamic environments.
[0005] This application provides a control method for an autonomous obstacle avoidance mobile robot based on fused perception. The method includes: acquiring a multimodal environmental perception dataset by performing multi-source heterogeneous sensing based on the mobile robot's travel environment; spatiotemporally aligning the multimodal environmental perception dataset to generate a multi-source fused feature map; performing dynamic and static analysis based on the multi-source fused feature map to construct a real-time environmental semantic map; performing dynamic obstacle avoidance learning based on the real-time environmental semantic map to construct an obstacle avoidance action sequence to simulate and control the mobile robot to avoid obstacles; updating the real-time environmental semantic map based on the simulated obstacle avoidance results; and constructing an autonomous navigation closed-loop control strategy.
[0006] In a possible implementation, the multimodal environment perception dataset is spatiotemporally aligned to generate a multi-source fusion feature map. The method includes: parsing the multimodal environment perception dataset to extract 3D point cloud data, depth vision image data, and inertial measurement data. The 3D point cloud data includes a first sampling clock parameter, the depth vision image data includes a second sampling clock parameter, and the inertial measurement data includes a third sampling clock parameter. Using the third sampling clock parameter as a global time reference value, the first and second sampling clock parameters are interpolated according to the global time reference value to generate a time-synchronized dataset. Angular velocity and acceleration information from the inertial measurement data are extracted and pre-integrated to obtain multi-frame pose transformation data. Data sensing and localization are performed on the 3D point cloud data and the depth vision image based on the time-synchronized dataset to construct a global map coordinate system. The multi-frame pose transformation data is mapped to the global map coordinate system for spatial alignment to obtain spatially aligned 3D point cloud data and spatially aligned depth vision image data. The spatially aligned 3D point cloud data and spatially aligned depth vision image data are correlated pixel-by-pixel to construct the multi-source fusion feature map.
[0007] In a possible implementation, the multi-source fusion feature map is constructed by pixel-by-pixel associating the spatially aligned 3D point cloud data and the spatially aligned depth visual image data. The method includes: performing semantic segmentation on the spatially aligned depth visual image data to generate a semantic segmentation map; extracting pixels from the semantic segmentation map and performing cross-modal attention aggregation on the spatially aligned 3D point cloud data to calculate multiple association weight values; filtering the semantic segmentation map according to the multiple association weight values to determine multiple target point cloud subsets; and extracting the 3D coordinate data and semantic category data of the multiple target point cloud subsets and concatenating them to construct the multi-source fusion feature map.
[0008] In a possible implementation, a real-time environmental semantic map is constructed based on dynamic and static analysis of the multi-source fusion feature map. The method includes: maintaining a global static semantic map by introducing a three-dimensional voxel mesh based on the target region; comparing and analyzing the multi-source fusion feature map with the global static semantic map to obtain static observation data and dynamic observation data; incrementally updating the global static semantic map based on the static observation data to generate a global static semantic update map; performing dynamic target detection based on the dynamic observation data to generate a dynamic target list; and fusing the global static semantic update map with the dynamic target list to construct the real-time environmental semantic map.
[0009] In a possible implementation, the multi-source fusion feature map is compared and analyzed with the global static semantic map to obtain static observation data and dynamic observation data. The method includes: projecting the multi-source fusion feature map onto the global static semantic map for consistency verification and calculating multiple feature point matching coefficients; setting a preset feature consistency threshold and comparing the multiple feature point matching coefficients with the feature consistency threshold; when the multiple feature point matching coefficients are greater than or equal to the feature consistency threshold, they are identified as static observation points; static target processing is performed based on the static observation points to generate static observation data; when the multiple feature point matching coefficients are less than the feature consistency threshold, they are identified as potential dynamic points; spatial clustering association matching is performed based on the potential dynamic points to determine dynamic observation points; dynamic target processing is performed based on the dynamic observation points to generate dynamic observation data.
[0010] In a possible implementation, spatial clustering and association matching are performed based on the potential dynamic points to determine dynamic observation points. The method includes: calculating multiple point distance values based on the potential dynamic points; performing region growth on the potential dynamic points according to the multiple point distance values to obtain point growth coefficients; clustering the potential dynamic points according to the point growth coefficients to construct multiple point cloud clusters; performing multidimensional analysis based on the multiple point cloud clusters to obtain shape descriptors and motion descriptors for the multiple point cloud clusters; performing similarity matching between the shape descriptors and the motion descriptors and the potential dynamic points; and determining the dynamic observation points based on the similarity matching results.
[0011] In a possible implementation, dynamic obstacle avoidance learning is performed based on the real-time environmental semantic map to construct an obstacle avoidance action sequence. The method includes: extracting semantic occupancy information from the global static semantic update map; using the real-time position information of the mobile robot as the center and combining it with the semantic occupancy information for association mapping to construct a static risk potential field; performing dynamic trajectory prediction based on the dynamic target list to determine target motion trajectory prediction information; and constructing a dynamic risk corridor by spreading the target motion trajectory prediction information along the time axis; spatiotemporally superimposing the static risk potential field and the dynamic risk corridor to construct a composite risk parameter set; constructing a continuous action space based on the real-time position information of the mobile robot; and mapping the composite risk parameter set to the continuous action space for dynamic obstacle avoidance search learning to construct the obstacle avoidance action sequence.
[0012] This application also provides a fusion-perception autonomous obstacle avoidance mobile robot control system, the system comprising: a data acquisition module for acquiring multi-source heterogeneous sensor data based on the mobile robot's travel environment to obtain a multimodal environmental perception dataset; a feature analysis module for spatiotemporally aligning the multimodal environmental perception dataset to generate a multi-source fusion feature map, performing dynamic and static analysis based on the multi-source fusion feature map to construct a real-time environmental semantic map; and a control strategy construction module for performing dynamic obstacle avoidance learning based on the real-time environmental semantic map, constructing an obstacle avoidance action sequence to simulate and control the mobile robot to avoid obstacles, updating the real-time environmental semantic map based on the simulated obstacle avoidance results, and constructing an autonomous navigation closed-loop control strategy.
[0013] This application proposes a method and system for autonomous obstacle avoidance mobile robot control based on fused perception. The method involves acquiring a multi-modal environmental perception dataset through multi-source heterogeneous sensing of the mobile robot's travel environment. This dataset is then spatiotemporally aligned to generate a multi-source fused feature map, which is subjected to dynamic and static analysis to construct a real-time environmental semantic map. Dynamic obstacle avoidance learning is then performed to construct an obstacle avoidance action sequence, simulating obstacle avoidance control of the mobile robot. The real-time environmental semantic map is updated based on the simulated obstacle avoidance results, thus constructing an autonomous navigation closed-loop control strategy. This approach addresses the technical problems of insufficient closed-loop feedback between environmental perception and dynamic obstacle avoidance decision-making, as well as the inadequate real-time performance and adaptability of mobile robots in complex dynamic environments. It effectively improves the accuracy of environmental perception, the real-time performance of obstacle avoidance decision-making, and the autonomous adaptive control capability of mobile robots in dynamic environments. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0015] Figure 1 This is a schematic flowchart of the autonomous obstacle avoidance mobile robot control method based on fused perception provided in an embodiment of this application.
[0016] Figure 2 This is a schematic diagram of the structure of the autonomous obstacle avoidance mobile robot control system based on fusion perception, provided in an embodiment of this application.
[0017] Figure labeling: Data acquisition module 10, feature analysis module 20, control strategy construction module 30. Detailed Implementation
[0018] To further illustrate the technical means and effects adopted by the present invention in order to achieve the intended purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments, based on the specific implementation methods, structures, features and effects of the present invention.
[0019] This application provides a control method for an autonomous obstacle avoidance mobile robot based on fused perception, such as... Figure 1 As shown, the method includes:
[0020] Step S100: Perform multi-source heterogeneous sensing acquisition based on the mobile robot's travel environment to obtain a multimodal environmental perception dataset.
[0021] Preferably, during its movement, the mobile robot simultaneously collects different types of data from the same environment using various types of sensors installed on its body. These include, but are not limited to, LiDAR collecting 3D point cloud data, which includes the distance and orientation of obstacles, as well as the geometric shape information of object surfaces; depth cameras or vision cameras collecting depth vision image data, which includes the color, texture, and pixel-level depth information of the environment; and inertial measurement units (IMUs) collecting inertial measurement data, which includes the three-axis angular velocity information and three-axis acceleration information of the mobile robot body. By combining raw data from different sources and different modalities, a multimodal environmental perception dataset is obtained.
[0022] Step S200: Spatiotemporally align the multimodal environment perception dataset to generate a multi-source fusion feature map, and perform dynamic and static analysis based on the multi-source fusion feature map to construct a real-time environmental semantic map.
[0023] Step S200 further includes: parsing the multimodal environment perception dataset to extract 3D point cloud data, depth vision image data, and inertial measurement data, wherein the 3D point cloud data includes a first sampling clock parameter, the depth vision image data includes a second sampling clock parameter, and the inertial measurement data includes a third sampling clock parameter; using the third sampling clock parameter as a global time reference value, interpolating the first sampling clock parameter and the second sampling clock parameter according to the global time reference value to generate a time synchronization dataset; extracting angular velocity and acceleration information from the inertial measurement data for pre-integration calculation to obtain multi-frame pose transformation data; performing data sensing and localization on the 3D point cloud data and the depth vision image based on the time synchronization dataset to construct a global map coordinate system; mapping the multi-frame pose transformation data to the global map coordinate system for spatial alignment to obtain 3D point cloud spatial alignment data and depth vision image spatial alignment data; and associating the 3D point cloud spatial alignment data and the depth vision image spatial alignment data pixel by pixel to construct the multi-source fusion feature map.
[0024] Preferably, the multimodal environment perception dataset is parsed to extract 3D point cloud data, depth vision image data, and inertial measurement data, respectively. The 3D point cloud data, depth vision image data, and inertial measurement data each contain a first sampling clock parameter, a second sampling clock parameter, and a third sampling clock parameter, i.e., they are accompanied by their own sampling time stamps. Then, the third sampling clock parameter of the inertial measurement data is selected as the global time reference value, i.e., a unified time standard. The sampling time points of the point cloud data and image data are calculated according to the global time reference value, including the calculation of the first sampling clock parameter and the second sampling clock parameter through time interpolation. That is, the data at the intermediate time point is estimated between the data at known time points, and a time synchronization dataset is generated. Angular velocity and acceleration information are extracted from inertial measurement data and pre-integrated to calculate rotation angles by continuously accumulating angular velocity with time increments, and velocity and position changes by accumulating acceleration with time increments. This allows for the calculation of the position and attitude changes of the mobile robot at different times, resulting in multi-frame pose transformation data. Next, time-synchronized datasets are used to perform data sensing and localization on 3D point cloud data and depth vision images to determine the specific position and orientation of each frame in the real world, thus constructing a unified global map coordinate system. The multi-frame pose transformation data is then mapped to the global map coordinate system for spatial alignment. Coordinate transformations are performed on the point cloud data and image data to ensure all data are in the same spatial reference system. The output results are spatially aligned 3D point cloud data and depth vision image data. Finally, the 3D point cloud data and depth vision image data are correlated pixel-by-pixel. For each pixel in the image, its corresponding 3D spatial point in the point cloud data is found, and geometric information is merged with color / texture information to generate a multi-source fusion feature map. Each location in this map simultaneously possesses 3D spatial coordinates and visual semantic features.
[0025] Furthermore, step S200 also includes: performing semantic segmentation on the spatial alignment data of the depth visual image to generate a semantic segmentation map; extracting pixels from the semantic segmentation map to perform cross-modal attention aggregation on the spatial alignment data of the three-dimensional point cloud and calculating multiple association weight values; filtering the semantic segmentation map according to the multiple association weight values to determine multiple target point cloud subsets; extracting the three-dimensional coordinate data and semantic category data of the multiple target point cloud subsets and splicing them together to construct the multi-source fusion feature map.
[0026] Preferably, the input is spatially aligned depth vision image data. Each pixel in the image is classified and identified to determine the obstacle it belongs to, and a semantic segmentation map is output. Different colored regions in this map represent different categories, including pedestrians, vehicles, and static facilities such as streetlights, roadblocks, and building walls. Then, the semantic segmentation map and spatially aligned 3D point cloud data are input. Each pixel in the semantic segmentation map is used as a query reference, and cross-modal attention is calculated between it and spatial points in the 3D point cloud. The correlation or importance between each image pixel and each point cloud point is evaluated, and multiple association weight values are output. Each weight value represents the correlation between the image pixel and the point cloud point. The degree of matching between points is determined, i.e., the strength of information association between each 3D point and each semantic category. Selection criteria are set, such as retaining pixel-point pairs with weights higher than a threshold. Pixels highly correlated with point cloud data are selected from the semantic segmentation map according to multiple association weight values, and the corresponding subsets of point cloud data are determined. Multiple target point cloud subsets are output, each subset corresponding to a certain category or region selected in the semantic segmentation map. Finally, the 3D coordinate data and semantic category data of the multiple target point cloud subsets are extracted and concatenated to form a multi-source fusion feature map. Each point in this fusion feature map simultaneously possesses accurate 3D spatial location information and a clear obstacle semantic category label.
[0027] Furthermore, step S200 also includes: maintaining a three-dimensional voxel mesh based on the target region to construct a global static semantic map; comparing and analyzing the multi-source fusion feature map with the global static semantic map to obtain static observation data and dynamic observation data; incrementally updating the global static semantic map based on the static observation data to generate a global static semantic update map; performing dynamic target detection based on the dynamic observation data to generate a dynamic target list; and fusing the global static semantic update map with the dynamic target list to construct the real-time environmental semantic map.
[0028] Preferably, a pre-defined target area is input. Within the spatial range of this target area, a three-dimensional voxel mesh is introduced, dividing the space into small cubic grids as basic storage units to maintain and record static information in the environment, such as the location and semantic category of the ground, walls, and fixed facilities. A global static semantic map representing the static background of the area is output. Then, the fused feature map at the current moment is compared and analyzed point by point or region by region to obtain static observation data. The part that matches the static map in the current fused feature map is the background part. The dynamic observation data is the part that does not match the static map in the current fused feature map, i.e., the part of moving objects or newly appearing obstacles. Then, the global static semantic map is incrementally updated using static observation data, that is, only the changed parts are added or modified, such as updating the ground occupancy or correcting the semantic information of a voxel, to generate a global static semantic update map. Next, the dynamic observation data is processed by clustering, tracking and recognition to detect independent dynamic targets, such as moving pedestrians and moving vehicles, and a dynamic target list is generated, in which each entry describes the position, size, speed and semantic category of a dynamic target. Finally, the static background map and the dynamic target list are overlaid and merged as two layers, where the static layer provides the basic structure of the environment and the dynamic layer provides the information of moving obstacles at the current moment, to construct the final real-time environmental semantic map, which contains both the long-term stable information of the static environment and the instantaneous change information of the dynamic targets, to guide the mobile robot in obstacle avoidance navigation.
[0029] Furthermore, step S200 also includes: projecting the multi-source fused feature map onto the global static semantic map for consistency verification, calculating multiple feature point matching coefficients; setting a feature consistency threshold, comparing the multiple feature point matching coefficients with the feature consistency threshold; when the multiple feature point matching coefficients are greater than or equal to the feature consistency threshold, they are identified as static observation points; performing static target processing based on the static observation points to generate static observation data; when the multiple feature point matching coefficients are less than the feature consistency threshold, they are identified as potential dynamic points; performing spatial clustering association matching based on the potential dynamic points to determine dynamic observation points; and performing dynamic target processing based on the dynamic observation points to generate dynamic observation data.
[0030] Preferably, the multi-source fusion feature map includes a 3D point cloud with semantic labels at the current time. This point cloud is projected onto the coordinate system of the global static semantic map. In the overlapping area after projection, the consistency of each feature point (or voxel) with the existing information of the corresponding position in the static map is checked. A feature point matching coefficient is calculated for each feature point to quantify the degree of matching between the current observation point and the corresponding point in the static map, such as whether the position overlaps or whether the semantic category is consistent. A feature consistency threshold is preset as a judgment standard, and the matching coefficient of each feature point is numerically compared with the threshold. When the feature point matching coefficient is greater than or equal to the feature consistency threshold, the feature point is marked as a static observation point, indicating that the point belongs to the static background part of the environment, such as the ground, walls, or fixed facilities. All data marked as static observation points are aggregated and processed to generate static observation data. When the feature point matching coefficient is less than the feature consistency threshold, the feature point is marked as a potential dynamic point, indicating that the point may belong to a moving object, a newly appearing obstacle, or a temporary object not present in the static map. For data marked as potential dynamic points, further spatial clustering and association matching are performed. That is, points that are close together are grouped together based on spatial distance and neighborhood relationships. Through clustering and association matching, a set of points belonging to the same moving or changing entity is finally identified from the potential dynamic points and determined as dynamic observation points. Finally, all data determined as dynamic observation points are aggregated and processed to generate dynamic observation data.
[0031] Furthermore, step S200 also includes: calculating multiple point distance values based on the potential dynamic points; performing region growth on the potential dynamic points according to the multiple point distance values to obtain point growth coefficients; clustering the potential dynamic points according to the point growth coefficients to construct multiple point cloud clusters; performing multidimensional analysis based on the multiple point cloud clusters to obtain shape descriptors and motion descriptors for the multiple point cloud clusters; performing similarity matching between the shape descriptors and the motion descriptors and the potential dynamic points; and determining the dynamic observation points based on the similarity matching results.
[0032] Preferably, all potential dynamic points are input, and the spatial distance between each potential dynamic point and its neighboring points is calculated. Based on the spatial distance, starting from a seed point, neighboring points with a distance less than a set threshold are continuously incorporated into the same region, i.e., region growth is performed. The point growth coefficient corresponding to each growth region is output, such as the number of points in the region, the region radius, density, etc., to describe the growth degree of the continuous region. The potential dynamic points are clustered according to the point growth coefficient, and spatially adjacent points with the same growth coefficient are grouped into the same category to determine multiple point cloud clusters. Each point cloud cluster represents a spatially independent set of points that may be dynamic objects. Then, multidimensional processing is performed on each point cloud cluster. Feature analysis extracts shape descriptors to describe the geometric shape of point cloud clusters, such as aspect ratio, volume, flatness, and columnarity. Motion descriptors describe the motion characteristics of point cloud clusters, such as velocity vectors, acceleration, consistency of motion direction, and displacement magnitude. Finally, Euclidean distance is used to match the shape and motion descriptors with potential dynamic points. The higher the matching degree, the more the point cloud cluster conforms to the characteristics of real dynamic objects, such as the shape and motion pattern of people or vehicles. Based on the similarity matching results, reliable dynamic observation points are selected from the potential dynamic points and assigned to real dynamic targets, such as pedestrians and vehicles.
[0033] Step S300: Based on the real-time environmental semantic map, perform dynamic obstacle avoidance learning, construct an obstacle avoidance action sequence to simulate and control the mobile robot to avoid obstacles, update the real-time environmental semantic map according to the simulated obstacle avoidance results, and construct an autonomous navigation closed-loop control strategy.
[0034] Step S500 further includes: extracting semantic occupancy information from the global static semantic update map; using the real-time position information of the mobile robot as the center and combining it with the semantic occupancy information for association mapping to construct a static risk potential field; performing dynamic trajectory prediction based on the dynamic target list to determine target motion trajectory prediction information; and constructing a dynamic risk corridor by spreading the target motion trajectory prediction information along the time axis; spatiotemporally superimposing the static risk potential field and the dynamic risk corridor to construct a composite risk parameter set; constructing a continuous action space based on the real-time position information of the mobile robot; and mapping the composite risk parameter set to the continuous action space for dynamic obstacle avoidance search and learning to construct the obstacle avoidance action sequence.
[0035] Preferably, the global static semantic update map includes the location and semantic category of static objects such as the ground, walls, and fixed facilities. Semantic occupancy information is extracted from this map, that is, the spatial location occupied by static objects and the semantic category of the objects. For example, walls have a high risk factor and the ground has a high safety factor. Centered on the real-time position of the mobile robot, the occupancy information is mapped to a risk value in space. The closer to the static object, the higher the semantic risk level, and the greater the risk value. A static risk potential field is generated to represent the risk distribution of static obstacles in the environment to the robot. Dynamic trajectory prediction is performed on each dynamic target in the dynamic target list to estimate the path it may take in the future, that is, the target motion trajectory prediction information. Then, the target motion trajectory prediction information is diffused according to the time axis. For example, the risk range of points on the predicted trajectory expands over time, forming strip-shaped risk areas. Dynamic risk corridors are determined to represent the spatial areas that dynamic targets may occupy at various times in the future and their risk distribution.
[0036] Preferably, the static risk potential field and the dynamic risk corridor are spatiotemporally superimposed. At each spatial location and each future time point, the combined risk value of static obstacle risk and dynamic target prediction risk is comprehensively considered, and a composite risk parameter set is output to describe the comprehensive risk faced by the robot at different times and locations. Then, a continuous action space is constructed with the robot's real-time position as the starting point, that is, the robot can continuously change motion parameters such as speed, direction, and angular velocity. The composite risk parameter set is mapped into this action space as a cost function to evaluate the merits of each action. Then, dynamic obstacle avoidance search learning is performed through optimization algorithms or reinforcement learning to find a sequence of actions from the current position to the target point with the lowest cumulative risk and the smoothest motion. This sequence is the obstacle avoidance action sequence, which consists of multiple consecutive motion commands, such as linear velocity and angular velocity that change with time, to guide the mobile robot to execute sequentially to avoid all static and dynamic obstacles.
[0037] Preferably, before actually sending instructions to the robot for physical execution, the robot's motion process is simulated in a simulation or virtual environment according to the obstacle avoidance action sequence. During the simulation, it is observed whether the robot collides with obstacles and whether the obstacle avoidance path is smooth and efficient. The simulated obstacle avoidance results are output, including feedback information such as collision situations, path deviation, and risk accumulation values during the simulation. Then, the simulated obstacle avoidance results are analyzed. For example, if the risk value estimate for a certain area is too low, leading to a simulated collision, or if the predicted trajectory of a dynamic target is inconsistent with the actual simulated motion, the real-time environmental semantic map is corrected. This includes adjusting the parameters of the static risk potential field, correcting the predicted trajectory of dynamic targets, supplementing missing obstacle information in the map, or removing outdated dynamic targets. An updated real-time environmental semantic map is output. Next, the updated real-time environmental semantic map and the obstacle avoidance action sequence are input and integrated into a closed-loop control loop: environment perception → map construction / update → obstacle avoidance action sequence planning → simulation execution → map update based on simulation results → replanning. The entire loop requires no manual intervention and is autonomously cyclically run by the robot, outputting a complete autonomous navigation closed-loop control strategy. This enables the robot to achieve reliable autonomous obstacle avoidance navigation in complex dynamic environments through continuous self-simulation and self-correction.
[0038] In the above text, refer to Figure 1 A method for controlling an autonomous obstacle-avoiding mobile robot based on fused perception, according to embodiments of the present invention, is described in detail. Next, reference will be made to... Figure 2 A control system for an autonomous obstacle avoidance mobile robot based on fused perception, according to an embodiment of the present invention, is described.
[0039] The autonomous obstacle avoidance mobile robot control system based on fusion perception according to embodiments of the present invention addresses the technical problems in the prior art, namely, the lack of closed-loop feedback between environmental perception and dynamic obstacle avoidance decision-making, and the insufficient real-time performance and adaptability of mobile robots in complex dynamic environments. It achieves the technical effect of improving the accuracy of environmental perception, the real-time performance of obstacle avoidance decision-making, and the autonomous adaptive control capability of mobile robots in dynamic environments. Figure 2 As shown, the autonomous obstacle avoidance mobile robot control system with fused perception includes: a data acquisition module 10, a feature analysis module 20, and a control strategy construction module 30.
[0040] The data acquisition module 10 is used to acquire multi-source heterogeneous sensor data based on the mobile robot's travel environment to obtain a multimodal environment perception dataset; the feature analysis module 20 is used to perform spatiotemporal alignment of the multimodal environment perception dataset to generate a multi-source fusion feature map, and to perform dynamic and static analysis based on the multi-source fusion feature map to construct a real-time environmental semantic map; the control strategy construction module 30 is used to perform dynamic obstacle avoidance learning based on the real-time environmental semantic map, construct an obstacle avoidance action sequence to simulate and control the mobile robot to avoid obstacles, update the real-time environmental semantic map according to the simulated obstacle avoidance results, and construct an autonomous navigation closed-loop control strategy.
[0041] The specific configuration of the feature analysis module 20 will be described in detail below. The feature analysis module 20 further includes: parsing the multimodal environment perception dataset to extract 3D point cloud data, depth vision image data, and inertial measurement data, wherein the 3D point cloud data includes a first sampling clock parameter, the depth vision image data includes a second sampling clock parameter, and the inertial measurement data includes a third sampling clock parameter; using the third sampling clock parameter as a global time reference value, interpolating the first and second sampling clock parameters according to the global time reference value to generate a time synchronization dataset; extracting angular velocity and acceleration information from the inertial measurement data for pre-integration calculation to obtain multi-frame pose transformation data; performing data sensing and localization on the 3D point cloud data and the depth vision image based on the time synchronization dataset to construct a global map coordinate system; mapping the multi-frame pose transformation data to the global map coordinate system for spatial alignment to obtain 3D point cloud spatial alignment data and depth vision image spatial alignment data; and performing pixel-by-pixel association on the 3D point cloud spatial alignment data and the depth vision image spatial alignment data to construct the multi-source fusion feature map.
[0042] The specific configuration of the feature analysis module 20 will be described in detail below. The feature analysis module 20 further includes: performing semantic segmentation on the spatially aligned data of the depth visual image to generate a semantic segmentation map; extracting pixels from the semantic segmentation map and performing cross-modal attention aggregation on the spatially aligned data of the 3D point cloud to calculate multiple association weight values; filtering the semantic segmentation map according to the multiple association weight values to determine multiple target point cloud subsets; and extracting the 3D coordinate data and semantic category data of the multiple target point cloud subsets and concatenating them to construct the multi-source fusion feature map.
[0043] The specific configuration of the feature analysis module 20 will be described in detail below. The feature analysis module 20 further includes: maintaining a three-dimensional voxel mesh based on the target region to construct a global static semantic map; comparing and analyzing the multi-source fused feature map with the global static semantic map to obtain static observation data and dynamic observation data; incrementally updating the global static semantic map based on the static observation data to generate a global static semantic update map; performing dynamic target detection based on the dynamic observation data to generate a dynamic target list; and performing layer fusion of the global static semantic update map and the dynamic target list to construct the real-time environmental semantic map.
[0044] The specific configuration of the feature analysis module 20 will be described in detail below. The feature analysis module 20 further includes: projecting the multi-source fused feature map onto the global static semantic map for consistency verification, and calculating multiple feature point matching coefficients; setting a preset feature consistency threshold, and comparing the multiple feature point matching coefficients with the feature consistency threshold; when the multiple feature point matching coefficients are greater than or equal to the feature consistency threshold, they are identified as static observation points; static target processing is performed based on the static observation points to generate static observation data; when the multiple feature point matching coefficients are less than the feature consistency threshold, they are identified as potential dynamic points; spatial clustering association matching is performed based on the potential dynamic points to determine dynamic observation points; dynamic target processing is performed based on the dynamic observation points to generate dynamic observation data.
[0045] The specific configuration of the feature analysis module 20 will be described in detail below. The feature analysis module 20 further includes: calculating multiple point distance values based on the potential dynamic points; performing region growing on the potential dynamic points according to the multiple point distance values to obtain point growth coefficients; clustering the potential dynamic points according to the point growth coefficients to construct multiple point cloud clusters; performing multidimensional analysis based on the multiple point cloud clusters to obtain shape descriptors and motion descriptors for the multiple point cloud clusters; performing similarity matching between the shape descriptors, the motion descriptors, and the potential dynamic points; and determining the dynamic observation points based on the similarity matching results.
[0046] The specific configuration of the control strategy construction module 30 will be described in detail below. The control strategy construction module 30 further includes: extracting semantic occupancy information from the global static semantic update map; using the real-time position information of the mobile robot as the center and combining it with the semantic occupancy information for association mapping to construct a static risk potential field; performing dynamic trajectory prediction based on the dynamic target list to determine target motion trajectory prediction information; and constructing a dynamic risk corridor by spreading the target motion trajectory prediction information along the time axis; spatiotemporally superimposing the static risk potential field and the dynamic risk corridor to construct a composite risk parameter set; constructing a continuous action space based on the real-time position information of the mobile robot; mapping the composite risk parameter set to the continuous action space for dynamic obstacle avoidance search and learning to construct the obstacle avoidance action sequence.
[0047] The autonomous obstacle avoidance mobile robot control system with fusion perception provided in this embodiment of the invention can execute the autonomous obstacle avoidance mobile robot control method with fusion perception provided in this embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A control method for an autonomous obstacle avoidance mobile robot based on fused perception, characterized in that, The method includes: Multi-source heterogeneous sensing data acquisition is performed based on the mobile robot's travel environment to obtain a multimodal environmental perception dataset; The multimodal environment perception dataset is spatiotemporally aligned to generate a multi-source fusion feature map. Based on the multi-source fusion feature map, dynamic and static analysis is performed to construct a real-time environmental semantic map. Dynamic obstacle avoidance learning is performed based on the real-time environmental semantic map, an obstacle avoidance action sequence is constructed to simulate and control the mobile robot to avoid obstacles, the real-time environmental semantic map is updated according to the simulated obstacle avoidance results, and an autonomous navigation closed-loop control strategy is constructed.
2. The autonomous obstacle avoidance mobile robot control method based on fused perception as described in claim 1, characterized in that, The method involves spatiotemporally aligning the multimodal environment-aware dataset to generate a multi-source fusion feature map, including: Based on the multimodal environment perception dataset, 3D point cloud data, depth vision image data, and inertial measurement data are extracted. The 3D point cloud data includes a first sampling clock parameter, the depth vision image data includes a second sampling clock parameter, and the inertial measurement data includes a third sampling clock parameter. Using the third sampling clock parameter as a global time reference value, time interpolation is performed on the first sampling clock parameter and the second sampling clock parameter according to the global time reference value to generate a time synchronization dataset. Angular velocity and acceleration information from inertial measurement data are extracted and pre-integrated to obtain multi-frame pose transformation data; Based on the time-synchronized dataset, data sensing and localization are performed on the 3D point cloud data and the depth visual image to construct a global map coordinate system; The multi-frame pose transformation data is mapped to the global map coordinate system for spatial alignment to obtain 3D point cloud spatial alignment data and depth vision image spatial alignment data. The spatially aligned 3D point cloud data and the spatially aligned depth visual image data are correlated pixel by pixel to construct the multi-source fusion feature map.
3. The autonomous obstacle avoidance mobile robot control method based on fused perception as described in claim 2, characterized in that, The method involves associating the spatially aligned 3D point cloud data and the spatially aligned depth visual image data pixel-by-pixel to construct the multi-source fusion feature map, including: Semantic segmentation is performed on the spatially aligned data of the depth visual image to generate a semantic segmentation map; The pixels of the semantic segmentation map are extracted and cross-modal attention aggregation is performed on the spatial alignment data of the 3D point cloud to calculate multiple association weight values; The semantic segmentation graph is filtered according to the multiple association weight values to determine multiple subsets of target point clouds; The 3D coordinate data and semantic category data of the multiple target point cloud subsets are extracted and concatenated to construct the multi-source fusion feature map.
4. The autonomous obstacle avoidance mobile robot control method based on fused perception as described in claim 1, characterized in that, Based on the multi-source fusion feature map, dynamic and static analysis is performed to construct a real-time environmental semantic map. The method includes: A global static semantic map is constructed by introducing a three-dimensional voxel mesh based on the target region for maintenance. The multi-source fusion feature map is compared and analyzed with the global static semantic map to obtain static observation data and dynamic observation data; Based on the static observation data, the global static semantic map is incrementally updated to generate a global static semantic updated map. Based on the dynamic observation data, dynamic target detection is performed to generate a dynamic target list; The global static semantic update map and the dynamic target list are merged into layers to construct the real-time environmental semantic map.
5. The autonomous obstacle avoidance mobile robot control method based on fused perception as described in claim 4, characterized in that, The method involves comparing and analyzing the multi-source fused feature map with the global static semantic map to obtain static observation data and dynamic observation data. The multi-source fusion feature map is projected onto the global static semantic map for consistency verification, and the matching coefficients of multiple feature points are calculated. A preset feature consistency threshold is set, and the matching coefficients of the multiple feature points are compared with the feature consistency threshold. If the matching coefficient of the plurality of feature points is greater than or equal to the feature consistency threshold, then it is identified as a static observation point; Static target processing is performed based on the static observation points to generate static observation data; If the matching coefficients of the multiple feature points are less than the feature consistency threshold, they are identified as potential dynamic points. Based on the potential dynamic points, spatial clustering and association matching are performed to determine the dynamic observation points; Dynamic target processing is performed based on the dynamic observation points to generate dynamic observation data.
6. The autonomous obstacle avoidance mobile robot control method based on fused perception as described in claim 5, characterized in that, Based on the potential dynamic points, spatial clustering and association matching are performed to determine dynamic observation points. The method includes: Based on the potential dynamic points, calculate multiple point distance values, and perform region growth on the potential dynamic points according to the multiple point distance values to obtain the point growth coefficient; The potential dynamic points are clustered according to the point growth coefficient to construct multiple point cloud clusters; Multidimensional analysis is performed on the multiple point cloud clusters to obtain shape descriptors and motion descriptors for the multiple point cloud clusters; The shape descriptor, the motion descriptor, and the potential dynamic point are matched for similarity, and the dynamic observation point is determined based on the similarity matching result.
7. The autonomous obstacle avoidance mobile robot control method based on fused perception as described in claim 4, characterized in that, Based on the real-time environmental semantic map, dynamic obstacle avoidance learning is performed to construct an obstacle avoidance action sequence. The method includes: Extract the semantic occupancy information of the global static semantic update map, and use the real-time location information of the mobile robot as the center to perform association mapping with the semantic occupancy information to construct a static risk potential field; Based on the dynamic target list, dynamic trajectory prediction is performed to determine the target motion trajectory prediction information. The target motion trajectory prediction information is then diffused according to the time axis to construct a dynamic risk corridor. The static risk potential field and the dynamic risk corridor are spatiotemporally superimposed to construct a composite risk parameter set. A continuous action space is constructed based on the real-time location information of the mobile robot. The composite risk parameter set is mapped to the continuous action space for dynamic obstacle avoidance search and learning, and the obstacle avoidance action sequence is constructed.
8. A control system for an autonomous obstacle avoidance mobile robot based on fused perception, characterized in that, The system is used to implement the autonomous obstacle avoidance mobile robot control method according to any one of claims 1 to 7, and the system includes: The data acquisition module is used to collect data from multiple heterogeneous sensors based on the mobile robot's travel environment to obtain a multimodal environmental perception dataset. The feature analysis module is used to perform spatiotemporal alignment of the multimodal environment perception dataset, generate a multi-source fusion feature map, perform dynamic and static analysis based on the multi-source fusion feature map, and construct a real-time environmental semantic map. The control strategy construction module is used to perform dynamic obstacle avoidance learning based on the real-time environmental semantic map, construct an obstacle avoidance action sequence to simulate the control of the mobile robot to avoid obstacles, update the real-time environmental semantic map according to the simulated obstacle avoidance results, and construct an autonomous navigation closed-loop control strategy.