Autonomous navigation robot system based on multi-modal perception

Through the autonomous navigation robot system with multimodal perception and decision-making, the problem of blind spots and insufficient interaction under a single sensor is solved, efficient navigation and natural language interaction in complex environments are achieved, and the robot's autonomy and user experience in complex environments are improved.

CN120576757APending Publication Date: 2025-09-02SENDAO ENERGY (HANGZHOU) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510691439.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing navigation robots rely on a single sensor, which has problems with environmental occlusion, lighting changes and detection blind spots, and it is difficult for user interaction to achieve real-time flexibility.

Method used

The visual perception module, lidar module, voice command module and central processing unit are adopted to generate navigation paths and obstacle avoidance strategies through multimodal data fusion algorithm and path planning algorithm, and combined with self-attention mechanism, Transformer model and model prediction control, multimodal collaborative perception and decision-making are achieved.

Benefits of technology

Achieve high robust navigation in an environment of uneven lighting, complex obstacles and dynamic personnel interference, supports natural language interaction, and improve user experience and system flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120576757A_ABST
    Figure CN120576757A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot navigation and artificial intelligence, in particular to an autonomous navigation robot system based on multi-modal perception, which comprises a visual perception module, a laser radar module, a voice instruction module, a central processing unit and an execution control unit, the visual perception module is used for collecting an environment RGB-D image; the laser radar module is used for acquiring an environment distance point cloud; the voice instruction module is used for receiving and analyzing a user voice instruction; the central processing unit generates a navigation path and an obstacle avoidance strategy through a multi-modal data fusion algorithm and a path planning algorithm; the execution control unit drives the robot chassis to move according to the navigation path and the obstacle avoidance strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot navigation and artificial intelligence technology, and in particular to an autonomous navigation robot system based on multimodal perception. Background Art

[0002] The widespread application of service robots in industrial automation, intelligent security, logistics and warehousing, and home care has placed higher demands on their autonomous navigation and human-robot interaction capabilities in complex environments. Existing navigation robots often rely on single sensors, such as lidar, vision, or ultrasound, each of which presents challenges with environmental occlusion, lighting variations, and detection blind spots. Furthermore, when responding to user commands, they often rely on preset paths or simple semantic matching, making it difficult to achieve real-time and flexible interaction. Summary of the Invention

[0003] In order to overcome the above technical problems at least to a certain extent, the present application provides an autonomous navigation robot system based on multimodal perception.

[0004] The scheme of this application is as follows:

[0005] An autonomous navigation robot system based on multimodal perception, comprising: a visual perception module, a laser radar module, a voice command module, a central processing unit, and an execution control unit;

[0006] The visual perception module is used to collect environmental RGB-D images;

[0007] The laser radar module is used to obtain the environmental distance point cloud;

[0008] The voice command module is used to receive and analyze user voice commands;

[0009] The central processing unit generates a navigation path and obstacle avoidance strategy through a multimodal data fusion algorithm and a path planning algorithm;

[0010] The execution control unit drives the robot chassis to move according to the navigation path and obstacle avoidance strategy.

[0011] Preferably, the multimodal data fusion algorithm includes: utilizing a self-attention mechanism to perform spatiotemporal alignment of visual and lidar data, and performing adaptive fusion in combination with sensor confidence weights.

[0012] Preferably, the path planning algorithm includes: constructing a grid map based on the fused environmental data, using the A* algorithm to perform global path planning, and using model predictive control to achieve local dynamic obstacle avoidance.

[0013] Preferably, the voice command module includes a far-field microphone array and an end-to-end semantic parsing unit based on a pre-trained Transformer model, which is used to convert natural language instructions into navigation targets or auxiliary execution commands.

[0014] Preferably, it also includes:

[0015] The calibration module is used to automatically calibrate the internal and external parameters of the visual perception module and the lidar module based on the standard calibration plate and environmental characteristics during system startup or operation, and feed the calibration results back to the central processing unit to ensure the consistency of the spatial coordinates of the multimodal data source.

[0016] Preferably, the central processing unit further comprises:

[0017] The simultaneous localization and mapping unit is used to construct a dense environmental map in real time based on the fused environmental point cloud and RGB-D data, and to estimate its own position and posture in combination with the odometry information. The map constructed by the simultaneous localization and mapping unit is used for subsequent path planning and historical playback analysis.

[0018] Preferably, the execution control unit includes:

[0019] The multimodal motion control submodule adopts a composite control strategy based on model predictive control and PID. Based on the navigation path and real-time obstacle avoidance strategy, it predicts the robot's motion trajectory and adjusts the speed and steering angle of each drive wheel to meet dynamic constraints and improve motion smoothness.

[0020] Preferably, it also includes:

[0021] A wireless communication module and a remote monitoring unit. The wireless communication module is used to upload the robot's operating status, environmental map, path planning results, and alarm information to a cloud server via Wi-Fi or 5G network;

[0022] The remote monitoring unit is based on a mobile application or web interface, allowing users to view the robot's location, map view and voice command execution in real time, and supports a one-click call emergency stop function.

[0023] The technical solution provided by this application may have the following beneficial effects:

[0024] Through multimodal collaboration of vision, radar, and voice, this technical solution achieves highly robust navigation in environments with uneven lighting, complex obstacles, and dynamic human interference; it also supports natural language interaction, improving user experience and system flexibility.

[0025] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0027] Figure 1 This is a structural diagram of an autonomous navigation robot system based on multimodal perception provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0029] Figure 1 This is a schematic diagram of the structure of an autonomous navigation robot system based on multimodal perception provided by an embodiment of the present application, with reference to Figure 1 , an autonomous navigation robot system based on multimodal perception, including: a visual perception module, a lidar module, a voice command module, a central processing unit and an execution control unit;

[0030] The visual perception module is used to collect RGB-D images of the environment;

[0031] The laser radar module is used to obtain the environmental distance point cloud;

[0032] The voice command module is used to receive and analyze user voice commands;

[0033] The central processing unit generates navigation paths and obstacle avoidance strategies through multimodal data fusion algorithms and path planning algorithms;

[0034] The execution control unit drives the robot chassis to move according to the navigation path and obstacle avoidance strategy.

[0035] Through multimodal collaboration of vision, radar, and voice, this technical solution achieves highly robust navigation in environments with uneven lighting, complex obstacles, and dynamic human interference; it also supports natural language interaction, improving user experience and system flexibility.

[0036] It should be noted that the multimodal data fusion algorithm includes: using the self-attention mechanism to align the visual and lidar data in time and space, and combining the sensor confidence weights for adaptive fusion.

[0037] The attention mechanism in the Transformer is used to match and align RGB-D image features with point cloud features in both temporal and spatial dimensions. The weight of each sensor in the fusion process is dynamically adjusted based on its signal-to-noise ratio or historical accuracy in the current environment.

[0038] When sensor data quality degrades (e.g., lens obstruction, radar reflection distortion), its impact is automatically reduced to ensure overall perception accuracy. Self-attention captures long-range dependencies and complex spatial relationships, improving recognition of dynamic / complex scenes (crowds, furniture, water, etc.). The same attention framework allows for seamless integration of multiple modalities (e.g., infrared, sonar), ensuring strong system scalability.

[0039] It should be noted that the path planning algorithm includes: building a grid map based on the fused environmental data, using the A* algorithm to perform global path planning, and using model predictive control to achieve local dynamic obstacle avoidance.

[0040] Grid Map: Projects the 3D environment onto a ground grid, marking obstacles and traversable areas.

[0041] A* global planning: Find the optimal (shortest or lowest cost) path on a static grid map.

[0042] MPC local obstacle avoidance: During the robot's movement, it predicts the trajectory of nearby moving obstacles in real time and optimizes its speed and direction at the next moment online.

[0043] In this technical solution, A* ensures overall path quality, while MPC ensures dynamic response speed and safety. MPC integrates dynamic constraints during obstacle avoidance to avoid sudden stops or detours, reducing energy consumption and improving comfort. MPC's online optimization solution is minimal, meeting millisecond-level response requirements.

[0044] It should be noted that the voice command module includes a far-field microphone array and an end-to-end semantic parsing unit based on a pre-trained Transformer model, which is used to convert natural language instructions into navigation targets or auxiliary execution commands.

[0045] The Transformer's ability to model long-term contextual dependencies enables accurate parsing of complex commands and long sentences. The combination of a far-field array and an end-to-end model ensures more stable command capture and understanding in noisy environments. Compared to the traditional ASR→NLU→DM process, the end-to-end approach simplifies the system architecture, reducing latency and error accumulation.

[0046] It should be noted that the system also includes:

[0047] The calibration module is used to automatically calibrate the internal and external parameters of the visual perception module and the lidar module based on the standard calibration plate and environmental characteristics during system startup or operation, and feed the calibration results back to the central processing unit to ensure the consistency of the spatial coordinates of the multimodal data source.

[0048] When the robot detects a known calibration plate or natural environment features (corners, lines), it automatically performs geometric alignment calculations between the camera and the radar.

[0049] After the calibration is completed, the results are updated to the fusion algorithm in real time to ensure the spatial consistency of multimodal data.

[0050] It should be noted that the central processing unit further includes:

[0051] The simultaneous localization and mapping unit is used to build a dense environmental map in real time based on the fused environmental point cloud and RGB-D data, and to estimate its own position in combination with odometry information. The map constructed by the simultaneous localization and mapping unit is used for subsequent path planning and historical playback analysis.

[0052] Multimodal data complement each other, generating maps that are more complete and detailed, with significantly reduced positioning errors.

[0053] When the light is extreme or a single sensor fails, SLAM operation can still be maintained by relying on other modalities.

[0054] The constructed dense map can be used for offline playback, path optimization, or digital twin system integration.

[0055] It should be noted that the execution control unit includes:

[0056] The multimodal motion control submodule adopts a composite control strategy based on model predictive control and PID. Based on the navigation path and real-time obstacle avoidance strategy, it predicts the robot's motion trajectory and adjusts the speed and steering angle of each drive wheel to meet dynamic constraints and improve motion smoothness.

[0057] MPC prediction: solves the optimal control sequence based on path and obstacle information in the short future time;

[0058] PID regulation: real-time correction of deviations caused by model errors or external disturbances to ensure tracking accuracy.

[0059] MPC-planned trajectories take dynamics and obstacle avoidance into account, while PID feedback correction ensures execution accuracy and prevents sudden braking or spinouts. On non-ideal surfaces like slopes and gravel, PID quickly responds to unexpected situations like skidding or spins. Predictive control reduces unnecessary acceleration and braking, improving range.

[0060] It should be noted that the system also includes:

[0061] Wireless communication module and remote monitoring unit. The wireless communication module is used to upload the robot's operating status, environmental map, path planning results, and alarm information to the cloud server via Wi-Fi or 5G network;

[0062] The remote monitoring unit is based on a mobile application or web interface, allowing users to view the robot's location, map view, and voice command execution in real time, and supports a one-click emergency stop function.

[0063] Through the wireless communication module, operators and maintenance personnel can monitor the robot's health and location at any time, proactively detecting and addressing anomalies. Emergency stop functions can be triggered remotely, minimizing losses from accidents. Historical operating data stored in the cloud can be used for performance analysis, fault diagnosis, and algorithm upgrades.

[0064] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0065] It should be noted that, in the description of this application, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" refers to at least two.

[0066] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0067] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0068] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0069] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0070] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0071] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0072] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. An autonomous navigation robot system based on multimodal perception, characterized in that: include: Visual perception module, lidar module, voice command module, central processing unit and execution control unit; The visual perception module is used to collect environmental RGB-D images; The laser radar module is used to obtain the environmental distance point cloud; The voice command module is used to receive and analyze user voice commands; The central processing unit generates a navigation path and obstacle avoidance strategy through a multimodal data fusion algorithm and a path planning algorithm; The execution control unit drives the robot chassis to move according to the navigation path and obstacle avoidance strategy.

2. The system according to claim 1, wherein: The multimodal data fusion algorithm includes: using a self-attention mechanism to align visual and lidar data in time and space, and combining sensor confidence weights for adaptive fusion.

3. The system according to claim 1, wherein: The path planning algorithm includes: building a grid map based on the fused environmental data, using the A* algorithm to perform global path planning, and using model predictive control to achieve local dynamic obstacle avoidance.

4. The system according to claim 1, wherein: The voice command module includes a far-field microphone array and an end-to-end semantic parsing unit based on a pre-trained Transformer model, which is used to convert natural language instructions into navigation targets or auxiliary execution commands.

5. The system according to claim 1, wherein: Also includes: The calibration module is used to automatically calibrate the internal and external parameters of the visual perception module and the lidar module based on the standard calibration plate and environmental characteristics during system startup or operation, and feed the calibration results back to the central processing unit to ensure the consistency of the spatial coordinates of the multimodal data source.

6. The system according to any one of claims 1 to 4, characterized in that The central processing unit further comprises: The simultaneous localization and mapping unit is used to construct a dense environmental map in real time based on the fused environmental point cloud and RGB-D data, and to estimate its own position and posture in combination with the odometry information. The map constructed by the simultaneous localization and mapping unit is used for subsequent path planning and historical playback analysis.

7. The system according to claim 1, wherein: The execution control unit includes: The multimodal motion control submodule adopts a composite control strategy based on model predictive control and PID. Based on the navigation path and real-time obstacle avoidance strategy, it predicts the robot's motion trajectory and adjusts the speed and steering angle of each drive wheel to meet dynamic constraints and improve motion smoothness.

8. The system according to any one of claims 1 to 4, characterized in that Also includes: A wireless communication module and a remote monitoring unit. The wireless communication module is used to upload the robot's operating status, environmental map, path planning results, and alarm information to a cloud server via Wi-Fi or 5G network; The remote monitoring unit is based on a mobile application or web interface, allowing users to view the robot's location, map view and voice command execution in real time, and supports a one-click call emergency stop function.

Citation Information

Cited By

  • Mobile robot control device giving consideration to field navigation patrol and transportation load

    CN120821235A

  • Robot automatic navigation method and system based on depth vision fusion

    CN121140802A

  • Mobile robot navigation control method and system based on natural language

    CN121541649A

  • Natural language based mobile robot navigation control method and system

    CN121541649B