A visual and auditory fusion perception and navigation method for mobile robots
By calibrating the parameters of visual and auditory sensors and integrating information, a navigation grid map is constructed to enable autonomous environmental perception and navigation of mobile robots. This solves the problem of single interaction means in existing technologies and provides a natural and intelligent robot interaction method.
Patent Information
- Application Number
- CN202211614647.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-12-15
AI Technical Summary
The existing mobile robot systems have a single means of interaction, voice control is easily affected by noise and far-field environment, and visual information assistance fails to effectively improve the accuracy and convenience of interaction.
The parameters of visual and auditory sensors are calibrated to construct a navigation grid map. Three-dimensional convolution and long short-term memory networks are combined for gesture recognition. Visual and auditory sensors are used to obtain salient target objects, achieving end-to-end visual and auditory information fusion.
It provides a more intelligent and natural way of robot interaction, realizes autonomous environmental perception and navigation, and improves the accuracy and convenience of interaction. It is suitable for sweeping robots, service robots, and inspection robots, etc.
Smart Images

Figure CN116380061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision technology, and in particular to a visual and auditory fusion perception and navigation method for a mobile robot. Background Art
[0002] Vision and auditory information are the primary sources of information for humans to perceive the external environment and are also a crucial foundation for robots to possess artificial intelligence. Unlike existing AI technologies, which process visual and auditory information as two separate research areas, human perception and cognition of the objective world is achieved through the fusion of a vast amount of diverse information acquired through multiple senses, including sight and hearing.
[0003] For low-speed mobile robot systems, providing more natural interaction methods is a direct manifestation of their advanced intelligence. Existing mobile robot systems often provide relatively simple interaction methods, mostly using voice or interactive screens to issue commands. Deeper audio-visual fusion interaction methods have not been implemented in mobile robot systems. Single voice control commands are often affected by environmental factors such as noise and far-field, which will significantly reduce the recall rate and accuracy of voice wake-up. The extensive use of visual sensors will make it possible to use visual information to assist voice and improve the accuracy and convenience of interaction. It has also become an important breakthrough direction for overcoming the disadvantages of single-modal information processing. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a mobile robot visual and auditory fusion perception and navigation method, which can interact with the robot navigation system in a more intelligent and natural way.
[0005] The technical solution adopted by the present invention to solve the technical problem is to provide a mobile robot visual and auditory fusion perception and navigation method, including the following steps:
[0006] Calibrate parameters of the mobile robot's visual sensor system and auditory sensor system;
[0007] Construct navigation grid maps using visual sensor systems and calibrated parameters;
[0008] Using a visual sensor system to acquire a video sequence of the interacting objects, and based on a gesture recognition method based on a three-dimensional convolutional and long short-term memory network, using an attention mechanism and multi-scale feature fusion, to achieve end-to-end gesture behavior recognition with the video sequence as input;
[0009] The target object of interest is extracted from the video sequence and tracked, and a sequence of the target object with significance is obtained by using an auditory sensor system and a visual sensor system.
[0010] The parameter calibration of the visual sensor system and the auditory sensor system of the mobile robot includes:
[0011] Use the calibration plate to calibrate and correct the internal and external parameters of the visual sensor system;
[0012] Placing a sound-emitting unit at a certain position near the calibration plate so that the sound-emitting unit can measure the positional relationship with the calibration plate;
[0013] After the sound unit wakes up the auditory sensor system and obtains the sound angle of the sound unit, the external parameter relationship between the auditory sensor system and the visual sensor system is obtained based on the internal parameters of the visual sensor system and the external parameter relationship between the visual sensor system and the calibration board.
[0014] The method of constructing a navigation grid map using a visual sensor system and calibrated parameters includes:
[0015] Use a visual sensor system to obtain depth information of the scene;
[0016] Based on the depth information, the visual stereo matching technology is used to obtain the disparity map of the visual sensor system, and the depth image is constructed by combining the intrinsic and extrinsic parameters of the visual sensor system. A single-frame 3D point cloud image is obtained based on the depth image;
[0017] Visual SLAM technology is used to calculate the posture parameters of the mobile robot. Continuous frames of three-dimensional point clouds are spliced and fused according to the depth image and the posture parameters of the mobile robot. A three-dimensional point cloud map of the scene is constructed during the movement, and a navigation grid map is established by setting the maximum obstacle height information.
[0018] After establishing the navigation grid map, it also includes non-interventional autonomous environmental perception and map construction based on predefined objective functions, and constructing environmental element maps through the original images and three-dimensional point cloud maps of the scenes collected by the visual sensor system.
[0019] The end-to-end gesture behavior recognition specifically involves learning the short-term spatiotemporal features of gestures through a three-dimensional convolutional neural network, and learning the long-term spatiotemporal features of gestures through a convolutional LSTM network based on the extracted short-term spatiotemporal features, so as to mine the implicit gesture motion trajectory and posture change information between frames in the input video sequence.
[0020] The step of extracting and tracking the target object of interest from the video sequence, and obtaining a sequence of the target object with significance using an auditory sensor system and a visual sensor system, comprises:
[0021] Use multi-target detection and tracking methods to extract and track the target objects of interest from the video sequence;
[0022] The sound source localization method of the auditory sensor system and the line of sight direction detection method of the visual sensor system are used to assign saliency levels to the target objects respectively.
[0023] The target object saliency results obtained by the auditory sensor system and the target object saliency results obtained by the visual sensor system are fused to obtain the final salient target object sequence.
[0024] During fusion, when the difference between the angle between the mobile robot and the target object inferred by the visual sensor system and the angle between the mobile robot and the target object inferred by the auditory sensor system exceeds a threshold, the target object inferred by the visual sensor system is used as the main factor, the target object is reselected and determined, and the execution of this interactive task is abandoned.
[0025] Beneficial effects
[0026] Due to the adoption of the above-mentioned technical solution, the present invention has the following advantages and positive effects compared to the existing technology: the present invention can perceive the three-dimensional information of the environment, and on this basis establish capabilities such as human gesture recognition, human eye gaze detection, and robot navigation maps. The present invention provides the ability to autonomously search for routes in unknown environments based on binocular or multi-viewing vision, achieving rapid environmental perception and map construction. The present invention implements a method for establishing robot navigation target points by fusing multiple visual and auditory information such as microphone array orientation, environmental visual features, human gesture recognition, and human eye gaze detection. The fusion of multi-source information can provide a more natural human-computer interaction method. The present invention can provide a more intelligent and human-friendly natural interaction method for various types of robots such as sweeping robots, service robots, and inspection robots for autonomous robot movement. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Schematic diagram of a method for visual and auditory fusion perception and navigation of a mobile robot according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of calibration of a visual sensor and an auditory sensor in an embodiment of the present invention;
[0029] Figure 3 Schematic diagram of a mobile robot interaction system combining voice, sight, and gestures in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.
[0031] The embodiment of the present invention relates to a method for visual and auditory fusion perception and navigation of a mobile robot, which can provide a more intelligent and natural interaction method close to human habits for various types of robots such as sweeping robots, service robots and inspection robots for autonomous movement of robots. The mobile robot in this embodiment includes a visual sensor system and an auditory sensor system, wherein the visual sensor system can be a binocular camera (composed of two separate monocular cameras) or a multi-camera, and the auditory sensor system can be a microphone array. Figure 1 As shown, the implementation steps are as follows:
[0032] Step 1: Calibrate the parameters of the mobile robot's visual sensor system and auditory sensor system. During calibration, first, use a calibration plate to calibrate and correct the visual sensor system's internal and external parameters. Then, place a sound-generating unit near the calibration plate so that it can measure its positional relationship with the calibration plate. Finally, the sound-generating unit wakes up the auditory sensor system and obtains the sound-generating unit's sound angle. Based on the visual sensor system's internal parameters and the external parameter relationship between the visual sensor system and the calibration plate, the external parameter relationship between the auditory sensor system and the visual sensor system is determined.
[0033] Step 2: Construct a navigation grid map using the visual sensor system and calibrated parameters. After completing step 1, the visual sensor system constructs a 3D point cloud map of the environment through stereo matching and visual positioning techniques, and further constructs a navigation grid map. Specifically, the visual sensor system acquires depth information of the scene; based on the depth information, it uses visual stereo matching techniques to obtain the visual sensor system's disparity map, and combines the intrinsic and extrinsic parameters of the visual sensor system to construct a depth image. A single-frame 3D point cloud map is then obtained based on the depth image; the mobile robot's pose parameters are calculated using visual SLAM technology; consecutive frames of 3D point clouds are spliced and fused based on the depth image and the mobile robot's pose parameters. A 3D point cloud map of the scene is constructed during movement, and a navigation grid map is established by setting maximum obstacle height information. It is worth noting that after constructing the navigation grid map, autonomous environmental perception and map construction can be performed without intervention based on predefined objective functions (e.g., maximum coverage, minimum search efficiency, specific target search efficiency). A map of environmental elements is constructed using the original image collected by the visual sensor system and the 3D point cloud map of the scene.
[0034] Step 3: Initiate and execute the audiovisual fusion task. After completing the visual depth map and 3D point cloud map provided in Step 2, a binocular camera is used to detect key visual interaction information such as the interactive object's gestures, eye gaze, and environmental visual features. Furthermore, combined with the interactive object's voice wake-up and command output, the microphone array is used to determine the direction of the interactive object through sound source orientation. Combining visual and auditory interaction information, the target point location set by the interactive object is determined through fusion setting and cross-comparison, completing the autonomous mobile robot's characteristic tasks such as command issuance, target point setting, and task intervention.
[0035] Step 4: Autonomous Navigation and Perception of the Mobile Robot. After completing Step 3, the mobile robot can autonomously construct a 3D scene map and navigation map of the unfamiliar environment without relying on remote control. After obtaining the navigation map, the autonomous perception robot can achieve autonomous movement, obstacle avoidance, and walking tasks at any point. During autonomous walking, the interactive object can issue any command, interrupting the interactive object's interactive command, and returning to Step 3 to interact naturally with the robot.
[0036] It is not difficult to find that a low-speed mobile robot equipped with a visual sensor system and an auditory sensor system will simultaneously perform visual tasks and auditory tasks after obtaining the external parameter relationship between the visual sensor system and the auditory sensor system. The visual task relies on the depth map constructed by binocular and multi-eye sensors, and uses this as a basis for visual navigation tasks such as visual SLAM and visual three-dimensional reconstruction; at the same time, it relies on RGB and depth maps to establish visual interaction tasks such as gesture detection of interactive objects and human eye sight detection. It should be pointed out that the execution of the auditory task only begins after voice awakening and recognition, and the auditory task relies on the sound source orientation of the microphone array to obtain the specific position information of the interactor. Based on the output of the visual and auditory tasks, the main purpose of the fusion of visual and auditory information provided by this embodiment is to provide the interactive object with a more accurate robot navigation target point and correct interpretation of task instructions. Relying on the fusion output of visual and auditory tasks, the main advantages of this implementation are: achieving two-way visual and auditory wake-up, assisting hearing to reduce redundant information processing, achieving accurate auditory wake-up, and avoiding misoperation during single auditory wake-up; combining line of sight direction detection with voice wake-up to perceive objects of interest, and integrating specific behavior recognition such as gestures with voice processing to distinguish noise and wake up the voice system to process target requests, achieving more accurate wake-up and recognition needs; more naturally configuring temporary task points of mobile robots, realizing the joint interaction of gestures, line of sight, and voice, guiding the robot to reach the designated target location to carry out task operations, etc.
[0037] The present invention will be further explained below by taking a sweeping robot as an example.
[0038] First, use Figure 2The parameters of the visual sensor system and auditory sensor system of the sweeping robot are calibrated in this way. In this embodiment, the visual sensor system is a binocular vision system, and the auditory sensor system is a microphone array. In the calibration method of this embodiment, the internal and external parameters of the binocular vision system are first calibrated and corrected using a chessboard. When the visual sensor system is a multi-eye vision system, the external parameter information between the multi-eye cameras also needs to be calibrated. Then a specific sound unit is used to a specific position on the chessboard. Its coordinates relative to the first corner point of the chessboard are X(x, y, 0). When the external parameters of the chessboard and the binocular vision system are measured to be (R, T), and the directional angle θ between the sound source and the microphone array, it can be inferred that the relative external parameter angle between the binocular vision system and the microphone array is R θ -θ, where R θ is the rotation angle component of the extrinsic rotation in the microphone array plane.
[0039] The main purpose of the binocular vision system is to obtain depth information of the scene. Binocular stereo matching technology can be used to obtain the left and right eye disparity map of the binocular vision system. Combined with the internal and external parameter information of the microphone array, the left and right object depth images can be constructed, and further a single-frame 3D point cloud map can be constructed. Binocular vision SLAM technology is then used to calculate the position parameters of the sweeping robot. After obtaining the single-frame binocular depth map, the camera pose parameter information obtained by SLAM technology can be combined to stitch and fuse the 3D point clouds of consecutive frames. Finally, a 3D point cloud map of the scene is constructed during the movement. By setting the maximum obstacle height information, a navigation grid map can be established.
[0040] For gesture recognition in visual tasks: This implementation example adopts deep learning technology, a gesture recognition method based on three-dimensional convolution and long short-term memory (LSTM) network, and uses attention mechanism and multi-scale feature fusion to realize end-to-end gesture behavior recognition with video sequence as input. This embodiment first learns the short-term and spatial features of gestures through a three-dimensional convolutional neural network, and then learns the long-term and spatial features of gestures through a convolutional LSTM network based on the extracted short-term and spatial features, and fully explores the implicit gesture motion trajectory and posture change information between the input sequence frames. Through data-driven network autonomous learning, it can better solve the problems in traditional gesture recognition such as the failure of the inter-frame correlation model due to excessive hand movement, which in turn causes tracking loss.
[0041] For human gaze detection in visual tasks: This embodiment uses a salient object sequence acquisition technique based on sound source orientation and gaze direction detection. First, a multi-target detection and tracking method is used to extract and track the target objects of interest from the video sequence (e.g., the MOTS algorithm is used to detect and track target objects). Then, a microphone array-based sound source localization method and a binocular vision system-based gaze direction detection method are used to assign saliency levels to the target objects. Finally, the target object saliency results obtained from sound source localization and gaze direction detection are combined to obtain a final salient target object sequence.
[0042] The sweeping robot of this embodiment can autonomously construct a three-dimensional map and a navigation grid map of the environment without relying on remote control, and in the process can use audio-visual fusion means to interact and update temporary or sudden task points, and ultimately change the target location of the task. Figure 3 As shown in the figure, after using voice orientation to obtain the angle θ1 between the interaction object and the robot, the binocular vision can obtain the three-dimensional coordinates (x0, y0, z0) of the interaction object relative to the robot, and then reversely infer the polar coordinates (r, θ2, ), if the difference between the visually inferred angles θ2 and θ1 between the interactive object and the machine exceeds the threshold δ, that is:
[0043] ||θ1-θ2||≥δ
[0044] It is judged that the interaction object inferred by vision conflicts with the interaction object inferred by hearing, and the interaction object inferred by vision should be given priority. The interaction object should be selected and determined again, and the execution of this interaction task should be abandoned.
[0045] The autonomous perception and navigation tasks of this embodiment include autonomous search and mapping, updating temporary task points for audiovisual fusion, collecting perception elements, and setting arbitrary navigation task points. It should be noted that autonomous search and mapping relies on a preset objective function, performing local path planning according to the set maximum coverage and minimum time target search, and recording the current path range for global path search. The collection of perception elements includes scene visual perception elements such as visual target detection, visual target recognition, semantic recognition, and panoptic segmentation, and the collection of these elements is not limited to online or offline methods.
[0046] It can be seen that the sweeping robot using this embodiment can realize cleaning within the specified area through the audio-visual fusion system. After the interactive object sends a voice wake-up interaction command, it can further clarify the demarcation of the task area through gestures, so as to quickly and accurately reach the designated location for cleaning operations, avoiding the operation of designated points on the map, and making the interaction of sudden task instructions more accurate, convenient and natural.
Claims
1. A mobile robot visual and auditory fusion perception and navigation method, characterized in that: The following steps are included: Calibrate parameters of the mobile robot's visual sensor system and auditory sensor system; Construct navigation grid maps using visual sensor systems and calibrated parameters; Using a visual sensor system to acquire a video sequence of the interacting objects, and based on a gesture recognition method based on a three-dimensional convolutional and long short-term memory network, using an attention mechanism and multi-scale feature fusion, to achieve end-to-end gesture behavior recognition with the video sequence as input; Extracting and tracking a target object of interest from the video sequence, and obtaining a sequence of salient target objects using an auditory sensor system and a visual sensor system, specifically comprising: Use multi-target detection and tracking methods to extract and track the target objects of interest from the video sequence; The sound source localization method of the auditory sensor system and the line of sight direction detection method of the visual sensor system are used to assign saliency levels to the target objects respectively. The target object saliency results obtained by the auditory sensor system and the target object saliency results obtained by the visual sensor system are fused to obtain the final salient target object sequence.
2. The mobile robot visual and auditory fusion perception and navigation method according to claim 1, characterized in that: The parameter calibration of the visual sensor system and the auditory sensor system of the mobile robot includes: Use the calibration plate to calibrate and correct the internal and external parameters of the visual sensor system; Placing a sound-emitting unit at a certain position near the calibration plate so that the sound-emitting unit can measure the positional relationship with the calibration plate; After the sound unit wakes up the auditory sensor system and obtains the sound angle of the sound unit, the external parameter relationship between the auditory sensor system and the visual sensor system is obtained based on the internal parameters of the visual sensor system and the external parameter relationship between the visual sensor system and the calibration board.
3. The mobile robot visual and auditory fusion perception and navigation method according to claim 1, characterized in that: The method of constructing a navigation grid map using a visual sensor system and calibrated parameters includes: Use a visual sensor system to obtain depth information of the scene; Based on the depth information, the visual stereo matching technology is used to obtain the disparity map of the visual sensor system, and the depth image is constructed by combining the intrinsic and extrinsic parameters of the visual sensor system. A single-frame 3D point cloud image is obtained based on the depth image; Visual SLAM technology is used to calculate the posture parameters of the mobile robot. Continuous frames of three-dimensional point clouds are spliced and fused according to the depth image and the posture parameters of the mobile robot. A three-dimensional point cloud map of the scene is constructed during the movement, and a navigation grid map is established by setting the maximum obstacle height information.
4. The mobile robot visual and auditory fusion perception and navigation method according to claim 3, characterized in that: After establishing the navigation grid map, it also includes non-interventional autonomous environmental perception and map construction based on predefined objective functions, and constructing environmental element maps through the original images and three-dimensional point cloud maps of the scenes collected by the visual sensor system.
5. The mobile robot visual and auditory fusion perception and navigation method according to claim 1, characterized in that: The end-to-end gesture behavior recognition specifically involves learning the short-term spatiotemporal features of gestures through a three-dimensional convolutional neural network, and learning the long-term spatiotemporal features of gestures through a convolutional LSTM network based on the extracted short-term spatiotemporal features, so as to mine the implicit gesture motion trajectory and posture change information between frames in the input video sequence.
6. The mobile robot visual and auditory fusion perception and navigation method according to claim 1, characterized in that: During fusion, when the difference between the angle between the mobile robot and the target object inferred by the visual sensor system and the angle between the mobile robot and the target object inferred by the auditory sensor system exceeds a threshold, the target object inferred by the visual sensor system is used as the main factor, the target object is reselected and determined, and the execution of this interactive task is abandoned.
Citation Information
Patent Citations
Visual language indoor navigation method and system, terminal and application
CN112710310A