An Active Auditory Localization Method for Mapless Navigation

Through the active auditory positioning method of reinforced learning training, combined with lidar and auditory information, the problems of obstacle occlusion and field of view limitation in robot map-free navigation are solved, and efficient and accurate target positioning is achieved, suitable for outdoor environments.

CN114563011BActive Publication Date: 2025-08-01PEKING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210079214.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-08-01
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

In the prior art, robot map-free navigation methods are difficult to accurately locate targets when facing obstacle occlusion and field of view limitations, and traditional methods are costly or rely on additional equipment and cannot be effectively applied to outdoor environments.

Method used

Adopt active auditory positioning method based on reinforcement learning, through training the robot navigation model, combining lidar ranging information and auditory direction information, the target is positioned using the sound source, and the output speed command is used for collision-free navigation.

Benefits of technology

It achieves efficient and accurate target positioning in complex environments, reduces dependence on additional equipment, and is suitable for map-free navigation in outdoor environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114563011B_ABST
    Figure CN114563011B_ABST
Patent Text Reader

Abstract

The present invention discloses an active auditory localization method for mapless navigation. The steps include: 1) training a mobile robot navigation model on a simulation platform through a reinforcement learning method; 2) the mobile robot collects the ranging information of the lidar at the current moment, the auditory orientation information obtained based on the target position, and the pose information of the mobile robot odometer according to a set time step; wherein, the lidar is mounted on the mobile robot; 3) inputting the ranging information, the auditory orientation information and the pose information into the mobile robot navigation model trained in step 1) to infer the speed command at the current moment, and the mobile robot navigates to the target position according to the speed command. The present invention adopts a more reliable and effective target localization method, which has high application value for mapless navigation in real scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information science, and relates to an auditory localization method, in particular to an active auditory localization method for mobile robot mapless navigation. Background Art

[0002] So far, robots have created great value in the fields of industrial manufacturing, home service, interstellar exploration, military reconnaissance, etc. Compared with the visual information mobile robot navigation method, the robot navigation based on auditory perception has advantages in protecting privacy. In addition, when the target is not in the robot's field of view or is blocked by obstacles, auditory localization can provide additional information to help the robot determine the target. The autonomous navigation of a mobile robot refers to the process in which the mobile robot perceives the external environment through sensors and combines its own state to reach the target point without collision. Only when the mobile robot has flexible, efficient and robust navigation capabilities can it be better applied in industries, service industries and military, etc.

[0003] The navigation technology of robots can be divided into two categories: map-dependent and map-independent. The map-dependent navigation technology means that the robot needs to build a map of the environment as accurately as possible before navigation. The disadvantage of this method is that the robot needs to spend a long time building the map, and the map is required to be accurate enough to help the robot locate during navigation. The map-independent navigation technology is also called mapless navigation. Traditional algorithms include the dynamic window method, D* algorithm, vector histogram algorithm, etc. With the rise of deep learning, learning-based methods have gradually become a popular research direction for mapless navigation methods. The main method is to model the navigation process of the robot based on reinforcement learning and imitation learning. However, when applying the learned navigation strategy to the real environment, an inevitable problem is how to determine the relative position of the target. Previous work has shown that the methods based on wifi localization and visible light communication have low costs, but require external receivers for the target and the indoor environment to have wifi hotspots or LED lights. The vision-based target localization method has strong flexibility and can process various targets according to semantic types; however, there are problems such as obstacle occlusion and field of view range, and the real-time performance is not good.

[0004] As far as we know, the method based on active auditory localization has not been introduced into the research of mobile robot mapless navigation. The auditory-based localization method can solve the problem of obstacle occlusion, and at the same time does not require a signal receiver. It can be applied to outdoor environments and can also assist the vision-based localization method. Summary of the Invention

[0005] The object of the present invention is to provide an active auditory localization method and apply it to mapless navigation technology. By training the navigation model of the robot through a navigation strategy based on reinforcement learning, and adopting the method of active auditory localization to obtain the continuously converging relative position of the target during the actual navigation process, a more accurate and robust navigation model can be obtained.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] An active auditory localization method for mapless navigation, the steps of which include:

[0008] 1) Training the navigation model of the mobile robot through the reinforcement learning method on the simulation platform;

[0009] 2) The mobile robot collects the ranging information of the lidar at the current moment, the auditory orientation information obtained based on the target position, and the pose information of the mobile robot odometer according to the set time step; wherein, the lidar is mounted on the mobile robot;

[0010] 3) Inputting the ranging information, auditory orientation information and pose information into the navigation model of the mobile robot trained in step 1) to infer the speed command at the current moment, and the mobile robot navigates to the target position according to the speed command.

[0011] Further, the navigation model of the mobile robot includes an Actor network and a Critic network; wherein, the Actor network is used to output the action that can maximize the reward according to the observed state, the state includes the ranging information, auditory orientation information and pose information, and the action is the linear velocity and angular velocity of the mobile robot; the Critic network is used to output the value of <state, action> according to the action information output by the Actor network and the observed information of the current state.

[0012] Further, the method for training the navigation model of the mobile robot through the reinforcement learning method is: first, build different simulation environments, randomly set multiple obstacles and target points in the simulation environment, and then use the set reward formula to motivate the mobile robot to reach the target point.

[0013] Further, the reward calculation formula is: r(s t ,a t ,s t+1 )=α1dis(p t ,p t+1 )+α2(dis(p t ,p target )- dis(p t+1 ,p target))+α3×success+α4×collision; where dis(p t ,p t+1 ) is to calculate the position point p at time t t At time (t+1) the position point p t+1 The displacement between (dis(p t ,p target )-dis(p t+1 ,p target )) is to calculate the position point p at time t+1 t+1 Relative to the position point p at time t t Approaching the target point p target The degree of approach to the target, success means that the target point has been successfully reached, and collision means that a collision has occurred; the coefficients α1, α2, and α3 are all positive, and the coefficient α4 is a negative number; s t is the state at time t, a t is the action at time t (control of linear velocity and angular velocity), s t+1 is the state at time t+1.

[0014] Furthermore, the auditory orientation information includes a 2-dimensional direction vector; the posture information includes 2-dimensional position information and 1-dimensional angle information; the target position includes 2-dimensional target position information; and the speed instruction includes linear velocity and angular velocity.

[0015] Furthermore, whether a collision occurs is determined based on the ranging information of the laser radar; if the minimum value in the ranging information is less than a set threshold, it is determined that the mobile robot has collided with an obstacle.

[0016] Furthermore, the parameters of the randomized obstacle include: the shape and type of the obstacle, the position of the obstacle, and the size of the obstacle.

[0017] Furthermore, auditory directional information is obtained based on an active auditory localization method.

[0018] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the above method.

[0019] A computer-readable storage medium stores a computer program thereon, wherein the computer program implements the steps of the above method when executed by a processor.

[0020] The present invention constructs a mobile robot navigation model based on reinforcement learning. The input of the mobile robot navigation model is the ranging information of the laser radar carried by the mobile robot, the auditory orientation information obtained based on the sound source at the target position, and the position information of the mobile robot odometer. The output is the speed instruction that the mobile robot needs to execute. The training of the model includes training on a simulation platform and training in an actual environment. After the model training is completed, the mobile robot collects the ranging information of the laser radar at the current moment, the auditory orientation information, and the position information of the mobile robot odometer at a certain time step, and uses this information as the model input to infer the speed instruction to be executed, and finally navigates to the target position.

[0021] Furthermore, the method for training the robot navigation model through reinforcement learning in simulation involves constructing different simulation environments based on various real-world indoor layouts, including rectangular environments of 10m×10m, 10m×5m, and 5m×5m, as well as circular areas with radii of 5m and 10m. Using a randomized algorithm, we placed obstacles of varying parameters and shapes at various locations within these environments, with 20 obstacle types per environment. The parameters that required randomization included obstacle shape, location, and size. Each time, we randomly selected a target point within the environment that fell outside the obstacles. Each time the robot explored the environment and reached the target point, we designed a reward formula to incentivize the mobile robot to reach the target point without collision.

[0022] The calculation formula of the immediate reward function of the present invention is:

[0023] r(s t ,a t ,s t+1 )=α1dis(p t ,p t+1 )+α2(dis(p t ,p target )-dis(p t+1 ,p target ))+α3×success+α4×collision;s t s t+1 Represents the state at time t and time t+1, including the robot's posture information, sensor input, and robot speed information. The calculation of immediate feedback consists of four items. The first item is to calculate the position point p at time t. t At time (t+1) the position point p t+1 The displacement between dis(p t ,p t+1 ), the second term calculates the position point p at time t+1 t+1 Relative to the position point p at time t tApproaching the target point p target The degree of approaching the target (dis(p t , p target ) - dis(p t+1 , p target ))). The third term calculates whether the target point has been successfully reached, and the fourth term calculates whether a collision has occurred; the coefficients of the first three terms are positive, and the coefficient of the last term is negative. We collect data from 5 types of shaped environments asynchronously and store them in the experience pool, and use the method of reinforcement learning to train the control model based on the incentive of the reward.

[0024] Furthermore, the lidar information includes 360 - dimensional ranging information; the information of auditory orientation includes a 2 - dimensional direction vector with a modulus of 1; the odometer information includes 2 - dimensional position information and 1 - dimensional angle information; the target position includes 2 - dimensional target position information; the velocity command includes linear velocity and angular velocity. We adopt the current state - of - the - art reinforcement learning algorithm TD3, which is an improved version based on the DDPG algorithm. Improvements have been made in terms of policy action smoothing, the update frequency of the policy network, and the overestimation problem of the state - value function, resulting in a significant performance improvement compared to DDPG. If a collision occurs or the target is reached, it is considered that the training for this round has ended.

[0025] To reduce the difficulty of transferring from simulation to reality, we restored the real indoor situation and robot configuration in the simulation environment as much as possible. Considering that the complexity of the real environment is difficult to be expressed in the simulation environment, we simply reduced the dimension of the ranging information during the simulation training, that is, we uniformly selected 10 dimensions from the 360 - dimensional lidar ranging signals; similarly, only these 10 - dimensional lidar information is used in the real environment. It should be further noted that there are slight differences between the training in the simulation environment and the training in the real environment. In terms of the calculation of the reward, the simulation environment can obtain an unbiased pose of the mobile robot; while on the mobile robot, the pose needs to be calculated through odometry, which is biased due to the existence of cumulative errors. However, considering the operation in a small indoor scenario, this error is tolerable. In terms of collision detection, the simulation environment can detect by checking the intersection of the robot and obstacles because it has all the information; in the actual scenario, we solve it by setting a collision distance threshold. That is, in the 360 - dimensional ranging information of the lidar, if the minimum value is less than the threshold, it is considered that a collision has occurred, and the training for this round is stopped simultaneously.

[0026] Furthermore, during the process of target localization in a real environment, we adopt the method of active auditory localization. This method reduces the uncertainty of auditory localization by continuously estimating the direction of arrival (DOA) of the sound source and actively moving, in combination with odometer information. Since the navigation process is one of continuously avoiding obstacles and approaching the sound source, and the process of approaching is also one where the direct sound is amplified and the reflected sound is reduced, our DOA estimation will become more and more accurate, and ultimately, navigation in the real environment can be achieved.

[0027] Compared with the prior art, the positive effects of the present invention are as follows:

[0028] The present invention adopts a more reliable and effective target localization method in the way of obtaining environmental target information. In an environment serving people, the sounds generated by people's activities are a clue worth utilizing. At the same time, the target localization method based on active sound source localization can achieve a good fusion effect with other localization methods such as visual localization. It has high application value for mapless navigation in real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram of the active auditory navigation of a mobile robot;

[0030] Figure 2 is a schematic diagram of the orientation of a spherical microphone array in different directions;

[0031] Figure 3 is a schematic diagram of the error result of active auditory localization. DETAILED DESCRIPTION OF THE INVENTION

[0032] In order to enable a mobile robot to achieve collision-free navigation in an actual unknown scenario, the present invention proposes an active sound source localization technology for mapless navigation. Through reinforcement learning, the present invention proposes an end-to-end navigation model oriented to the robot platform and the target. This model can learn a complex strategy: the robot selects a moving mode according to environmental information, which includes the original 2D laser ranging result and the target position. At the same time, in order to apply the model trained in the simulation environment to the real environment, we set up active auditory localization to determine the relative position of the target, Figure 1 showing the method by which the robot continuously determines the target position by adjusting its own pose during navigation. Figure 2 is the measurement error of determining the target position in different directions by a spherical microphone array. In order to quantitatively evaluate the performance of active auditory localization, we compared the positioning accuracies of different methods, as shown in Figure 3 , and it can be seen that the method based on active auditory localization has a more accurate positioning accuracy.

[0033] (1) Data acquisition: In the technical solution used in the present invention, it relies on a navigation data set. Since there is no open-source navigation data set currently, we need to construct our own data set. In the Gazebo simulation environment, different simulation environments are built according to the layouts of various real indoor environments.

[0034] (2) Model construction: As Figure 3 shown, the TD3 network structure we adopted includes an Actor network (policy network) and a Critic network (valuation network). Among them, the Actor network is responsible for outputting actions that can maximize the reward based on the observation of the state. For this patent, its input is the ranging information of the lidar, the information of auditory orientation, and the pose information of the mobile robot's odometer. After being processed by the neural network, its output is the linear velocity and angular velocity of the mobile robot (i.e., the action that maximizes the cumulative reward). Among them, the input of the Critic network is the action information output by the Actor network and the observation information of the current state, and the output is the evaluation of the value function (i.e., the cumulative reward) of this <state, action>. For this patent, its input includes two parts. One part is the linear velocity and angular velocity of the mobile robot output by the Actor network, and the other part is the observation of the state, including the ranging information of the lidar, the information of auditory orientation, and the pose information of the mobile robot's odometer, and its output is a fractional value.

[0035] (3) Migration from the simulation model to the physical environment: We use the HOA coding method to determine the direction of the sound source target, and then continuously determine the position of the sound source target through the active movement of the robot. The navigation strategy learning based on reinforcement learning can obtain continuous action instructions according to the position of the sound source target, and finally enable the robot to navigate to the real target position without collision.

[0036] The above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Those of ordinary skill in the art can modify or equivalently replace the technical solution of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. An active auditory localization method for mapless navigation, the steps of which include: 1) Train a mobile robot navigation model through reinforcement learning on a simulation platform; among them, the method of training the mobile robot navigation model through reinforcement learning is: first build different simulation environments, randomly set multiple obstacles and target points in the simulation environment, and then use the set reward formula to motivate the mobile robot to reach the target point; the reward formula is: r(s t ,a t ,s t+1 ) = α1dis(p t ,p t+1 ) + α2(dis(p t ,p target ) - dis(p t+1 ,p target )) + α3×success + α4×collision; where dis(p t ,p t+1 ) is to calculate the displacement between the position point p t at time t and the position point p t+1 at (t + 1) time, (dis(p t ,p target ) - dis(p t+1 ,p target )) is to calculate the degree of approaching the target of the position point p t+1 at time t + 1 relative to the position point p t at time t approaching the target point p target , success represents successfully reaching the target point, and collision represents a collision; the coefficients α1, α2, and α3 are all positive numbers, and the coefficient α4 is a negative number; s t is the state at time t, a t is the action at time t, and s t+1 is the state at time t + 1; 2) The mobile robot collects the ranging information of the lidar at the current moment, the auditory orientation information obtained based on the target position, and the pose information of the mobile robot odometer according to the set time step; wherein, the lidar is mounted on the mobile robot; 3) Input the ranging information, auditory orientation information and pose information into the mobile robot navigation model trained in step 1) to infer the speed command at the current moment, and the mobile robot navigates to the target position according to the speed command.

2. The method according to claim 1, characterized in that The mobile robot navigation model includes an Actor network and a Critic network; wherein, the Actor network is used to output an action that can maximize the reward according to the observed state, the state includes the ranging information, auditory orientation information and pose information, and the action is the linear velocity and angular velocity of the mobile robot; the Critic network is used to output the value of <state, action> according to the action information output by the Actor network and the observed information of the current state.

3. The method according to claim 1, characterized in that, The auditory orientation information includes a 2D direction vector; the pose information includes 2D position information and 1D angle information; the target position includes 2D target position information; the speed command includes linear velocity and angular velocity.

4. The method according to claim 1, characterized in that, Determine whether a collision occurs according to the ranging information of the lidar; if the minimum value in the ranging information is less than the set threshold, it is determined that the mobile robot has collided with an obstacle.

5. The method according to claim 1, wherein Randomizing the parameters of the obstacle includes: the type of obstacle shape, the position of the obstacle, and the size of the obstacle.

6. The method according to claim 1, wherein Obtain auditory orientation information based on the active auditory localization method.

7. A server, characterized in that, Comprising a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intelligent mobile platform map-free autonomous navigation method based on deep reinforcement learning

    CN111141300A

  • Robot map-free navigation method based on time sequence information modeling

    CN112857370A

  • Robot system and positioning navigation method

    WO2021254367A1