Mobile robot autonomous navigation method and system for cross-regional map-free scene
By introducing auxiliary exploration tasks and internal reward mechanisms into mobile robot navigation, combined with external rewards, the problem of mobile robots being difficult to achieve efficient navigation in cross-region scenarios without maps is solved, efficient exploration and path optimization are achieved, and the success rate of navigation tasks is improved.
Patent Information
- Application Number
- CN202510067667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
In cross-region scenarios without maps, it is difficult for mobile robots to achieve efficient path planning and target arrival, traditional navigation methods are difficult to adapt to dynamic changes and unknown complex scenarios, and traditional DRL methods have problems of strategy overfitting and local optimal solutions.
A cross-regional map-free scenario autonomous navigation method is adopted. By introducing auxiliary exploration tasks and an intrinsic reward mechanism, combining external rewards and intrinsic reward mechanisms, the exploration ability of the mobile robot in the local optimal area is improved, and the intrinsic rewards are calculated through the episodic memory bank and comparator network, the mobile robot is encouraged to actively explore the unseen area.
It realizes efficient exploration and path optimization of mobile robots in unknown environments, avoids local optimal problems, improves the training efficiency and adaptability of the algorithm under sparse reward conditions, and significantly enhances the success rate of navigation tasks.
Smart Images

Figure CN119984267A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous navigation of mobile robots, and in particular to an autonomous navigation method and system for mobile robots in cross-regional map-free scenarios. Background Art
[0002] Mobile robots are an important branch of mobile robot technology. They have autonomous motion capabilities and can perform tasks in the environment without relying on fixed trajectories or preset paths. Compared with traditional industrial mobile robots, mobile robots have high flexibility and adaptability. They are widely used in logistics and transportation, environmental monitoring, security inspection, medical services and disaster relief, and have significant academic value and engineering significance. When mobile robots perform tasks such as workshop operations, biological sampling, and environmental mapping, the real-time information obtained by sensors, planning reasonable paths, and executing motion control are the basic guarantees for the smooth completion of related tasks. As a core issue in mobile robot research, autonomous navigation technology directly affects the intelligence level of mobile robots and the efficiency of task execution. However, in mapless cross-regional scenarios, mobile robot autonomous navigation tasks face significant technical challenges. Such tasks require mobile robots to complete efficient path planning and target arrival in unknown environments, relying only on local perception information obtained by sensors. Traditional navigation methods mainly rely on path planning algorithms based on grid maps or classic heuristic planning algorithms (such as A*, Dijkstra, etc.), but these methods usually assume that the prior information of the environment is known and are difficult to adapt to dynamic changes and unknown complex scenarios. In addition, model-based navigation methods are sensitive to environmental uncertainties and are easily limited by local optimal solutions, making it difficult to achieve general navigation capabilities across regions.
[0003] In recent years, deep reinforcement learning (DRL) has provided a new solution for mobile robot navigation tasks due to its powerful learning ability in high-dimensional state space. By combining deep neural networks, DRL can directly learn policy mapping from high-dimensional sensor data (such as lidar, camera, etc.) and generate autonomous navigation behavior. However, traditional DRL methods still have some key problems when solving mapless cross-regional tasks: on the one hand, external reward signals usually rely on goal-oriented designs, such as shortening the target distance or completing the feedback of the task status; however, such rewards are often sparse and difficult to balance multi-objective requirements (such as obstacle avoidance and target arrival), which can easily lead to overfitting of specific paths and ignoring the global optimal solution. On the other hand, traditional algorithms mainly rely on external signals to drive policy updates, lack the ability to actively explore based on environmental changes, and are difficult to effectively jump out of the local optimal area. Therefore, in cross-regional mapless scenarios, mobile robots are prone to fall into local optimal areas and cannot actively jump out of inefficient strategies. Summary of the invention
[0004] The present disclosure provides a method and system for autonomous navigation of a mobile robot in a cross-region map-free scenario to solve the problems existing in related technologies. The technical solution is as follows:
[0005] In a first aspect, the present disclosure provides a method for autonomous navigation of a mobile robot in a cross-region map-free scenario, comprising the following steps:
[0006] The mobile robot observes the state information s at time step t t , the state information s t Input into the Actor network;
[0007] The Actor network sends the state information s t After forward propagation, according to the strategy π θ (s t ) Output an action selection a t , and send the action selection a t Give the mobile robot, the mobile robot performs the action and selects a t After that, the environment changes to a new state s t+1 , where θ is the network parameter of Actor;
[0008] Calculate the extrinsic reward r e,t and intrinsic reward r i,t , for each moment t, the extrinsic reward r e,t is defined as r e,t =r g +r c +r s , where r g is the goal-oriented part reward, r c Partial reward for safe obstacle avoidance, r s To assist exploration rewards; intrinsic rewards i,t The state information s t The input is generated by the intrinsic reward mechanism structure, where the intrinsic reward mechanism structure consists of four parts: feature extraction network, comparator network, episodic memory library and intrinsic reward estimation model;
[0009] The resulting quaternion <s t ,a t ,r t ,s t+1 >Store in the experience replay unit, the Critic network is based on the four-tuple in the experience replay unit <s t ,a t ,r t ,s t+1 > Approximate action value function Qφ (s t ,a t ) trains the Actor network and the Critic network and updates its parameters, where φ is the Critic network parameter, r t =r e,t +r i,t ;
[0010] The trained Actor network and Critic network models are used to guide the mobile robot to perform real-time cross-regional map-free autonomous navigation tasks.
[0011] Optionally, the status information s t Contains the scanning information of the single-line two-dimensional laser radar, the mobile robot's target point position information and the mobile robot's own position information.
[0012] Optionally, the action selects a t Includes the speed and angular velocity of the mobile robot.
[0013] Optionally, the intrinsic reward r i,t The generation steps are as follows:
[0014] The mobile robot observes the current lidar scanning data and the state information s of the target point at each time step t. t , after being processed by the feature extraction network, the state feature φ(s t );
[0015] The comparator network is used to evaluate the state feature s i and j Reachability R(s) within k steps i ,s j )=f(φ(s i ),φ(s j )), by training a multi-layer fully connected neural network f to predict the state pair (s i ,s j ) and uses the Sigmoid activation function to map the output to [0,1], representing the accessibility score R(s i ,s j ), a score close to 1 indicates that the states are relatively “close” or “reachable”, and a score close to 0 indicates that the two states are far away or unreachable.
[0016] At the current time step t, the mobile robot maintains a context memory M t-1 , used to memorize the historical observation information of the mobile robot; the current state feature φ(s t ) and episodic memory bank M t-1 The reachability score R(s) between the elements int ,M t-1 ), when the accessibility score is lower than the threshold τ, φ(s t ) is embedded in the episodic memory bank and defines the intrinsic reward r i,t =β·R(s t ,M t-1 ), where β is the intrinsic reward coefficient.
[0017] Optionally, the network structures of the Actor network and the Critic network are composed of a 512*512*512*2 fully connected layer and an activation function Relu, the parameter optimizer is Adam, and the capacity of the experience playback unit is 2×10 6 , respectively calculate the network loss function to update the corresponding parameters θ, φ.
[0018] Optionally, before the trained Actor network and Critic network models are used to guide the mobile robot to perform a real-time cross-region map-free autonomous navigation task, the following steps are also included:
[0019] Evaluation indicators are used to evaluate the safety and reliability of the mobile robot in the map-free cross-region navigation task, wherein the evaluation indicators include at least one of the following indicators: navigation success rate, collision rate, timeout rate or execution step number.
[0020] Optionally, the goal-oriented part reward is defined as follows:
[0021]
[0022] Among them, P t-1 and P t is the position of the mobile robot at the last moment and time step, P target represents the location of the target point, c1 is the reward coefficient when approaching the target point, and c2 is the positive reward obtained when the mobile robot reaches the target point; when the range between the mobile robot and the target point is less than a certain threshold d goal When , it is considered to have reached the target point and a large reward value is given;
[0023] Safety obstacle avoidance part reward c The definition is as follows:
[0024]
[0025] Among them, P obs,i It is represented as the position information of the ith obstacle, c3 is the reward coefficient of the safe obstacle avoidance part, and c4 is the penalty value obtained after the collision; when the range between the mobile robot and the obstacle is less than a certain threshold d collision When , it is considered to have collided with the obstacle and a large penalty value is given;
[0026] Assisted Exploration Reward s The definition is as follows:
[0027]
[0028] Among them, S new is the newly added area scanned by the mobile robot radar after a single action is executed, S lidar It represents the field of view that the mobile robot can scan; c5 is the reward coefficient of the auxiliary exploration part, which is used to adjust the weight of the auxiliary reward to balance the exploration and goal-oriented rewards.
[0029] In a second aspect, the embodiments of the present disclosure further provide a mobile robot autonomous navigation system for cross-regional map-free scenarios, including:
[0030] Sensor unit, used to observe the mobile robot time step t to obtain state information s t , the state information s t Input into the Actor network;
[0031] Action selection unit, the Actor network sends the state information s t After forward propagation, according to the strategy π θ (s t ) Output an action selection a t , and send the action selection a t Give the mobile robot, the mobile robot performs the action and selects a t After that, the environment changes to a new state s t+1 , where θ is the network parameter of Actor;
[0032] External reward and intrinsic reward calculation unit, calculate the external reward r e,t and intrinsic reward r i,t , for each moment t, the extrinsic reward r e,t is defined as r e,t =r g +r c +r s , where r g is the goal-oriented part reward, r c Partial reward for safe obstacle avoidance, r s To assist exploration rewards; intrinsic rewards i,t The state information s t The input is generated by the intrinsic reward mechanism structure, where the intrinsic reward mechanism structure consists of four parts: feature extraction network, comparator network, episodic memory library and intrinsic reward estimation model;
[0033] The Actor network and Critic network training units will obtain the four-tuple <s t ,a t ,r t ,s t+1 >Store in the experience replay unit, the Critic network is based on the four-tuple in the experience replay unit <s t ,a t ,r t ,s t+1 > Approximate action value function Q φ (s t ,a t ) trains the Actor network and the Critic network and updates its parameters, where φ is the Critic network parameter, r t =r e,t +r i,t ;
[0034] The control execution unit uses the trained Actor network and Critic network models to guide the mobile robot to perform real-time cross-regional map-free autonomous navigation tasks.
[0035] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned mobile robot autonomous navigation method when executing the computer program.
[0036] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned mobile robot autonomous navigation method when the program is executed by a processor.
[0037] The advantages or beneficial effects of the above technical solution include at least:
[0038] By introducing auxiliary exploration tasks and combining external rewards with intrinsic reward mechanisms, the exploration ability of the mobile robot in the local optimal area is improved while accelerating the algorithm training, guiding it to move in the direction close to the target point and get rid of the oscillation state, so as to achieve efficient exploration and path optimization of the mobile robot in an unknown environment. In addition, an intrinsic reward mechanism driven by situational curiosity is proposed, and the concept of state reachability is introduced as an exploration driving force to encourage the mobile robot to actively explore unseen areas, thereby effectively avoiding the local optimal problem and improving the training efficiency and adaptability of the algorithm under sparse reward conditions. The disclosed embodiment not only overcomes the limitations of traditional external reward sparsity, but also significantly enhances the algorithm's exploration ability, convergence efficiency and navigation task success rate, providing an innovative solution for autonomous navigation in map-free cross-regional scenarios.
[0039] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present disclosure will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0041] Figure 1 It is a flow chart of the autonomous navigation method of a mobile robot in a cross-region map-free scenario in an embodiment of the present disclosure;
[0042] Figure 2 It is a schematic diagram of the framework of the map-free cross-region navigation method of the mobile robot in the embodiment of the present disclosure;
[0043] Figure 3 A schematic diagram of auxiliary exploration reward in an embodiment of the present disclosure;
[0044] Figure 4 A schematic diagram of an intrinsic reward mechanism of a situation-based curiosity mechanism in an embodiment of the present disclosure;
[0045] Figure 5 is a schematic diagram of the structure of a feature extraction network in an embodiment of the present disclosure;
[0046] Figure 6 is a schematic diagram of the structure of a comparator network in an embodiment of the present disclosure;
[0047] Figure 7 The block diagram is a mobile robot autonomous navigation system for cross-regional map-free scenarios in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0048] In the following, only some exemplary embodiments are briefly described. As those skilled in the art will appreciate, the described embodiments may be modified in various ways without departing from the spirit or scope of the present disclosure. Therefore, the drawings and descriptions are considered to be exemplary and non-restrictive in nature.
[0049] The present disclosure provides a method for autonomous navigation of a mobile robot in a cross-regional map-free scenario. Figure 1 As shown, Figure 1 The flowchart of the mobile robot autonomous navigation method includes the following steps:
[0050] S10, the mobile robot observes and obtains state information s at time step t t , the state information s t Input into the Actor network;
[0051] Figure 2 The framework diagram of the mapless cross-region navigation method for mobile robots integrates a storage unit and a sensor unit. The sensor unit includes an inertial odometer, a single-line two-dimensional laser radar, and a GPS, etc., which are used to obtain the target point orientation, its own position and posture, and local observable environment information. It can collect the observation information, position information, and collision detection information of the mobile robot in real time, and provide accurate perception input for navigation decisions. At time step t, the mobile robot observes the time step state information s through the sensor unit. t In this embodiment, the status information is defined as follows:
[0052] s t =(s lidar ,s target ), where s lidar is the single-line two-dimensional laser radar data, s target is the orientation information of the target point, which is calculated by the mobile robot's own position and the target point orientation information.
[0053] S20, the Actor network sends the state information s t After forward propagation, according to the strategy π θ (s t ) Output an action selection a t , and send the action selection a t Give the mobile robot, the mobile robot performs the action and selects a t After that, the environment changes to a new state s t+1 , where θ is the network parameter of Actor;
[0054] In this embodiment, action selection a t It can be defined as the velocity v and angular velocity ω of the mobile robot: t =(v,ω), in order to ensure the safe navigation of the mobile robot, an upper limit is set on the output speed and angular velocity.
[0055] S30. Calculate the external reward r e,t and intrinsic reward r i,t , for each time step t, the extrinsic reward r e,t is defined as r e,t =r g +r c +r s , where r gis the goal-oriented part reward, r c Partial reward for safe obstacle avoidance, r s To assist exploration rewards; intrinsic rewards i,t The state information s t The input is generated by the intrinsic reward mechanism structure, where the intrinsic reward mechanism structure consists of four parts: feature extraction network, comparator network, episodic memory library and intrinsic reward estimation model;
[0056] In this embodiment, the goal-oriented part of the reward can be defined as follows:
[0057]
[0058] Among them, P t-1 and P t is the position of the mobile robot at the last moment and time step, P target represents the location of the target point, c1 is the reward coefficient when approaching the target point, and c2 is the positive reward obtained when the mobile robot reaches the target point. goal When , we believe that it has reached the target point and give it a larger reward value.
[0059] In this embodiment, the safe obstacle avoidance part reward r c It can be defined as follows:
[0060]
[0061] Among them, P obs,i It is represented as the position information of the ith obstacle, c3 is the reward coefficient of the safe obstacle avoidance part, and c4 is the penalty value obtained after the collision. collision When , we consider that it collides with an obstacle and give it a larger penalty value.
[0062] In this embodiment, the auxiliary exploration reward r s It can be defined as follows:
[0063]
[0064] Among them, S new is the newly added area scanned by the mobile robot radar after a single action is executed, S lidar Represents the area of the field of view that the mobile robot can scan. c5 is the reward coefficient of the auxiliary exploration part, which is used to adjust the weight of the auxiliary reward to balance the exploration and goal-oriented rewards, so as to avoid deviation from the main navigation task due to excessive exploration, such as Figure 3 Shown is a schematic diagram of assisted exploration rewards.
[0065] In this embodiment, Figure 4 The figure shows the schematic diagram of the intrinsic reward mechanism, which converts the current state information s of the mobile robot into t As input, generate the intrinsic reward r i,t At each time step t, the mobile robot observes the current laser radar scanning data and the target point position information s t =(s lidar ,s goal ); these data are processed separately through two independent network branches, and their features are integrated to form a feature extraction network φ. The extracted state features are expressed as φ(s t ), specific network structure, such as Figure 5 The figure shows the structure of the feature extraction network. Furthermore, a comparator network is designed to calculate the reachability R(s) of the state information observed by the mobile robot at two different times within k steps. i ,s j )=f(φ(s i ),φ(s j )), by training a multi-layer fully connected neural network f to predict a pair of states (s i ,s j ) is the accessibility of the mobile robot. t ), a buffer named episodic memory bank M is designed in the intrinsic reward mechanism, with a capacity of N. The memory bank is initialized before the start of the round and is updated to M at time step t. t-1 Specifically, at each time step, the current observation s is calculated using the feature extraction network and the comparator network. t and episodic memory bank M t-1 The reachability score R(s) between the elements in t ,M t-1 )=max{f(φ(s t ),m)|m∈M t-1}∈[0,1]. This score is the maximum value obtained by comparing the current observation value with the feature state values in the memory bank to avoid completely ignoring the contribution of old memory elements. If the reachability score is lower than the threshold, that is, it is far from the historical state of the mobile robot, the current observation value is embedded in the context memory bank. The specific formula is as follows:
[0066]
[0067] Where τ is the accessibility threshold parameter, and the episodic memory bank will be reset after each round.
[0068] Furthermore, when the accessibility score R(s t ,M t-1)<τ, it means that the current state cannot be reached from any state in the memory bank through k steps, which means that the state has enough "novelty" and is therefore added to the episodic memory bank. Therefore, the intrinsic reward estimation model is defined as:
[0069]
[0070] Among them, r i,t is the intrinsic reward obtained by the mobile robot at time t, and β is the intrinsic reward coefficient.
[0071] S40, the obtained quaternion <s t ,a t ,r t ,s t+1 >Store in the experience replay unit, the Critic network is based on the four-tuple in the experience replay unit <s t ,a t ,r t ,s t+1 > Approximate action value function Q φ (s t ,a t ) trains the Actor network and the Critic network and updates its parameters, where φ is the Critic network parameter, r t =r e,t +r i,t ;
[0072] In this embodiment, the data generated by the above state conversion process is t ,a t ,r t ,s t+1 > As an experience data, it is stored in the experience playback unit of the algorithm, where r t =r e,t +r i,t , and enter the next state transition process.
[0073] As an optional embodiment, the network structures of the Actor network and the Critic network are composed of a 512*512*512*2 fully connected layer and an activation function Relu, the parameter optimizer is Adam, and the experience playback unit capacity is 2×10 6 , respectively calculate the network loss function to update the corresponding parameters θ, φ. A batch of experience data is randomly selected from the experience replay unit and sent to the Critic and Actor networks for training. The Critic network approximates the action value function Q according to the data in the experience replay unit. φ (s t ,a t), where φ is the parameter of the Critic network. During a training process, the parameters θ and φ of the Actor and Critic networks are also updated by the Adam optimizer to ensure that the action decision output by the mobile robot tends to the optimal strategy.
[0074] Repeat the above steps until a better strategy is trained, and save the model parameters of the Actor network and the Cr itic network.
[0075] S50. Use the trained Actor network and Critic network models to guide the mobile robot to perform real-time cross-regional map-free autonomous navigation tasks.
[0076] In this embodiment, at the beginning of each round of training, the starting point of the mobile robot is randomly set in the environment, the size and structure of different areas in the room are randomly set, and the target point position is randomly set. At each time step, the mobile robot will select control instructions within the action space to move. If the mobile robot collides with an obstacle or boundary, the round ends with failure and a new round of training begins; if the mobile robot reaches the target point and obtains a large reward, the round ends with successful arrival and a new round of training begins. After each training session, the reward value, number of execution steps, and training success rate obtained by the mobile robot in that round are recorded. When the above indicators tend to be stable, it proves that the algorithm converges, and the mobile robot can make autonomous decisions in a cross-regional scenario without a map, ensuring the safety and stability of the navigation task. The algorithm and network model parameters are deployed to the mobile robot platform to realize its autonomous navigation task in a cross-regional scenario without a map.
[0077] As an optional embodiment, the state information s t Contains the scanning information of the single-line two-dimensional laser radar, the mobile robot's target point position information and the mobile robot's own position information.
[0078] As an optional embodiment, the action selection a t Includes the speed and angular velocity of the mobile robot.
[0079] As an optional embodiment, the intrinsic reward r i,t The generation steps are as follows:
[0080] The mobile robot observes the current lidar scanning data and the state information s of the target point at each time step t. t , after being processed by the feature extraction network, the state feature φ(s t );
[0081] like Figure 6 The diagram shows a comparator network used to evaluate the state characteristic s.i and j Reachability R(s) within k steps i ,s j )=f(φ(s i ),φ(s j )), by training a multi-layer fully connected neural network f to predict the state pair (s i ,s j ) and uses the Sigmoid activation function to map the output to [0,1], representing the accessibility score R(s i ,s j ), a score close to 1 indicates that the states are relatively “close” or “reachable”, and a score close to 0 indicates that the two states are far away or unreachable.
[0082] At the current time step t, the mobile robot maintains a context memory M t-1 , used to memorize the historical observation information of the mobile robot; the current state feature φ(s t ) and episodic memory bank M t-1 The reachability score R(s) between the elements in t ,M t-1 ), when the accessibility score is lower than the threshold τ, φ(s t ) is embedded in the episodic memory bank and defines the intrinsic reward r i,t =β·R(s t ,M t-1 ), where β is the intrinsic reward coefficient.
[0083] As an optional embodiment, before the trained Actor network and Critic network models are used to guide the mobile robot to perform a real-time cross-region map-free autonomous navigation task, the following steps are also included:
[0084] Evaluation indicators are used to evaluate the safety and reliability of the mobile robot in the map-free cross-region navigation task, wherein the evaluation indicators include at least one of the following indicators: navigation success rate, collision rate, timeout rate or execution step number.
[0085] The disclosed embodiment also provides a mobile robot autonomous navigation system 100 for cross-regional map-free scenarios, such as Figure 7 As shown, including:
[0086] Sensor unit 1 is used to observe the mobile robot at time step t to obtain state information s t , the state information s t Input into the Actor network;
[0087] Action selection unit 2, the Actor network sends the state information s tAfter forward propagation, according to the strategy π θ (s t ) Output an action selection a t , and send the action selection a t Give the mobile robot, the mobile robot performs the action and selects a t After that, the environment changes to a new state s t+1 , where θ is the network parameter of Actor;
[0088] External reward and intrinsic reward calculation unit 3, calculate the external reward r e,t and intrinsic reward r i,t , for each moment t, the extrinsic reward r e,t is defined as r e,t =r g +r c +r s , where r g is the goal-oriented part reward, r c Partial reward for safe obstacle avoidance, r s To assist exploration rewards; intrinsic rewards i,t The state information s t The input is generated by the intrinsic reward mechanism structure, where the intrinsic reward mechanism structure consists of four parts: feature extraction network, comparator network, episodic memory library and intrinsic reward estimation model;
[0089] Actor network and Critic network training unit 4, the obtained four-tuple <s t ,a t ,r t ,s t+1 >Store in the experience replay unit, the Critic network is based on the four-tuple in the experience replay unit <s t ,a t ,r t ,s t+1 > Approximate action value function Q φ (s t ,a t ) trains the Actor network and the Critic network and updates its parameters, where φ is the Critic network parameter, r t =r e,t +r i,t ;
[0090] The control execution unit 5 uses the trained Actor network and Critic network models to guide the mobile robot to perform real-time cross-region map-free autonomous navigation tasks.
[0091] In the absence of any contradiction, the above-mentioned units in the system 100 of the embodiment of the present disclosure can implement any implementation of the above-mentioned corresponding methods.
[0092] The disclosed embodiment also provides an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above-mentioned cross-regional map-free mobile robot autonomous navigation method. The electronic device can be provided as a terminal, a server, or other forms of equipment.
[0093] The present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned mobile robot autonomous navigation method for cross-regional map-free scenarios is implemented. The computer-readable storage medium may be a non-volatile computer-readable storage medium.
[0094] Those skilled in the art will understand that in the above-mentioned cross-regional map-free scene mobile robot autonomous navigation method and system in the specific implementation mode, the writing order of each step does not mean a strict execution order and constitutes any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0095] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0096] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0097] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of various changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A mobile robot autonomous navigation method for cross-regional map-free scenes, characterized in that: The steps include: The mobile robot observes the state information s at time step t t , the state information s t Input into the Actor network; The Actor network sends the state information s t After forward propagation, according to the strategy π θ (s t ) Output an action selection a t , and send the action selection a t Give the mobile robot, the mobile robot performs the action and selects a t After that, the environment changes to a new state s t+1 , where θ is the network parameter of Actor; Calculate the extrinsic reward r e,t and intrinsic reward r i,t , for each moment t, the extrinsic reward r e,t is defined as r e,t =r g +r c +r s , where r g is the goal-oriented part reward, r c Partial reward for safe obstacle avoidance, r s To assist exploration rewards; intrinsic rewards i,t The state information s t The input is generated by the intrinsic reward mechanism structure, where the intrinsic reward mechanism structure consists of four parts: feature extraction network, comparator network, episodic memory library and intrinsic reward estimation model; The resulting quaternion <s t ,a t ,r t ,s t+1 >Store in the experience replay unit, the Critic network is based on the four-tuple in the experience replay unit <s t ,a t ,r t ,s t+1 > Approximate action value function Q φ (s t ,a t ) trains the Actor network and the Critic network and updates its parameters, where φ is the Critic network parameter, r t =r e,t +r i,t ; The trained Actor network and Critic network models are used to guide the mobile robot to perform real-time cross-regional map-free autonomous navigation tasks.
2. The mobile robot autonomous navigation method according to claim 1, characterized in that: The status information t Contains the scanning information of the single-line two-dimensional laser radar, the mobile robot's target point position information and the mobile robot's own position information.
3. The mobile robot autonomous navigation method according to claim 1 or 2, characterized in that: The action selects a t Includes the speed and angular velocity of the mobile robot.
4. The mobile robot autonomous navigation method according to claim 1 or 2, characterized in that: The intrinsic reward r i,t The generation steps are as follows: The mobile robot observes the current lidar scanning data and the state information s of the target point at each time step t. t , after being processed by the feature extraction network, the state feature φ(s t ); The comparator network is used to evaluate the state feature s i and j The reachability R(s) within k steps i ,s j )=f(φ(s i ),φ(s j )), by training a multi-layer fully connected neural network f to predict the state pair (s i ,s j ) and uses the Sigmoid activation function to map the output to [0,1], representing the reachability score R(s i ,s j ), a score close to 1 indicates that the states are relatively "close" or "reachable", and a score close to 0 indicates that the two states are far away or unreachable; At the current time step t, the mobile robot maintains a context memory M t-1 , used to memorize the historical observation information of the mobile robot; the current state feature φ(s t ) and episodic memory bank M t-1 The reachability score R(s) between the elements in t ,M t-1 ), when the accessibility score is lower than the threshold τ, φ(s t ) is embedded in the episodic memory bank and defines the intrinsic reward r i,t =β·R(s t ,M t-1 ), where β is the intrinsic reward coefficient.
5. The mobile robot autonomous navigation method according to claim 1 or 2, characterized in that: The network structures of the Actor network and the Critic network are composed of a 512*512*512*2 fully connected layer and an activation function Relu. The parameter optimizer is Adam, and the capacity of the experience playback unit is 2×10 6 , respectively calculate the network loss function to update the corresponding parameters θ, φ.
6. The mobile robot autonomous navigation method according to claim 1 or 2, characterized in that: Before the trained Actor network and Critic network models are used to guide the mobile robot to perform real-time cross-region map-free autonomous navigation tasks, the following steps are also included: Evaluation indicators are used to evaluate the safety and reliability of the mobile robot in the map-free cross-region navigation task, wherein the evaluation indicators include at least one of the following indicators: navigation success rate, collision rate, timeout rate or execution step number.
7. The mobile robot autonomous navigation method according to claim 1 or 2, characterized in that: The goal-oriented part of the reward is defined as follows: Among them, P t-1 and P t is the position of the mobile robot at the last moment and time step, P target represents the location of the target point, c1 is the reward coefficient when approaching the target point, and c2 is the positive reward obtained when the mobile robot reaches the target point; when the range between the mobile robot and the target point is less than a certain threshold d goal When , it is considered to have reached the target point and a large reward value is given; Safety obstacle avoidance part reward c The definition is as follows: Among them, P obs,i It is represented as the position information of the ith obstacle, c3 is the reward coefficient of the safe obstacle avoidance part, and c4 is the penalty value obtained after the collision; when the range between the mobile robot and the obstacle is less than a certain threshold d collision When , it is considered to have collided with the obstacle and a large penalty value is given; Assisted Exploration Reward s The definition is as follows: Among them, S new is the newly added area scanned by the mobile robot radar after a single action is executed, S lidar It represents the field of view that the mobile robot can scan; c5 is the reward coefficient of the auxiliary exploration part, which is used to adjust the weight of the auxiliary reward to balance the exploration and goal-oriented rewards.
8. The autonomous navigation system of mobile robots in cross-regional map-free scenarios is characterized by: include: Sensor unit, used to observe the mobile robot time step t to obtain state information s t , the state information s t Input into the Actor network; Action selection unit, the Actor network sends the state information s t After forward propagation, according to the strategy π θ (s t ) Output an action selection a t , and send the action selection a t Give the mobile robot, the mobile robot performs the action and selects a t After that, the environment changes to a new state s t+1 , where θ is the network parameter of Actor; External reward and intrinsic reward calculation unit, calculate the external reward r e,t and intrinsic reward r i,t , for each moment t, the extrinsic reward r e,t is defined as r e,t =r g +r c +r s , where r g is the goal-oriented part reward, r c Partial reward for safe obstacle avoidance, r s To assist exploration rewards; intrinsic rewards i,t The state information s t The input is generated by the intrinsic reward mechanism structure, where the intrinsic reward mechanism structure consists of four parts: feature extraction network, comparator network, episodic memory library and intrinsic reward estimation model; The Actor network and Critic network training units will obtain the four-tuple <s t ,a t ,r t ,s t+1 >Store in the experience replay unit, the Critic network is based on the four-tuple in the experience replay unit <s t ,a t ,r t ,s t+1 > Approximate action value function Q φ (s t ,a t ) trains the Actor network and the Critic network and updates its parameters, where φ is the Critic network parameter, r t =r e,t +r i,t ; The control execution unit uses the trained Actor network and Critic network models to guide the mobile robot to perform real-time cross-regional map-free autonomous navigation tasks.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the mobile robot autonomous navigation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the mobile robot autonomous navigation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Visual target navigation method and device
CN114413910A
Multi-agent target collaborative search method and system
CN115952736A
Award obtaining method based on scene memory deep Q network
CN117852618A
Methods and systems for training an untrained policy network of an autonomous agent to generate a dual-action policy network
GB202310954D0
Continuous control with deep reinforcement learning
US20170024643A1