Navigation method, device, robot and storage medium implemented based on a knowledge-guided mapless navigation model

By introducing knowledge systems and multiple action rules into the DDPG algorithm, the navigation model is optimized, and the problem of reduced navigation accuracy in complex and dynamic environments is solved, achieving more efficient and stable navigation effects.

CN119779312BActive Publication Date: 2025-07-25SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510028410.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-25
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing map-free navigation method based on DDPG algorithm has reduced navigation accuracy in complex and dynamic environments, low learning efficiency and unstable, making it difficult to adapt to unknown environments.

Method used

The knowledge system is introduced to train the deep deterministic strategy gradient algorithm model, combining multiple action rules and generalizers, and optimize the navigation model by integrating guided action instructions and policy action instructions.

Benefits of technology

Improve the adaptability and learning efficiency of navigation models in complex and dynamic environments, ensuring the accuracy and stability of navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119779312B_ABST
    Figure CN119779312B_ABST
Patent Text Reader

Abstract

The present invention provides a navigation method, device, robot and storage medium implemented based on a knowledge-guided mapless navigation model. The mapless navigation model is obtained by training a pre-constructed DDPG algorithm model based on multiple action rules in a knowledge system, and has stronger generalization ability compared with using only the DDPG algorithm for navigation. During the training process, fusing the guiding action instructions and the policy action instructions can reduce the randomness of action selection of the DDPG algorithm model, so as to quickly obtain valuable data, improve the learning efficiency, and in an environment with sparse rewards, the mobile robot can interact with the environment under the guidance of knowledge instead of randomly, avoiding falling into local optima and ensuring easy convergence in an environment with sparse rewards; and inputting both the policy action instructions and the combined action instructions obtained by fusion into a preset loss function reduces the uncertainty of the corresponding loss function when using only the DDPG algorithm, thus making the learning process more stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot navigation, and in particular, to a navigation method, device, robot, and storage medium implemented based on a knowledge-guided mapless navigation model. Background Art

[0002] With the continuous development of mobile robot technology, navigation technology has become the key to realizing the autonomous movement of robots. Traditional robot navigation mostly adopts SLAM (Simultaneous Localization and Mapping) technology. The SLAM technology creates an environmental map in real time through sensor data (such as lidar, camera, etc.), and at the same time determines the position of the robot in the map, so as to achieve precise navigation. However, in practical applications, this method faces the following challenges:

[0003] (1) In complex and changeable scenarios, it is difficult for the SLAM technology to update the environmental map in real time, resulting in a decrease in navigation accuracy;

[0004] (2) In a dynamic environment, it is difficult for the SLAM technology to handle moving objects and is prone to positioning errors;

[0005] (3) In an unknown environment where it is impossible to pre-build a map, it is difficult for the SLAM technology to achieve efficient navigation.

[0006] Therefore, researchers began to explore navigation strategies that do not rely on map information, and thus proposed a mapless navigation method based on reinforcement learning (RL). For example, the DDPG (Deep Deterministic Policy Gradient) algorithm is adopted to enable a mobile robot to learn an optimal action strategy through interaction with the environment to maximize the cumulative reward. The advantages of this method are: it does not need to pre-build or rely on an environmental map and is applicable to unknown environments; it makes decisions directly based on sensor inputs (such as distance, speed, obstacle information, etc.) and immediate feedback to achieve end-to-end navigation. The navigation model obtained solely by the DDPG algorithm performs well in simple navigation scenarios similar to the virtual environment during training, but it is difficult to adapt to complex dynamic navigation scenarios. Summary of the Invention

[0007] The purpose of the present invention is to provide a navigation method, device, robot, and storage medium implemented based on a knowledge-guided mapless navigation model to improve the problems existing in the prior art.

[0008] The embodiments of the present invention can be implemented as follows:

[0009] In a first aspect, the present invention provides a navigation method implemented based on a mapless navigation model guided by knowledge, including:

[0010] Obtain the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment;

[0011] Convert the environmental state data and relative position data at the current moment and the action instruction at the previous moment into a current state vector;

[0012] Input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; the mapless navigation model is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system;

[0013] Use the action instruction at the current moment to control the mobile robot to move forward towards the target point.

[0014] In an optional implementation manner, the mapless navigation model is trained through the following method:

[0015] Based on the knowledge system, the deep deterministic policy gradient algorithm model, and the generalizer, control the intelligent agent to continuously interact with a preset virtual environment, and add each piece of interaction data generated by the interaction to an initially empty experience database;

[0016] When the number of interaction data in the experience database reaches Extract pieces of interaction data randomly from the experience database; is a positive integer;

[0017] Based on the pieces of interaction data and multiple action rules in the knowledge system, update the parameters of the deep deterministic policy gradient algorithm model and the generalizer;

[0018] Based on the knowledge system, the updated deep deterministic policy gradient algorithm model, and the updated generalizer, control the intelligent agent to continuously interact with the virtual environment, and add each piece of interaction data generated by the interaction to the experience database;

[0019] Judge whether the number of interaction rounds between the intelligent agent and the virtual environment reaches a preset value;

[0020] If not, return to the step of extracting pieces of interaction data from the experience database until the number of interaction rounds between the intelligent agent and the virtual environment reaches the preset value;

[0021] If so, control the agent to stop interacting with the virtual environment, and use the current Deep Deterministic Policy Gradient (DDPG) algorithm model as the mapless navigation model.

[0022] In an alternative embodiment, the step of updating the parameters of the DDPG algorithm model and the generalizer based on the multiple pieces of interaction data and multiple action rules in the knowledge system includes:

[0023] For the th piece of interaction data, obtain a state vector from the th piece of interaction data, and match the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction; where ;

[0024] Input the state vector into the DDPG algorithm model to obtain a policy action instruction and a state value;

[0025] Input the guiding action instruction and the policy action instruction into the generalizer for fusion to obtain a comprehensive action instruction;

[0026] Input the current fusion parameters of the generalizer, the state value, the policy action instruction, and the comprehensive action instruction into a preset loss function to obtain the comprehensive loss corresponding to the th piece of interaction data;

[0027] Based on the comprehensive loss corresponding to the multiple pieces of interaction data, calculate the expected loss, and update the model parameters of the DDPG algorithm model and the fusion parameters of the generalizer based on the expected loss.

[0028] In an alternative embodiment, the guiding action instruction includes a guiding linear velocity and a guiding angular velocity;

[0029] The multiple action rules of the knowledge system include at least one precise action rule, at least one fuzzy action rule, and at least one hybrid action rule, and each action rule has associated position data;

[0030] The precise action rule includes at least one precise condition and conclusion data, the fuzzy action rule includes at least one fuzzy condition and conclusion data, and the hybrid action rule includes at least one precise condition, at least one hybrid condition, and conclusion data; the conclusion data includes a preset linear velocity and a preset angular velocity;

[0031] The step of matching the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction includes:

[0032] For each of the precise action rules, obtain at least one vector value from the state vector based on the position data associated with the precise action rule; if the at least one vector value makes each precise condition of the precise action rule hold, then use the preset linear velocity and the preset angular velocity in the precise action rule as the target linear velocity and the target angular velocity respectively;

[0033] For each of the fuzzy action rules, obtain at least one vector value from the state vector based on the position data associated with the fuzzy action rule; calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the fuzzy action rule, multiply the obtained membership degrees to get the activation intensity, and multiply the obtained activation intensity by the preset linear velocity and the preset angular velocity in the fuzzy action rule respectively to get the target linear velocity and the target angular velocity;

[0034] For each of the hybrid action rules, obtain at least one vector value from the state vector based on the position data associated with the hybrid action rule; if the at least one vector value makes each precise condition of the hybrid action rule hold, then calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the hybrid action rule, multiply the obtained membership degrees to get the activation intensity, and multiply the obtained activation intensity by the preset linear velocity and the preset angular velocity in the hybrid action rule respectively to get the target linear velocity and the target angular velocity;

[0035] Add all the target linear velocities to obtain the guiding linear velocity, and intersect all the target angular velocities to obtain the guiding angular velocity.

[0036] In an alternative embodiment, the guiding action instruction includes a guiding linear velocity and a guiding angular velocity; the strategy action instruction includes a predicted linear velocity and a predicted angular velocity; the comprehensive action instruction includes a comprehensive linear velocity and a comprehensive angular velocity;

[0037] The step of inputting the guiding action instruction and the strategy action instruction into the generalizer for fusion to obtain a comprehensive action instruction includes:

[0038] Input the guiding linear velocity, the guiding angular velocity, the predicted linear velocity, and the predicted angular velocity into the generalizer;

[0039] Use the current fusion parameters to perform weighted processing on the predicted linear velocity and the guiding linear velocity to obtain a pending linear velocity, and add Gaussian noise to the pending linear velocity to obtain the comprehensive linear velocity;

[0040] The predicted angular velocity and the guidance angular velocity are weighted using the current fusion parameter to obtain a to-be-determined angular velocity, and Gaussian noise is added to the to-be-determined angular velocity to obtain the comprehensive angular velocity.

[0041] In an alternative embodiment, the policy action instruction includes a predicted linear velocity and a predicted angular velocity;

[0042] The calculation formula of the preset loss function is:

[0043]

[0044]

[0045]

[0046] Wherein, represents the comprehensive loss, represents the linear velocity loss, represents the angular velocity loss; is a fixed parameter indicating the credibility of the knowledge system; represents the state value, is the current fusion parameter of the generalizer, represents the predicted linear velocity, represents the comprehensive linear velocity, represents the predicted angular velocity, represents the comprehensive angular velocity.

[0047] In a second aspect, the present invention provides a navigation device implemented based on a knowledge-guided mapless navigation model, including:

[0048] A state acquisition module for acquiring environmental state data of the mobile robot at the current moment, relative position data between the current position and the target point, and the action instruction at the previous moment;

[0049] The state acquisition module is further configured to convert the environmental state data and relative position data at the current moment and the action instruction at the previous moment into a current state vector;

[0050] A prediction module for inputting the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; the mapless navigation model is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system;

[0051] A control module for controlling the mobile robot to move forward towards the target point by using the linear velocity and angular velocity in the action instruction at the current moment.

[0052] In a third aspect, the present invention provides a mobile robot, including: a memory and a processor, where the memory stores a software program, and when the mobile robot runs, the processor executes the software program to implement the navigation method as described in the foregoing first aspect.

[0053] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the navigation method as described in the foregoing first aspect.

[0054] Compared with the prior art, the embodiments of the present invention provide a navigation method, device, robot, and storage medium implemented based on a knowledge-guided mapless navigation model. The navigation method is as follows: First, obtain the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment; then convert the environmental state data and relative position data at the current moment and the action instruction at the previous moment into a current state vector; then input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; finally, use the action instruction at the current moment to control the mobile robot to move forward towards the target point. The mapless navigation model of the present invention is trained based on multiple action rules in a preset knowledge system for a pre-constructed deep deterministic policy gradient algorithm model, rather than simply using the DDPG algorithm. This enables the mapless navigation model to not only adapt to navigation scenarios similar to the training environment but also complex dynamic navigation scenarios, ensuring the accuracy of navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is a schematic flowchart of a navigation method provided by an embodiment of the present invention.

[0057] Figure 2 It is a schematic diagram of the training process of the mapless navigation model provided by an embodiment of the present invention.

[0058] Figure 3 It is a schematic diagram of a scenario where an intelligent agent interacts with a virtual environment provided by an embodiment of the present invention.

[0059] Figure 4A schematic structural diagram of a navigation device provided by an embodiment of the present invention.

[0060] Figure 5 A schematic structural diagram of a mobile robot provided by an embodiment of the present invention. Detailed implementation manners

[0061] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the accompanying drawings herein may be arranged and designed in a variety of different configurations.

[0062] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0063] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.

[0064] It should be noted that the features in the embodiments of the present invention may be combined with each other without conflict.

[0065] In the prior art, researchers have proposed a mapless navigation method based on reinforcement learning. For example, the DDPG algorithm is used to enable a mobile robot to learn an optimal action strategy through interaction with the environment to maximize the cumulative reward. The advantages of this method are as follows: it does not require prior construction or dependence on an environmental map and is applicable to unknown environments; it directly makes decisions based on sensor inputs (such as distance, speed, obstacle information, etc.) and immediate feedback to achieve end-to-end navigation. During the training process, the navigation model obtained solely by using the DDPG algorithm has the following problems:

[0066] (1) The learning efficiency is low. Especially in a sparse reward environment, it is difficult for the algorithm to quickly converge to an effective strategy;

[0067] (2) The stability of the learning process is poor and it is easily affected by noise and interference.

[0068] Moreover, the navigation model obtained solely by using the DDPG algorithm performs well in simple navigation scenarios similar to the virtual environment during training, but it is difficult to adapt to complex dynamic navigation scenarios.

[0069] Based on the discovery of the above technical problems, the inventors have proposed the following technical solutions through creative labor to solve or improve the above problems. It should be noted that the defects existing in the above solutions in the prior art are all the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the embodiments of the present application below for the above problems should both be the contributions made by the inventors to the present application during the invention creation process, and should not be understood as the technical content known to those skilled in the art.

[0070] In view of the deficiencies in the prior art, the inventors have proposed a mapless navigation model. In the training process of this mapless navigation model, a knowledge system is introduced as supervision, enabling the model to converge quickly during training and improving the learning efficiency. The following will be described in detail through embodiments and in conjunction with the accompanying drawings.

[0071] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a navigation method provided by an embodiment of the present invention. This navigation method is implemented based on a knowledge-guided mapless navigation model, and the execution subject of this navigation method is a mobile robot. This navigation method includes the following steps S201 to S204.

[0072] S201. Obtain the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment.

[0073] In this embodiment, the environmental state data may include the data collected and processed by various sensors carried by the mobile robot, such as obstacle point cloud data, acceleration data, image data, etc. The relative position data between the current position and the target point may include the straight-line distance and the deviation angle. The deviation angle is the angle between the line connecting the current position and the target point and the forward direction of the mobile robot. The action instruction at the previous moment includes the linear velocity and the angular velocity, where the linear velocity is the traveling speed of the mobile robot and the angular velocity is the rotation speed of the mobile robot.

[0074] S202. Convert the environmental state data, the relative position data at the current moment, and the action instruction at the previous moment into a current state vector.

[0075] Optionally, in the environmental state data, the distance data between the mobile robot and the obstacle can be determined using the obstacle point cloud data, and the obstacle determination data can be obtained by performing image recognition on the image data. This obstacle determination data indicates whether there are obstacles around the mobile robot. Therefore, the distance data, the acceleration data, the obstacle determination data, the relative position data, and the linear velocity and angular velocity at the previous moment can be converted into a current state vector.

[0076] S203. Input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment.

[0077] In this embodiment, the mapless navigation model is trained based on multiple action rules in a preset knowledge system for a pre - constructed deep deterministic policy gradient algorithm model. And an action rule reflects the linear velocity and angular velocity that a robot should adopt under specific conditions.

[0078] S204. Use the action instruction at the current moment to control the mobile robot to move forward towards the target point.

[0079] In this embodiment, the linear velocity and angular velocity in the action instruction at the current moment can be used to control the mobile robot to move forward towards the target point.

[0080] That is, starting from the starting point, the mobile robot repeatedly executes the above - mentioned steps S201 - S204 until it reaches the target point (i.e., the end point), which means completing a mapless navigation from the starting point to the end point.

[0081] The navigation method provided by the embodiment of the present invention first converts the acquired environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment into a current state vector; then inputs the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; finally, uses the action instruction at the current moment to control the mobile robot to move forward towards the target point. The mapless navigation model of the present invention is trained based on multiple action rules in a preset knowledge system for a pre - constructed deep deterministic policy gradient algorithm model, rather than simply using the DDPG algorithm. This enables the mapless navigation model to not only adapt to navigation scenarios similar to the training environment but also adapt to complex dynamic navigation scenarios, ensuring the accuracy of navigation.

[0082] The following introduces the training process of the mapless navigation model. This training process can be implemented by a computing device, such as a laptop, a tablet computer, a desktop computer, a server, etc.

[0083] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the training process of the mapless navigation model provided by the embodiment of the present invention. This training process may include the following steps S101 - S107.

[0084] S101. Based on the knowledge system, the deep deterministic policy gradient algorithm model, and the generalizer, control the intelligent agent to continuously interact with a preset virtual environment, and add each piece of interaction data generated by the interaction to an initially empty experience database.

[0085] In this embodiment, the experience database is initially empty, and the agent is a virtual robot that simulates a mobile robot for a computing device. During the continuous interaction between the agent and the preset virtual environment, an interaction round is as follows: in the virtual environment, the agent moves from the starting point towards the target point until it reaches the target point or collides with an obstacle in the virtual environment. And in each interaction round, multiple pieces of interaction data are generated and added to the experience database. An interaction data in the experience database is represented as:

[0086]

[0087] Among them, represents the state vector at the -th moment, represents the mixed action instruction output at the -th moment, represents the reward at the -th moment, represents the state vector at the -th moment. The state data includes environmental state data (such as obstacle point cloud data, acceleration data, image data, etc.), relative position data (linear distance and deflection angle) between the current position and the target point, and the mixed action instruction at the previous moment.

[0088] S102. When the number of interaction data in the experience database reaches , randomly sample pieces of interaction data from the experience database.

[0089] In this embodiment, is a positive integer and the size of can be set by itself. For example, can be 128 or 256. This example is only for illustration and is not limited here.

[0090] Taking as an example, when the number of interaction data in the experience database reaches 256, interaction data can be sampled and the deep deterministic policy gradient algorithm model and the generalizer can be trained and updated through the following step S103. Moreover, while training and updating, the agent and the virtual environment are still continuously interacting.

[0091] S103. Based on pieces of interaction data and multiple action rules in the knowledge system, update the parameters of the deep deterministic policy gradient algorithm model and the generalizer.

[0092] In this embodiment, using the sampled Based on the interaction data, the deep deterministic policy gradient algorithm model and the generalizer can be trained and updated once to obtain the updated deep deterministic policy gradient algorithm model and the updated generalizer.

[0093] S104. Based on the knowledge system, the updated deep deterministic policy gradient algorithm model, and the updated generalizer, control the agent to continuously interact with the virtual environment, and add each piece of interaction data generated by the interaction to the experience database.

[0094] In this embodiment, after one training update, it is necessary to continue to control the agent to continuously interact with the virtual environment based on the knowledge system, the updated deep deterministic policy gradient algorithm model, and the updated generalizer, and add each piece of interaction data generated by the interaction to the experience database.

[0095] Among them, assuming that the environmental state data only includes obstacle point cloud data, combined with Figure 3 , in the above steps S101 and S104, the process of each interaction between the agent and the virtual environment is as follows: obtain the obstacle point cloud data of the agent at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment, then convert the obstacle point cloud data at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment into a state vector, input the state vector into the deep deterministic policy gradient algorithm model to obtain a policy action instruction, and match the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction; then input the guiding action instruction and the policy action instruction into the generalizer for fusion to obtain a comprehensive action instruction, use the comprehensive action instruction to control the agent to move forward towards the target point, and at the same time, an interaction data will be generated during this interaction process and added to the experience database.

[0096] S105. Determine whether the number of interaction rounds between the agent and the virtual environment reaches a preset value.

[0097] In this embodiment, if the number of interaction rounds between the agent and the virtual environment does not reach the preset value, directly return to execute the process of sampling pieces of interaction data from the experience database in the above step S102 until the number of interaction rounds between the agent and the virtual environment reaches the preset value. If the number of interaction rounds between the agent and the virtual environment reaches the preset value, execute the following step S106.

[0098] Optionally, the preset value can be set by itself based on the actual situation. For example, the preset value can be 10000 or 15000. This example is only for illustration and is not limited here.

[0099] S106. Control the agent to stop interacting with the virtual environment and use the current Deep Deterministic Policy Gradient (DDPG) algorithm model as the mapless navigation model.

[0100] Through steps S101 - S106, a knowledge - guided mapless navigation model is obtained, which can be used to implement the mapless navigation task of a mobile robot.

[0101] In an alternative implementation, the process of "updating the parameters of the DDPG algorithm model and the generalizer based on multiple pieces of interaction data and multiple action rules in the knowledge system" in step S103 may include the following sub - steps S1031 - S1035.

[0102] S1031. For the th piece of interaction data, obtain the state vector from the th piece of interaction data, and match the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction; where, .

[0103] In this embodiment, the state vector obtained from the th piece of interaction data is . The guiding action instruction includes a guiding linear velocity and a guiding angular velocity.

[0104] Optionally, the construction of the knowledge system involves representing knowledge with fuzzy semantics and precise semantics. The multiple action rules of the knowledge system include at least one precise action rule, at least one fuzzy action rule, and at least one hybrid action rule. Each action rule has associated position data, which characterizes the position of at least one vector value associated with the action rule in the state vector. In the knowledge system, the conditions and conclusions of the three types of action rules are as follows:

[0105] 1. Precise action rule:

[0106] A precise action rule may include at least one precise condition and conclusion data, so a precise action rule can be expressed as:

[0107]

[0108] Where, represents the th precise action rule, represents the th precise condition of the th precise action rule; represents the conclusion data of the th precise action rule, respectively represent the The preset linear velocity and preset angular velocity of the precise action rules;

[0109] 2. Fuzzy action rules:

[0110] A fuzzy action rule can include at least one fuzzy condition and conclusion data, so a fuzzy action rule can be expressed as:

[0111]

[0112] Wherein, represents the th fuzzy action rule, represents the th fuzzy action rule's rd fuzzy condition; represents the conclusion data of the th fuzzy action rule, respectively represent the preset linear velocity and preset angular velocity of the th fuzzy action rule;

[0113] 3. Hybrid action rules:

[0114] A hybrid action rule can include at least one precise condition, at least one hybrid condition, and conclusion data, so a hybrid action rule can be expressed as:

[0115]

[0116] Wherein, represents the th hybrid action rule, represents the th hybrid action rule's fuzzy condition, the th hybrid action rule's th precise condition; respectively represent the preset linear velocity and preset angular velocity of the th hybrid action rule.

[0117] Therefore, in step S1031, the process of "matching the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction" can include the following sub-steps S10311~S10314.

[0118] S10311. For each precise action rule, obtain at least one vector value from the state vector based on the position data associated with the precise action rule; if the at least one vector value makes each precise condition of the precise action rule hold, then use the preset linear velocity and preset angular velocity in the precise action rule as the target linear velocity and target angular velocity respectively.

[0119] In this embodiment, for a precise action rule of a precise condition , due to its binary nature, the activation intensity of the precise condition can be expressed as:

[0120]

[0121] Therefore, the activation intensity of the precise action rule is: .

[0122] Therefore, for each precise action rule, after obtaining at least one vector value from the state vector based on the position data associated with the precise action rule; only when the at least one vector value makes each precise condition of the precise action rule hold (i.e., each precise condition is true), the precise action rule will be regarded as activated, and the preset linear velocity and preset angular velocity in the precise action rule will be used as the target linear velocity and target angular velocity respectively.

[0123] For example, "if there is an obstacle on the left, then turn right" is a precise rule in the human cognitive system. If a precise action rule for a mobile robot is determined based on this precise rule and added to the knowledge system, it needs to be expressed as: "if there is an obstacle on the left, then and ", this example is only for illustration and is not limited here.

[0124] S10312. For each fuzzy action rule, obtain at least one vector value from the state vector based on the position data associated with the fuzzy action rule; calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the fuzzy action rule, and multiply the obtained membership degrees to get the activation intensity, and multiply the obtained activation intensity by the preset linear velocity and preset angular velocity in the fuzzy action rule respectively to get the target linear velocity and target angular velocity.

[0125] In this embodiment, for a fuzzy action rule of a precise condition , due to its fuzziness, the membership degree is used as the activation intensity of this precise condition , expressed as:

[0126]

[0127] Among them, is the vector value corresponding to the subject variable in the exact condition ; is the fuzzy set associated with the exact condition ; represents the membership degree between the fuzzy set associated with the exact condition ;

[0128] Therefore, the activation strength of the fuzzy action rule is: .

[0129] Therefore, for each fuzzy action rule, after determining the activation strength of the fuzzy action rule, multiplying the activation strength by the preset linear velocity and the preset angular velocity of the fuzzy action rule respectively can obtain a set of target linear velocity and target angular velocity.

[0130] For example, "if the obstacle is very close, then stop" is a fuzzy rule in the human cognitive system. If a fuzzy action rule for a mobile robot is determined based on this fuzzy rule and added to the knowledge system, it is first necessary to define how close is considered close: assuming that the universe of discourse of the fuzzy set is 0 to 5 m, and the distance between 0 and 0.2 m is considered very close, then the center point and the standard deviation are determined, and the Gaussian membership function adopted is defined as:

[0131]

[0132] In this way, when the distance is closer to 0, the membership degree is closer to 1. When > 0.2 m, is approximately 0.98. Therefore, the fuzzy action rule can be expressed as: "if , then and ", and this example is only for illustration and is not limited here.

[0133] S10313. For each hybrid action rule, obtain at least one vector value from the state vector based on the position data associated with the hybrid action rule; if the at least one vector value makes each exact condition of the hybrid action rule hold, calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the hybrid action rule, multiply the obtained membership degrees to get the activation intensity, and multiply the obtained activation intensity by the preset linear velocity and the preset angular velocity in the hybrid action rule respectively to obtain the target linear velocity and the target angular velocity.

[0134] In this embodiment, for a hybrid action rule, the prerequisite for the activation of the hybrid action rule is that the activation intensity of each exact condition therein is 1 (i.e., each exact condition holds), then multiply the membership degrees of each fuzzy condition to obtain the activation intensity of the hybrid action rule, and finally multiply the activation intensity by the preset linear velocity and the preset angular velocity in the hybrid action rule respectively to obtain a set of target linear velocity and target angular velocity.

[0135] Among them, the calculation method of the membership degree of each fuzzy condition in the hybrid action rule is the same as that of each fuzzy condition in the above-mentioned fuzzy action rule, which will not be elaborated here.

[0136] For example, "if there is an obstacle on the left and close to the obstacle, then turn right" is a hybrid rule in the human cognitive system, where "there is an obstacle on the left" is an exact condition and "close to the obstacle" is a fuzzy condition. Then, in combination with the above example, a hybrid action rule determined based on this can be expressed as: "if there is an obstacle on the left and , then and ", this example is only for illustration and is not limited here.

[0137] S10314. Add all the target linear velocities to obtain the guiding linear velocity, and intersect all the target angular velocities to obtain the guiding angular velocity.

[0138] In this embodiment, for the th interaction data, input the state vector corresponding to the th interaction data into the knowledge system. Through the above steps S10311 - S10314, multiple target linear velocities and multiple target angular velocities can be determined. Add all the target linear velocities to obtain the guiding linear velocity, and intersect all the target angular velocities to obtain the guiding angular velocity. In this way, the guiding action instruction corresponding to the th interaction data is obtained.

[0139] S1032. Input the state vector into the deep deterministic policy gradient algorithm model to obtain the policy action instruction and the state value.

[0140] In this embodiment, the deep deterministic policy gradient algorithm model includes an Actor network and a Critic network. The state vector is first input into the Actor network to obtain a policy action instruction, which includes a predicted linear velocity and a predicted angular velocity. The policy action instruction is input into the Critic network to obtain a state value. The detailed processing process in the state vector deep deterministic policy gradient algorithm model is prior art and will not be elaborated here.

[0141] S1033. Input the guidance action instruction and the policy action instruction into a generalizer for fusion to obtain a comprehensive action instruction.

[0142] In this embodiment, the comprehensive action instruction output by the generalizer includes a comprehensive linear velocity and a comprehensive angular velocity.

[0143] Optionally, the sub-steps of step S1033 may include:

[0144] (1) Input the guidance linear velocity, guidance angular velocity, predicted linear velocity, and predicted angular velocity into the generalizer;

[0145] (2) Use the current fusion parameter to perform weighted processing on the predicted linear velocity and the guidance linear velocity to obtain a pending linear velocity, and add Gaussian noise to the pending linear velocity to obtain a comprehensive linear velocity;

[0146] (3) Use the current fusion parameter to perform weighted processing on the predicted angular velocity and the guidance angular velocity to obtain a pending angular velocity, and add Gaussian noise to the pending angular velocity to obtain a comprehensive angular velocity.

[0147] Among them, the calculation formula for the comprehensive linear velocity is as follows:

[0148]

[0149] Among them, is the comprehensive linear velocity, is the fusion parameter of the generalizer, is the guidance linear velocity, is the predicted linear velocity; represents Gaussian noise, is the mean of the Gaussian noise, is the variance of the Gaussian noise, is the input of the Gaussian noise function.

[0150] Among them, the calculation formula for the comprehensive angular velocity is as follows:

[0151]

[0152] Among them, is the comprehensive angular velocity, is the guidance angular velocity, To predict the angular velocity.

[0153] S1034. Input the current fusion parameters, state value, policy action instruction, and comprehensive action instruction of the generalizer into a preset loss function to obtain the comprehensive loss corresponding to the th interaction data.

[0154] In this embodiment, for each of the pieces of interaction data, execute steps S1031 to S1034 to obtain comprehensive losses. Among them, the calculation formula of the preset loss function is:

[0155]

[0156]

[0157]

[0158] Where represents the comprehensive loss, represents the linear velocity loss, represents the angular velocity loss; is a fixed parameter indicating the credibility of the knowledge system; represents the state value.

[0159] S1035. Calculate the expected loss based on the comprehensive losses corresponding to the pieces of interaction data, and update the model parameters of the deep deterministic policy gradient algorithm model and the fusion parameters of the generalizer based on the expected loss.

[0160] In this embodiment, based on the comprehensive losses corresponding to the pieces of interaction data, the expected loss (i.e., the average loss) can be calculated, and the model parameters of the deep deterministic policy gradient algorithm model and the fusion parameters of the generalizer can be updated using this expected loss. The specific update process is prior art and will not be elaborated here.

[0161] It can be seen from the above calculation formula of the preset loss function that if the fixed parameter is not 0, it means that even when the fusion parameters of the generalizer , there is always a certain proportion of supervision component in the update of the model parameters of the deep deterministic policy gradient algorithm model. Therefore, for a knowledge system with extremely high reliability, the fixed parameter can be set so that there is a certain proportion of supervision in the update of the model parameters of the deep deterministic policy gradient algorithm model.

[0162] It should be noted that the execution order of each step in the above method embodiments is not limited by the figures shown, and the execution order of each step is subject to the actual application situation.

[0163] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0164] The mapless navigation model of the present invention is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system, rather than simply using the DDPG algorithm. This enables the mapless navigation model to not only adapt to navigation scenarios similar to the training environment but also complex dynamic navigation scenarios, with stronger navigation generalization.

[0165] During the training process of the mapless navigation model of the present invention, while inputting the state vector into the deep deterministic policy gradient algorithm model to obtain the policy action instruction, the state vector is also input into the knowledge system for rule matching to obtain the guiding action instruction. Finally, the guiding action instruction and the policy action instruction are fused to obtain the comprehensive action instruction, reducing the randomness of action selection by the deep deterministic policy gradient algorithm model, thereby enabling valuable data to be quickly obtained and improving the learning efficiency.

[0166] During the training process of the mapless navigation model of the present invention, both the policy action instruction and the fused comprehensive action instruction are input into the preset loss function, reducing the uncertainty of the corresponding loss function when simply using the DDPG algorithm, thus making the learning process more stable.

[0167] Since the present invention fuses the guiding action instruction and the policy action instruction to obtain the comprehensive action instruction, in an environment with sparse rewards, the mobile robot can interact with the environment under the guidance of knowledge rather than randomly, avoiding falling into local optima and ensuring easy convergence in an environment with sparse rewards.

[0168] To execute the corresponding steps in the above method embodiments and each possible implementation manner, the following provides an implementation manner of a navigation device.

[0169] Please refer to Figure 4 , Figure 4 FIG. shows a schematic structural diagram of a navigation device provided by an embodiment of the present invention. The navigation device 200 includes: a state acquisition module 210, a prediction module 220, and a control module 230.

[0170] The state acquisition module 210 is configured to acquire the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment.

[0171] The status acquisition module 210 is further configured to convert the environmental status data and relative position data at the current moment, as well as the action instruction at the previous moment, into a current state vector;

[0172] The prediction module 220 is configured to input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; the mapless navigation model is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system;

[0173] The control module 230 is configured to use the linear velocity and angular velocity in the action instruction at the current moment to control the mobile robot to move forward towards the target point.

[0174] Optionally, the mapless navigation model is trained in the following manner: Based on the knowledge system, the pre-constructed deep deterministic policy gradient algorithm model, and the generalizer, controlling the intelligent agent to interact with a preset virtual environment to construct an experience database; the experience database is initially empty; when the number of interaction data in the experience database is greater than a certain value, randomly sample pieces of interaction data from the experience database; Based on pieces of interaction data and multiple action rules in the knowledge system, update the parameters of the deep deterministic policy gradient algorithm model and the generalizer; Based on the knowledge system, the updated deep deterministic policy gradient algorithm model, and the generalizer, control the intelligent agent to continue interacting with the virtual environment to expand the experience database; determine whether the number of interaction rounds between the intelligent agent and the virtual environment reaches a preset value; if not, return to the step of sampling pieces of interaction data from the experience database until the number of interaction rounds between the intelligent agent and the virtual environment reaches the preset value; if so, control the intelligent agent to stop interacting with the virtual environment and use the current deep deterministic policy gradient algorithm model as the mapless navigation model.

[0175] Optionally, the step of updating the parameters of the deep deterministic policy gradient algorithm model and the generalizer based on pieces of interaction data and multiple action rules in the knowledge system includes: For the th piece of interaction data, obtain the state vector from the th piece of interaction data, and match the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction; where ; input the state vector into the deep deterministic policy gradient algorithm model to obtain a policy action instruction and a state value; input the guiding action instruction and the policy action instruction into the generalizer for fusion to obtain a comprehensive action instruction; input the current fusion parameters of the generalizer, the state value, the policy action instruction, and the comprehensive action instruction into a preset loss function to obtain the The comprehensive loss corresponding to the interaction data; based on the comprehensive loss corresponding to the interaction data, calculate the expected loss, and update the model parameters of the deep deterministic policy gradient algorithm model and the fusion parameters of the generalizer based on the expected loss.

[0176] Optionally, the guiding action instruction includes a guiding linear velocity and a guiding angular velocity. The multiple action rules of the knowledge system include at least one exact action rule, at least one fuzzy action rule, and at least one hybrid action rule, and each action rule has associated position data; the exact action rule includes at least one exact condition and conclusion data, the fuzzy action rule includes at least one fuzzy condition and conclusion data, and the hybrid action rule includes at least one exact condition, at least one hybrid condition, and conclusion data; the conclusion data includes a preset linear velocity and a preset angular velocity. The steps of matching the state vector with the multiple action rules in the knowledge system to obtain the guiding action instruction include: for each exact action rule, obtain at least one vector value from the state vector based on the position data associated with the exact action rule; if at least one vector value makes each exact condition of the exact action rule hold, then use the preset linear velocity and preset angular velocity in the exact action rule as the target linear velocity and target angular velocity respectively; for each fuzzy action rule, obtain at least one vector value from the state vector based on the position data associated with the fuzzy action rule; calculate the membership degrees between at least one vector value and the fuzzy sets associated with each fuzzy condition of the fuzzy action rule, and multiply the obtained membership degrees to get the activation intensity, multiply the obtained activation intensity by the preset linear velocity and preset angular velocity in the fuzzy action rule respectively to get the target linear velocity and target angular velocity; for each hybrid action rule, obtain at least one vector value from the state vector based on the position data associated with the hybrid action rule; if at least one vector value makes each exact condition of the hybrid action rule hold, then calculate the membership degrees between at least one vector value and the fuzzy sets associated with each fuzzy condition of the hybrid action rule, and multiply the obtained membership degrees to get the activation intensity, multiply the obtained activation intensity by the preset linear velocity and preset angular velocity in the hybrid action rule respectively to get the target linear velocity and target angular velocity; add all the target linear velocities to get the guiding linear velocity, and intersect all the target angular velocities to get the guiding angular velocity.

[0177] Optionally, the guiding action instruction includes a guiding linear velocity and a guiding angular velocity; the strategic action instruction includes a predicted linear velocity and a predicted angular velocity; the comprehensive action instruction includes a comprehensive linear velocity and a comprehensive angular velocity. The step of inputting the guiding action instruction and the strategic action instruction into a generalization device for fusion to obtain the comprehensive action instruction includes: inputting the guiding linear velocity, the guiding angular velocity, the predicted linear velocity, and the predicted angular velocity into the generalization device; using the current fusion parameters to perform weighted processing on the predicted linear velocity and the guiding linear velocity to obtain a to-be-determined linear velocity, adding Gaussian noise to the to-be-determined linear velocity to obtain the comprehensive linear velocity; using the current fusion parameters to perform weighted processing on the predicted angular velocity and the guiding angular velocity to obtain a to-be-determined angular velocity, adding Gaussian noise to the to-be-determined angular velocity to obtain the comprehensive angular velocity.

[0178] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the navigation device 200 described above can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0179] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a mobile robot provided by an embodiment of the present invention. The mobile robot 300 includes a processor 310, a memory 320, and a bus 330. The processor 310 is connected to the memory 320 through the bus 330.

[0180] The memory 320 can be used to store software programs. For example, the software program corresponding to the navigation device 200 provided by the embodiment of the present invention. The processor 310 executes various functional applications and data processing by running the software program stored in the memory 320 to implement the navigation method provided by the embodiment of the present invention.

[0181] Among them, the memory 320 can be, but is not limited to: RAM (Random Access Memory), ROM (Read Only Memory), FLASH (Flash Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electric Erasable Programmable Read-Only Memory), etc.

[0182] The processor 310 may be an integrated circuit chip with signal processing capabilities. The processor 310 may be a general-purpose processor, including: CPU (Central Processing Unit), NP (Network Processor), SoC (System on Chip), etc.; it may also be: DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0183] It can be understood that Figure 5 The structure shown is only for illustration, and the mobile robot 300 may also include more or fewer components than Figure 5 shown therein, or have a different configuration from Figure 5 shown therein. Figure 5 Each component shown therein may be implemented by hardware, software, or a combination thereof.

[0184] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the navigation method disclosed in the above embodiment is implemented. The computer-readable storage medium may be, but is not limited to: various media such as USB flash drives, mobile hard disks, ROM, RAM, PROM, EPROM, EEPROM, FLASH magnetic disks, or optical discs that can store program codes.

[0185] In summary, the embodiment of the present invention provides a navigation method, device, robot, and storage medium implemented based on a knowledge-guided mapless navigation model. The navigation method is as follows: First, obtain the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment; then convert the environmental state data and relative position data at the current moment and the action instruction at the previous moment into a current state vector; then input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; finally, use the action instruction at the current moment to control the mobile robot to move forward towards the target point. The mapless navigation model of the present invention is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system, rather than simply using the DDPG algorithm. This enables the mapless navigation model to not only adapt to navigation scenarios similar to the training environment but also adapt to complex dynamic navigation scenarios, ensuring the accuracy of navigation.

[0186] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A navigation method implemented based on a knowledge-guided mapless navigation model, characterized in that, Including: Obtain the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment; Convert the environmental state data and relative position data at the current moment and the action instruction at the previous moment into a current state vector; Input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment; The mapless navigation model is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system; Use the action instruction at the current moment to control the mobile robot to move forward towards the target point; Among them, the mapless navigation model is obtained by training in the following manner: Based on the knowledge system, the deep deterministic policy gradient algorithm model, and the generalizer, control the intelligent agent to continuously interact with a preset virtual environment, and add each piece of interaction data generated by the interaction to an initially empty experience database; When the number of interaction data in the experience database reaches randomly sample pieces of interaction data from the experience database; is a positive integer; Based on the interactive data and multiple action rules in the knowledge system, update the parameters of the deep deterministic policy gradient algorithm model and the generalizer; Based on the knowledge system, the updated deep deterministic policy gradient algorithm model, and the updated generalizer, control the intelligent agent to continuously interact with the virtual environment, and add each piece of interaction data generated by the interaction to the experience database; Judge whether the number of interaction rounds between the intelligent agent and the virtual environment reaches a preset value; Otherwise, return the step of sampling the interaction data from the experience database until the number of interaction rounds between the agent and the virtual environment reaches the preset value; If so, control the intelligent agent to stop interacting with the virtual environment, and use the current deep deterministic policy gradient algorithm model as the mapless navigation model; Among them, the step of updating the parameters of the deep deterministic policy gradient algorithm model and the generalizer based on the multiple interaction data and multiple action rules in the knowledge system includes: For the th interaction data, obtain a state vector from the th interaction data, and match the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction; wherein, ; Input the state vector into the deep deterministic policy gradient algorithm model to obtain a policy action instruction and a state value; Input the guiding action instruction and the policy action instruction into the generalizer for fusion to obtain a comprehensive action instruction; Input the current fusion parameter of the generalizer, the state value, the policy action instruction, and the comprehensive action instruction into a preset loss function to obtain the comprehensive loss corresponding to the th interaction data; Based on Calculate the expected loss based on the comprehensive loss corresponding to the interactive data, and update the model parameters of the deep deterministic policy gradient algorithm model and the fusion parameters of the generalizer based on the expected loss; Among them, the guiding action instruction includes a guiding linear velocity and a guiding angular velocity; the multiple action rules of the knowledge system include at least one precise action rule, at least one fuzzy action rule, and at least one mixed action rule, and each action rule has associated position data; the precise action rule includes at least one precise condition and conclusion data, the fuzzy action rule includes at least one fuzzy condition and conclusion data, the mixed action rule includes at least one precise condition, at least one mixed condition, and conclusion data; the conclusion data includes a preset linear velocity and a preset angular velocity; The step of matching the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction includes: For each precise action rule, obtain at least one vector value from the state vector based on the position data associated with the precise action rule; if the at least one vector value makes each precise condition of the precise action rule hold, then use the preset linear velocity and preset angular velocity in the precise action rule as the target linear velocity and target angular velocity respectively; For each of the fuzzy action rules, obtain at least one vector value from the state vector based on the position data associated with the fuzzy action rule; calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the fuzzy action rule, and multiply the obtained membership degrees to get the activation intensity. Multiply the obtained activation intensity by the preset linear velocity and the preset angular velocity in the fuzzy action rule respectively to obtain the target linear velocity and the target angular velocity. For each of the hybrid action rules, obtain at least one vector value from the state vector based on the position data associated with the hybrid action rule; if the at least one vector value makes each exact condition of the hybrid action rule hold, calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the hybrid action rule, and multiply the obtained membership degrees to get the activation intensity. Multiply the obtained activation intensity by the preset linear velocity and the preset angular velocity in the hybrid action rule respectively to obtain the target linear velocity and the target angular velocity. Add all the target linear velocities to obtain the guiding linear velocity, and intersect all the target angular velocities to obtain the guiding angular velocity.

2. The method according to claim 1, wherein The policy action instruction includes a predicted linear velocity and a predicted angular velocity; the comprehensive action instruction includes a comprehensive linear velocity and a comprehensive angular velocity. The step of inputting the guiding action instruction and the policy action instruction into the generalizer for fusion to obtain a comprehensive action instruction includes: Input the guiding linear velocity, the guiding angular velocity, the predicted linear velocity, and the predicted angular velocity into the generalizer. Use the current fusion parameter to perform weighted processing on the predicted linear velocity and the guiding linear velocity to obtain a pending linear velocity, and add Gaussian noise to the pending linear velocity to obtain the comprehensive linear velocity. Use the current fusion parameter to perform weighted processing on the predicted angular velocity and the guiding angular velocity to obtain a pending angular velocity, and add the Gaussian noise to the pending angular velocity to obtain the comprehensive angular velocity.

3. The method according to claim 1, characterized in that The policy action instruction includes a predicted linear velocity and a predicted angular velocity. The calculation formula of the preset loss function is: Among them, represents the comprehensive loss, represents the linear velocity loss, represents the angular velocity loss; is a fixed parameter, indicating the credibility of the knowledge system; represents the state value, is the current fusion parameter of the generalizer, represents the predicted linear velocity, represents the comprehensive linear velocity, represents the predicted angular velocity, represents the comprehensive angular velocity.

4. A navigation device implemented based on a knowledge-guided mapless navigation model, characterized in that, including: A state acquisition module, configured to acquire the environmental state data of the mobile robot at the current moment, the relative position data between the current position and the target point, and the action instruction at the previous moment. The state acquisition module is further configured to convert the environmental state data and relative position data at the current moment and the action instruction at the previous moment into a current state vector. A prediction module, configured to input the current state vector into the trained mapless navigation model to obtain the action instruction at the current moment. The mapless navigation model is obtained by training a pre-constructed deep deterministic policy gradient algorithm model based on multiple action rules in a preset knowledge system. A control module, configured to use the linear velocity and angular velocity in the action instruction at the current moment to control the mobile robot to move forward towards the target point. Wherein, the mapless navigation model is trained in the following manner: Based on the knowledge system, the pre-constructed deep deterministic policy gradient algorithm model, and the generalizer, control the agent to interact with a preset virtual environment to construct an experience database; the experience database is initially empty; When the number of interaction data in the experience database is greater than , randomly sample pieces of interaction data from the experience database; Based on the interactive data items and multiple action rules in the knowledge system, update the parameters of the deep deterministic policy gradient algorithm model and the generalizer; Based on the knowledge system, the updated deep deterministic policy gradient algorithm model, and the generalizer, control the agent to continue to interact with the virtual environment to expand the experience database; Determine whether the number of interaction rounds between the agent and the virtual environment reaches a preset value; Otherwise, if returning the step of sampling pieces of interaction data from the experience database until the number of interaction rounds between the agent and the virtual environment reaches the preset value; If so, control the agent to stop interacting with the virtual environment and use the current deep deterministic policy gradient algorithm model as the mapless navigation model; Among them, the step of updating the parameters of the deep deterministic policy gradient algorithm model and the generalizer based on the multiple interaction data and multiple action rules in the knowledge system includes: For the th interaction data, obtain a state vector from the th interaction data, and match the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction; wherein, ; Input the state vector into the deep deterministic policy gradient algorithm model to obtain a policy action instruction and a state value; Input the guiding action instruction and the policy action instruction into the generalizer for fusion to obtain a comprehensive action instruction; Input the current fusion parameter of the generalizer, the state value, the policy action instruction, and the comprehensive action instruction into a preset loss function to obtain the comprehensive loss corresponding to the th interaction data; Based on calculate the expected loss based on the comprehensive loss corresponding to the interactive data, and update the model parameters of the deep deterministic policy gradient algorithm model and the fusion parameters of the generalizer based on the expected loss; Wherein, the guiding action instruction includes a guiding linear velocity and a guiding angular velocity; the multiple action rules of the knowledge system include at least one precise action rule, at least one fuzzy action rule, and at least one hybrid action rule, and each action rule has associated position data; the precise action rule includes at least one precise condition and conclusion data, the fuzzy action rule includes at least one fuzzy condition and conclusion data, and the hybrid action rule includes at least one precise condition, at least one hybrid condition, and conclusion data; the conclusion data includes a preset linear velocity and a preset angular velocity; The step of matching the state vector with multiple action rules in the knowledge system to obtain a guiding action instruction includes: For each precise action rule, obtain at least one vector value from the state vector based on the position data associated with the precise action rule; if the at least one vector value makes each precise condition of the precise action rule hold, then use the preset linear velocity and preset angular velocity in the precise action rule as the target linear velocity and target angular velocity respectively; For each fuzzy action rule, obtain at least one vector value from the state vector based on the position data associated with the fuzzy action rule; calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the fuzzy action rule, and multiply the obtained membership degrees to get the activation intensity, and multiply the obtained activation intensity by the preset linear velocity and preset angular velocity in the fuzzy action rule respectively to get the target linear velocity and target angular velocity; For each of the mixed action rules, obtain at least one vector value from the state vector based on the position data associated with the mixed action rule; if the at least one vector value makes each exact condition of the mixed action rule hold, calculate the membership degrees between the at least one vector value and the fuzzy sets associated with each fuzzy condition of the mixed action rule, and multiply the obtained membership degrees to get the activation intensity. Multiply the obtained activation intensity by the preset linear velocity and the preset angular velocity in the mixed action rule respectively to obtain the target linear velocity and the target angular velocity. Add all the target linear velocities to obtain the guiding linear velocity, and intersect all the target angular velocities to obtain the guiding angular velocity.

5. A mobile robot, characterized in that, Comprising: A memory and a processor, wherein the memory stores a software program, and when the mobile robot runs, the processor executes the software program to implement the navigation method according to any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the navigation method according to any one of claims 1-3 is implemented.

Citation Information

Patent Citations

  • Intelligent mobile platform map-free autonomous navigation method based on deep reinforcement learning

    CN111141300A

  • Method for obstacle avoidance of robot in the complex indoor scene based on monocular camera

    WO2022160430A1