SLAM implementation device and SLAM implementation method for determining compensation based on exploration compensation and development compensation
Patent Information
- Application Number
- KR1020240067626
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-24
Smart Images

Figure R1020240067626_ABST
Abstract
Description
Technology Field
[0001] The embodiment relates to a SLAM implementation device and a SLAM implementation method that determine a reward based on exploration rewards and development rewards. Background Technology
[0003] Reinforcement Learning (RL) is a field of machine learning (ML) focused on learning sequential decision-making processes. Agents, which also act as decision-makers, interact with the environment to learn the best course of action. This interaction can be represented as a Markov Decision Process (MDP). Combining Markov Decision Theory with Dynamic Programming (DP), MDPs represent a simple and mathematically ideal form of RL problems. Using this framework allows for the concise representation of essential functions.
[0004] In general, an MDP is defined by four tuples M = {S, A, P, R}. Here, S is a finite set of states in which an agent can be. A is the set of agent actions that the agent can perform while in state s ∈ S. Since an agent's actions or actions are a set of possible outcomes for performing a specific action, they can be understood as probability distributions; therefore, P represents the probability of transition. R is the reward function associated with performing a specific action.
[0005] Currently, most reinforcement learning-based active SLAM algorithms use reinforcement learning algorithms such as DDQN and DQN. These algorithms can only process individual robot movements and space. However, since the robot's actions in the active-SLAM process are continuous actions and a movement space, using individual movements results in discontinuous and unsmooth robot movement.
[0006] Furthermore, most methods utilizing reinforcement learning focus only on the robot's exploration or search process, while the development process is overlooked. If the focus is placed solely on the exploration process, the SLAM map may be incomplete and the accuracy may be low. Prior art literature
[0008] Korean Patent Publication No. 10-2010-0031992 (Published March 25, 2010) “Method for Creating an Environment Map of a Robot Using Neural Networks and Evolutionary Computation” The problem to be solved
[0009] The purpose of the embodiment is to design a reward function based on exploration and exploitation of the Deep Deterministic Policy Gradient (DDPG) algorithm and to propose a solution that applies DDPG to Active-SLAM tasks. means of solving the problem
[0011] A SLAM implementation device according to one embodiment is a device for implementing SLAM (Simultaneous Localization and Mapping) using reinforcement learning, comprising at least one sensor mounted on a robot, a memory, and a processor implemented to execute computer-readable commands, wherein the processor calculates a first reward based on map completeness created according to the movement of the robot, calculates an angle range where no obstacle detected by the sensor exists centered on the robot, calculates a second reward determined according to whether the robot moves to the angle range, calculates a third reward determined according to whether the robot moves to a path different from the initial path, calculates a reward function calculated based on the first reward, the second reward, and the third reward, learns an artificial neural network mounted on the robot using the reward function, and determines the movement path of the robot using the learned artificial neural network.
[0012] In addition, the map completeness mentioned above can be calculated using the following mathematical formula.
[0013] [Mathematical Formula]
[0014]
[0015] (Here, M represents map completeness, O is the number of cells determined to be occupied in the map, uo is the number of cells determined to be unoccupied, and represents the scaling factor, and X represents the total number of cells.)
[0016] Additionally, the second compensation can be calculated as a positive compensation if the robot moves within the angle range.
[0017] Additionally, the second compensation may be calculated as a negative compensation if the robot moves outside the angle range.
[0018] In addition, the third reward can be calculated as a positive reward if the robot moves to the other path.
[0019] In addition, the third compensation may be calculated as a negative compensation if the robot moves along the initial path.
[0020] In addition, the above compensation function can be calculated by linearly combining the first compensation, the second compensation, and the third compensation.
[0021] In addition, the movement path of the robot can be determined based on DDPG (Deep Deterministic Policy Gradient).
[0022] A SLAM implementation method according to another embodiment may include at least one processor and implement SLAM (Simultaneous Localization and Mapping) using reinforcement learning, the method comprising: a step of calculating a first reward based on map completeness created according to the movement of a robot; a step of calculating an angle range where no obstacles detected by a sensor mounted on the robot exist centered on the robot, and a step of calculating a second reward determined according to whether the robot moves to the angle range; a step of calculating a third reward determined according to whether the robot moves to a path different from an initial path; a step of calculating a reward function calculated based on the first reward, the second reward, and the third reward, and a step of learning an artificial neural network mounted on the robot using the reward function; and a step of determining the movement path of the robot using the learned artificial neural network.
[0023] In addition, the map completeness mentioned above can be calculated using the following mathematical formula.
[0024] [Mathematical Formula]
[0025]
[0026] (Here, M represents map completeness, O is the number of cells determined to be occupied in the map, uo is the number of cells determined to be unoccupied, and represents the scaling factor, and X represents the total number of cells.)
[0027] Additionally, the second compensation can be calculated as a positive compensation if the robot moves within the angle range.
[0028] Additionally, the second compensation may be calculated as a negative compensation if the robot moves outside the angle range.
[0029] In addition, the third reward can be calculated as a positive reward if the robot moves to the other path.
[0030] In addition, the third compensation may be calculated as a negative compensation if the robot moves along the initial path.
[0031] In addition, the above compensation function can be calculated by linearly combining the first compensation, the second compensation, and the third compensation.
[0032] In addition, the movement path of the robot can be determined based on DDPG (Deep Deterministic Policy Gradient).
[0033] A recording medium according to another embodiment tangibly embodies a program of instructions that can be executed by a digital processing device to implement Simultaneous Localization and Mapping (SLAM) using Reinforcement Learning, and as a recording medium readable by a digital processing device, a program for executing the above-described method on a computer can be recorded thereon. Effects of the invention
[0035] According to the embodiment, since a map can be created based on the robot's continuous actions and continuous spatial information, the accuracy of the SLAM map can be improved. Brief explanation of the drawing
[0037] Figure 1 shows a flowchart of the DDPG algorithm based on the Active SLAM framework. Figure 2 shows an example diagram of the configuration and occupancy of a grid map. FIG. 3a shows an example of setting an angle range in the exploration reward of an embodiment, and FIG. 3b shows an example of movement of a robot utilizing the exploration reward of an embodiment. Figure 4 shows the code used to implement the exploration reward of the embodiment. FIG. 5a shows the movement path of a robot utilizing exploration rewards in an embodiment, and FIG. 5b shows the movement path of a robot utilizing development rewards in an embodiment. Figure 6 shows the code used to implement the development reward of the embodiment. Figure 7 is an example diagram illustrating the difference in accuracy between maps created according to a comparative example and an embodiment. Specific details for implementing the invention
[0038] In describing the embodiments of this specification, detailed descriptions of known technologies related to this specification are omitted if it is determined that such descriptions would unnecessarily obscure the essence of this specification. Furthermore, the terms described below are defined in consideration of their functions within this specification, and these definitions may vary depending on the intentions or practices of the user or operator. Therefore, such definitions should be based on the content throughout this specification. Terms used in the detailed description are intended merely to describe the embodiments of this specification and should never be interpreted restrictively. Unless explicitly stated otherwise, expressions in the singular form include the meaning of the plural form. In this description, expressions such as "include" or "comprise" are intended to refer to certain characteristics, numbers, steps, actions, elements, parts thereof, or combinations thereof, and should not be interpreted as excluding the existence or possibility of one or more other characteristics, numbers, steps, actions, elements, parts thereof, or combinations thereof other than those described.
[0039] Terms containing ordinal numbers, such as “first,” “second,” etc., may be used to describe various components, but said components are not limited by said terms. These terms may be used solely in a nominal sense to distinguish one component from another, and their sequential meaning is determined not by such nomenclature but by the context of the description.
[0040] The term “and / or” is used to include all cases of any combination of the multiple items in question. For example, “and / or B” means including all three cases, such as “and B.”
[0041] When it is stated that one component is "connected" or "joined" to another component, it should be understood that while it may be directly connected or joined to that other component, there may also be other components in between.
[0042] Hereinafter, specific embodiments of the examples will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, devices, and / or objects described herein. However, this is merely illustrative and the examples are not limited thereto.
[0043] Figure 1 shows a flowchart of the DDPG algorithm based on the Active SLAM framework.
[0044] Referring to Fig. 1, the local SLAM module receives input from LiDAR data, IMU data, and robot motion control commands. These inputs can construct a map of the robot's local environment and estimate the robot's trajectory. The overall SLAM algorithm can merge maps from different local SLAM sessions to create an overall consistent map. The DPPG algorithm considers the correlation between multiple local maps and the transformation of the robot's attitude across different local maps. The loop closure module performs loop closure operations, such as adjusting the robot's attitude or optimizing the overall map, to minimize errors caused by loops and ensure map consistency.
[0045] First, the parameters of the artificial neural network are initialized. The agent selects an action according to the action policy. Random noise can be added to the selected action policy.
[0046] The compensation algorithm of the embodiment utilizes sensor data and map information from the SLAM stage. These data accurately describe the current state of the agent (robot) and can support a compensation function.
[0047] Figure 2 shows an example diagram of the configuration and occupancy of a grid map.
[0048] Referring to FIG. 2, the black areas represent non-empty, i.e., occupied grids or cells, and the white areas represent empty (unoccupied) grids or cells. Gray represents unknown areas.
[0049] The first reward of the embodiment refers to a reward based on map completeness.
[0050] Map completeness can be expressed by the following mathematical formula 1.
[0051]
[0052] Here, M represents the map completeness, O is the number of cells determined to be occupied in the map, and uo is the number of cells determined to be unoccupied. represents the scaling factor, and X represents the total number of cells. Map completeness M can be referenced as Mc. Scaling factor is a number between 0 and 1.
[0053] The first reward (r1) based on map completeness can be expressed by the following mathematical formula 2.
[0054]
[0055] Referring to mathematical formula 2, the first reward can be calculated in three ways.
[0056] Here, rmapdone is the scaling factor Rcrach is a positive scalar when the map is completed at a specific threshold indicated by 95%. Rcrach is a scalar when a collision occurs. When the distance between the robot and the obstacle is less than the minimum detection distance Lmin of the LiDAR, Rcrach can be a negative scalar. Otherwise, the first reward can be calculated as the map completion Mt at a specific time t and the map completion Mt-1 at the time prior to t, t-1.
[0057] Since there are no cases where the map completeness is 100% due to minor mapping errors, the threshold for map completeness in the example was set to 95%.
[0058] FIG. 3a shows an example of setting an angle range in the exploration reward of an embodiment, and FIG. 3b shows an example of robot movement utilizing the exploration reward of an embodiment. FIG. 4 shows the code used to implement the exploration reward of an embodiment. The exploration reward of an embodiment refers to a second reward (r2).
[0059] The motivation for introducing exploration rewards stems from the desire to induce robots to actively explore unknown environments, particularly areas where the LiDAR failed to detect obstacles in the previous moment. If we denote the current time as t and the previous time as t-1, the LiDAR data from the previous time t-1 is stored in the Robot Operating System (ROS).
[0060] Obstacles are distributed at angles ranging from 0 to 359 degrees. From the previous position, Pt-1, the robot identifies the obstacles surrounding it and can mark them on the map as a set of red dots. When an obstacle is out of the LiDAR's sensing range, the corresponding data can be represented as an Inf value. The exploration reward considers this angular range of Inf values.
[0061] In exploration rewards, an angle range N where the angle value is Inf is considered. That is, in the embodiment, N can be referenced as an angle range where no obstacles exist.
[0062] If the discovered angle belongs to N, a positive reward value is provided, and if the discovered angle does not belong to N, that is, if it is outside the angle range, a negative reward value is provided.
[0063] In other words, in the exploration reward (r2) of the example, the reward can be determined as shown in the following mathematical formula 3.
[0064]
[0065] Here, represents the distance between the current position Pt and the previous position Pt-1. τ represents the discount factor that adjusts the degree of reward and punishment.
[0066] Referring to Fig. 4, the exploration reward of the embodiment can be performed when the map completion rate (Mc) is 96% or less.
[0067] FIG. 5a shows the movement path of a robot utilizing exploration rewards in an embodiment, and FIG. 5b shows the movement path of a robot utilizing development rewards in an embodiment. FIG. 6 shows the code used to implement development rewards in an embodiment. The development reward in an embodiment refers to the third reward (r3).
[0068] The motivation for introducing development rewards stemmed from the desire to induce robots to deviate from their initial paths and explore alternative routes during the exploration process.
[0069] As can be seen in Fig. 5a, the robot continuously explores unknown areas during the exploration process, but nevertheless, unknown areas such as those marked in gray may occur. These gray areas can degrade the accuracy of the map. Fig. 5b shows the robot's expected path during the movement process based on development rewards. By utilizing the development rewards of the embodiment, the robot can increase the accuracy of the map and move along more diverse paths for complete area detection.
[0070] In the development reward process, if the robot successfully deviates from the initial path, the reward value gradually increases, causing the robot to reinforce this behavior. Conversely, if the robot revisits the previous initial path, it may receive a small penalty value.
[0071] In the development compensation (r3) of the example of Balhae again, the compensation can be determined as shown in the following mathematical formula 4.
[0072]
[0073] Here, represents the distance between the current position Pt and the previous position Pt-1. Ρ represents a discount factor that adjusts the degree of reward and punishment. The distance threshold between two path points is denoted by dismin and can be adjusted according to actual environmental conditions.
[0074] Referring to FIG. 6, the development reward of the embodiment can be performed when the map completion rate (Mc) is 96% or higher.
[0075] In the embodiment, the compensation function may represent the total sum of compensations. That is, in the embodiment, the compensation function for the robot's state st at time t can be expressed by the following mathematical formula 5.
[0076]
[0077] Here, α, β, and γ represent scalar weight factors for the three rewards r1, r2, and r3.
[0078] In the embodiment, Deep Deterministic Policy Gradient (DDPG) was selected as the path planning algorithm to control the robot (e.g., linear velocity and angular velocity) to search for a target location without colliding with obstacles. Here, the linear velocity is a continuous function restricted to the range [0, 1] to allow only forward motion, and the angular velocity is a constant function specified in the range [-1, 1] to rotate left or right. The state vector st can be expressed as o, which is 180 LiDAR data points in a 360-degree range, the action at-1 at time t-1, and the currently occupied map m, as shown in Equation 6 below.
[0079]
[0080] Figure 7 is an example diagram illustrating the difference in accuracy between maps created according to a comparative example and an embodiment.
[0081] FIG. 7a shows a map created according to the Manual Control (MC) method, FIG. 7b shows a map created according to the Frontier Detection based exploration (FD) method, FIG. 7c shows a map created according to the D3Qn- and D-opt-based SLAM (DDS) method, and FIG. 7d shows a map created according to the SLAM implementation method of the embodiment.
[0082] The four methods described above were applied to the SLAM mapping task in three scenarios: Env-1, Env-2, and Env-3. Referring to Fig. 7a, it can be seen that the map obtained through manual control shows clear and accurate edges. Referring to Figs. 7b and 7c, it was found that the FD and DDS methods increase the computational load and decrease map accuracy. Referring to Fig. 7d, it can be confirmed that the map created according to the SLAM implementation method of the embodiment shows map accuracy similar to or better than the map generated through manual control.
[0083] The foregoing description of this specification is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical concept or essential features of the embodiments. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.
[0084] The scope of the embodiments is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and the concept of equivalents thereof should be interpreted as being included within the scope of the embodiments.
Claims
Claim 1 A SLAM implementation device for implementing SLAM (Simultaneous Localization and Mapping) using reinforcement learning, comprising: at least one sensor mounted on a robot; a memory; and a processor implemented to execute computer-readable commands, wherein the processor calculates a first reward based on map completeness generated according to the movement of the robot, calculates an angle range centered on the robot where no obstacle detected by the sensor exists, calculates a second reward determined according to whether the robot moves into the angle range, calculates a third reward determined according to whether the robot moves along a path different from an initial path, calculates a reward function calculated based on the first reward, the second reward, and the third reward, learns an artificial neural network mounted on the robot using the reward function, and determines the movement path of the robot using the learned artificial neural network. Claim 2 In claim 1, the SLAM implementation device wherein the map completeness is calculated using the following mathematical formula.[Mathematical Formula] (Here, M represents map completeness, O is the number of cells determined to be occupied in the map, uo is the number of cells determined to be unoccupied, and represents the scaling factor, and X represents the total number of cells.) Claim 3 A SLAM implementation device according to claim 1, wherein the second compensation is calculated as a positive compensation when the robot moves within the angle range. Claim 4 In paragraph 3, the SLAM implementation device wherein the second compensation is calculated as a negative compensation when the robot moves outside the angle range. Claim 5 A SLAM implementation device according to claim 1, wherein the third compensation is calculated as a positive compensation when the robot moves to the other path. Claim 6 In claim 5, the above third compensation is a SLAM implementation device that is calculated as a negative compensation when the robot moves along the above initial path. Claim 7 In claim 1, the compensation function is a SLAM implementation device calculated by linearly combining the first compensation, the second compensation, and the third compensation. Claim 8 A SLAM implementation device according to claim 1, wherein the determination of the robot's movement path is determined based on DDPG (Deep Deterministic Policy Gradient). Claim 9 A SLAM implementation method comprising at least one processor and implementing SLAM (Simultaneous Localization and Mapping) using reinforcement learning, the method comprising: a step of calculating a first reward based on map completeness generated according to the movement of a robot; a step of calculating an angle range centered on the robot where no obstacles detected by a sensor mounted on the robot exist, and a step of calculating a second reward determined according to whether the robot moves into the angle range; a step of calculating a third reward determined according to whether the robot moves along a path different from an initial path; a step of calculating a reward function based on the first reward, the second reward, and the third reward, and training an artificial neural network mounted on the robot using the reward function; and a step of determining the movement path of the robot using the trained artificial neural network. Claim 10 In claim 9, the above map completeness is a SLAM implementation method calculated by the following mathematical formula.[Mathematical Formula] (Here, M represents map completeness, O is the number of cells determined to be occupied in the map, uo is the number of cells determined to be unoccupied, and represents the scaling factor, and X represents the total number of cells.) Claim 11 In claim 9, the SLAM implementation method wherein the second compensation is calculated as a positive compensation when the robot moves within the angle range. Claim 12 A SLAM implementation method according to claim 11, wherein the second compensation is calculated as a negative compensation when the robot moves outside the angle range. Claim 13 In claim 9, the above third reward is a SLAM implementation method in which the robot moves to the other path and is calculated as a positive reward. Claim 14 In paragraph 13, the above third compensation is a SLAM implementation method in which the robot moves along the above initial path and is calculated as a negative compensation. Claim 15 In claim 9, the compensation function is a SLAM implementation method calculated by linearly combining the first compensation, the second compensation, and the third compensation. Claim 16 In claim 9, the SLAM implementation method in which the movement path of the robot is determined based on DDPG (Deep Deterministic Policy Gradient). Claim 17 A computer-readable recording medium characterized by having a program executable on a digital processing device for performing a method of implementing Simultaneous Localization and Mapping (SLAM) using reinforcement learning, wherein the program performs the step of claim 9.
Citation Information
Patent Citations
Position recognition method for mobile object using convergence of sensors and apparatus thereof
KR1020140030955A
Method and device for real-time mapping and localization
KR1020180004151A