Method and system for controlling robot on basis of trajectory optimization based on reward or cost optimization

The method optimizes robot trajectories using gradient-based algorithms and STE for non-differentiable sections, addressing inefficiencies and safety issues in existing control systems by minimizing costs and maximizing rewards under hardware and safety constraints.

WO2025187883A1PCT designated stage Publication Date: 2025-09-11NAVER CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/013765
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-04
Filing Date
2024-09-11
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing robot control methods struggle to optimize trajectories efficiently while adhering to hardware constraints and safety requirements, leading to suboptimal performance and potential safety hazards.

Method used

A method and system that utilize trajectory optimization by determining a cost function and updating actions using gradient descent or ascent algorithms, with a Straight Through Estimator (STE)-based approach for non-differentiable sections, to minimize or maximize cost/reward under constraints such as hardware limitations and safety considerations.

Benefits of technology

Enables accurate and efficient robot control that optimizes movement and service provision by minimizing costs and maximizing rewards while satisfying essential constraints, ensuring safe and precise operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024013765_12092025_PF_FP_ABST
    Figure KR2024013765_12092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a method in which a cost function for determining the cost incurred or reward obtained as a robot performs possible actions in a first state of the robot is determined, the actions are updated to minimize / maximize the cost / reward on the basis of a gradient value calculated for the cost function under a constraint related to control of the robot, and the robot is controlled on the basis of at least one of the updated actions.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for controlling a robot based on trajectory optimization according to optimization of reward or cost

[0001] The description below relates to a method and system for controlling a robot by optimizing the trajectory of the robot by optimizing the cost incurred or the reward obtained as the robot performs an action in a work state.

[0002] Robots used in industrial settings or providing services in indoor spaces must be controlled to operate accurately according to their intended purpose. For example, robots used in industrial settings require components such as manipulators and robotic arms to be controlled to the precise position and / or angle appropriate to their intended purpose. Robots operating within spaces such as buildings to provide services must be controlled to avoid obstacles and move at optimal speeds to their intended destination.

[0003] Control of these robot movements can be achieved through trajectory optimization, whereby parameters for controlling the robot (e.g., its speed, angular velocity, rotation angle, etc.) are determined. Meanwhile, robots must be controlled within hard constraints, such as those imposed by their hardware specifications or those required for safety within the operating space.

[0004] Therefore, in order to ensure the safe operation of the robot and the efficient provision of services by the robot, it is important to perform trajectory optimization for the robot to control the robot so that it operates correctly for its purpose while satisfying various constraints.

[0005] Korean Patent Publication No. 10-2005-0024840 is a technology regarding a path planning method for an autonomous mobile robot, and discloses a method for planning an optimal path for a mobile robot moving autonomously at home or in the office to safely and quickly reach a target point while avoiding obstacles.

[0006] The information described above is for the purpose of understanding only and may contain matters that do not form part of the prior art and may not contain what the prior art would suggest to a person skilled in the art.

[0007] In a first state of a robot, a cost function is determined to determine a cost incurred or a reward obtained as the robot performs possible actions, and under constraints related to the control of the robot, actions are updated to minimize / maximize the cost / reward based on a gradient value calculated for the cost function, and a method of controlling the robot based on at least one of the updated actions can be provided.

[0008] A method can be provided to perform optimization by calculating the gradient using a gradient descent algorithm or a gradient ascent algorithm for a cost function, and to determine the gradient as a predetermined value using an STE (Straight Through Estimator)-based algorithm for a section where differentiation is not possible, and to update actions to optimize reward or cost using the gradient.

[0009] In one aspect, a robot control method performed by a robot or a robot control system controlling a robot is provided, comprising: identifying a first state of the robot; acquiring actions of the robot for predicting states of the robot after the first state; and determining a cost function for determining a cost incurred or a reward obtained as the robot performs the actions in the first state based on the first state, the states, and the actions; and updating the actions so as to minimize or maximize the cost or the reward based on a gradient value calculated for the cost function under a constraint related to the control of the robot, wherein the cost function includes an interval in which the actions cannot be differentiated due to the constraint, and the gradient of the interval in which the actions cannot be differentiated is determined to be a predetermined value.

[0010] The above robot control method may further include a step of controlling the robot based on at least one of the updated behaviors.

[0011] The control of the robot is to control the movement of the robot from the current location of the robot to the destination, and the first state represents the current location or the control state of the robot at the current location, and the states represent the locations to which the robot will move from the current location or the control state of the robot at the locations to which the robot will move, and each of the actions may include at least one of the speed of the robot, the acceleration of the speed, the angular velocity of a driving unit for moving the robot, and the angular acceleration of the angular velocity.

[0012] The above constraints may include at least one of a maximum speed of the robot, a maximum acceleration of the robot, an angular velocity of the robot, a maximum angular acceleration of the robot, and a maximum movement angle of a manipulator or robot arm included in the robot.

[0013] The updating step includes a step of determining the gradient of the non-differentiable section as the predetermined value using a Straight Through Estimator (STE)-based algorithm, and the updating step may update the actions to minimize or maximize the cost or the reward based on the gradient of the cost function using a gradient descent algorithm or a gradient ascent algorithm.

[0014] The gradient of the above non-differentiable interval can be determined as the above-mentioned predetermined value of 1.

[0015] The step of acquiring the robot's behaviors includes acquiring a certain number of behaviors through random sampling, and the step of updating includes a step of updating the behaviors by changing the values ​​of the behaviors in the direction or the opposite direction of the calculated gradient; and a step of determining the finally updated behaviors by repeating the step of updating the behaviors by changing the values ​​of the behaviors until the cost or reward calculated through the cost function converges, and the robot can be controlled based on the finally updated behaviors.

[0016] The step of controlling the robot may include the step of determining a second state of the robot following the first state based on the updated actions; and the step of controlling the robot from the first state to the second state.

[0017] The above robot control method may further include a step of obtaining actions of the robot by setting the second state to the first state; a step of determining the cost function; and a step of repeating the step of updating the actions.

[0018] The above actions are a sequence A of the actions of the robot in mathematical expression 1, and the first state is s t The following states of the above robot are defined by the model of mathematical expression 2,

[0019] [Mathematical Formula 1]

[0020]

[0021] [Equation 2]

[0022] ,

[0023] a represents the behavior of the robot, t represents a specific point in time, and the cost function is defined by mathematical expression 3.

[0024] [Equation 3]

[0025] ,

[0026] c(s t , a t ) is the first state s t In the above robot acts a t A model for determining the cost incurred by performing the above actions, and an updated sequence A of the above actions to minimize the above cost. * can be defined by mathematical formula 4.

[0027] [Equation 4]

[0028]

[0029] The above updated sequence A * middle act as a t The updated action of the above cost function is action a t Gradient differentiated by is determined based on the gradient is calculated by mathematical formula 5,

[0030] [Equation 5]

[0031] ,

[0032] Mathematical formula 5 can be calculated by mathematical formula 6.

[0033] [Equation 6]

[0034]

[0035] The first state above is s t is defined by mathematical formula 7,

[0036] [Equation 7]

[0037] ,

[0038] x t and y t is the position coordinate of the robot at time t, is the position angle of the robot at time t, v t is the velocity of the robot at time t, w t represents the angular velocity of the robot at time t, and the first state s t The following state of the robot s t+1 is defined by mathematical formula 8,

[0039] [Equation 8]

[0040] ,

[0041] The above dt is the timescale between time points t and t+1, and the above v t target is the target velocity of the robot at time t, and w t target is the target angular velocity of the robot at time t, and a v and the above a wAs the above constraints, they may be an acceleration constraint of the speed of the robot and an angular acceleration constraint of the angular velocity of the robot, respectively.

[0042] The above acceleration constraint may indicate an acceleration range or maximum acceleration within which the robot can operate, and the above angular acceleration constraint may indicate an angular acceleration range or maximum angular acceleration within which the robot can operate.

[0043] In the above cost function, the target speed is v t +a v *The gradient of the section corresponding to the case where dt is greater than is determined by the above-mentioned value, and the target angular velocity is w t +a w *The gradient of the section corresponding to the case greater than dt can be determined by the above-mentioned value.

[0044] The above-mentioned non-differentiable interval of the above-mentioned cost function may be an interval in which the above-mentioned cost function is discontinuous.

[0045] In another aspect, a computer system for a robot moving in space is provided, comprising at least one processor configured to execute computer-readable instructions, wherein the at least one processor identifies a first state of the robot, obtains actions of the robot for predicting states of the robot after the first state, determines a cost function for determining a cost incurred or a reward obtained as the robot performs the actions in the first state based on the first state, the states, and the actions, and updates the actions to minimize or maximize the cost or the reward based on a gradient value calculated for the cost function under constraints related to the control of the robot, wherein the cost function includes an interval in which the actions are not differentiable due to the constraints, and the gradient of the interval in which the actions are not differentiable is determined to be a predetermined value.

[0046] In an embodiment, when a robot is controlled under hard constraints such as constraints based on the hardware specifications of the robot or constraints for safety in the space in which it operates, a dynamic model for predicting the state of the robot and a compensation model for determining cost or reward can be created by reflecting these hard constraints, and an updated action of the robot that can optimize (i.e., minimize / maximize) the cost / reward can be appropriately determined.

[0047] In an embodiment, in performing trajectory optimization for controlling a robot, a gradient is calculated and optimization is performed using a gradient descent algorithm or a gradient ascent algorithm for a cost function that determines a cost or reward obtained according to an action according to a state, but for a section where differentiation is impossible due to constraints, a STE-based algorithm can be used to determine the gradient to a predetermined value, and thus, trajectory optimization for the robot can be accurately performed while satisfying constraints.

[0048] FIG. 1 illustrates a method for controlling a robot based on trajectory optimization according to optimization of reward or cost, according to one embodiment.

[0049] FIG. 2 is a block diagram illustrating a computer system that performs a method for controlling a robot based on trajectory optimization, according to one embodiment.

[0050] FIG. 3 is a block diagram illustrating a robot controlled according to trajectory optimization, according to one embodiment.

[0051] FIGS. 4A and 4B are block diagrams illustrating a robot control system for controlling a robot according to one embodiment.

[0052] FIG. 5 is a flowchart illustrating a method for controlling a robot based on trajectory optimization according to optimization of compensation or cost, according to one embodiment.

[0053] Figure 6 is a flowchart illustrating a method for controlling a robot based on updated actions (sequence of actions), according to an example.

[0054] FIG. 7 illustrates a method for controlling a robot based on trajectory optimization by updating action parameters such as speed using gradient calculation and decision for a cost function, according to an example.

[0055] Figure 8 shows a model of the state of a robot that is the target of trajectory optimization according to an example.

[0056] Hereinafter, the detailed description will be given with reference to the attached drawings.

[0057]

[0058] FIG. 1 illustrates a method for controlling a robot based on trajectory optimization according to optimization of reward or cost, according to one embodiment.

[0059] In Fig. 1, as an example of control of a robot (100), a method is illustrated in which the robot (100) is controlled to move toward a destination (50) while avoiding obstacles (30) within a space (10). This robot (100) may be a mobile robot that autonomously moves within a space (10) to provide a service within the space (10).

[0060] The space (10) in which the robot (100) moves (or drives) is a place where the robot (100) provides a service, and may represent, for example, an indoor and / or outdoor space.

[0061] A robot (100) driving within a space (10) may be a service robot used to provide a service within the space (10). For example, the robot (100) may be configured to provide a service at a predetermined location within the space (10) or to a predetermined user through autonomous driving, and the (respective) movement of the robot (100) and provision of the service may be controlled by a robot control system (120). The movement of the robot (100) to a destination (50) controlled through the aforementioned driving algorithm may be movement of the robot (100) to a predetermined location for providing such a service.

[0062] The robot control system (120) (or, the robot (100)) can obtain actions (40) (i.e., a sequence of actions) to reach the destination (50) while avoiding obstacles (30) according to trajectory optimization for the robot (100), and accordingly, the robot (100) can be optimally controlled to move to the destination (50). These actions (40) can include control parameters of the robot (e.g., parameters associated with the displacement of the robot (100), such as the moving speed and angle of the robot (100)) in each state of the robot (100) (e.g., the displacement of the robot (100)). In addition, the actions (40) can represent an optimal path for the robot (100) to move to the destination (50).

[0063] In order to determine actions (40) for performing optimal control of the robot (100), in an embodiment, the robot control system (120) (or the robot (100)) can identify a first state corresponding to the current state of the robot (100), and can obtain a plurality of actions that the robot (100) can perform in order to predict states after the first state. The robot control system (120) can determine a cost function (or reward function) for determining a cost incurred or a reward obtained when the robot (100) performs an action in the first state, and can update the obtained actions to actions that minimize (or maximize) the cost (or reward). The robot (100) can be controlled by at least one of the updated actions.

[0064] In an embodiment, actions that minimize (or maximize) the cost (or reward) can be determined using the gradient calculated for the cost function. For example, the robot control system (120) can calculate the gradient by differentiating the cost function and determine actions that minimize (or maximize) the cost (or reward) using a gradient descent algorithm or a gradient ascent algorithm.

[0065] Meanwhile, the robot (100) may need to be controlled subject to essential constraints (hard constraints), such as constraints based on the hardware specifications of the robot (100) or constraints for safety in the space (10) in which it operates. The above-described cost function and the dynamic model for predicting the state according to the behavior of the robot (100) may be constructed to take into account such constraints in the control of the robot (100). At this time, the constructed cost function may include a section where differentiation (toward the behavior) is impossible due to the above-described constraints. In other words, the cost function may include a discontinuous section.

[0066] The robot control system (120) can determine the gradient of the corresponding section as a predetermined value using an STE (Straight Through Estimator)-based algorithm for the section where the cost function cannot be differentiated.

[0067] In this embodiment, by utilizing an STE-based algorithm in conjunction with a gradient descent / ascent algorithm, the costs and rewards associated with performing actions of the robot (100) can be optimized, thereby accurately performing trajectory optimization of the robot (100). Accordingly, the robot (100) can be controlled based on actions (action sequences) updated according to trajectory optimization under constraints.

[0068] A specific method for optimizing cost and reward by calculating the gradient of the cost function and a specific method for optimizing the trajectory of the robot (100) according to the method are described in more detail with reference to FIGS. 2 to 7, which will be described later.

[0069] Meanwhile, the robot (100) is not limited to the mobile robot described with reference to FIG. 1, and may also be a robot that is fixed in a location. For example, the robot (100) is used in industrial sites and may include at least one manipulator and / or robot arm. In this case, the trajectory optimization of the robot (100) may be for the control of the manipulator and / or robot arm. In other words, according to the trajectory optimization, the manipulator and / or robot arm can be controlled to an optimal position and angle suitable for its purpose.

[0070] Additionally, the robot (100) is not limited to a robot that actually operates in space (10), but may also be a robot agent for virtually simulating the movements of the robot (100).

[0071] In the detailed description to be described later, embodiments are described with a focus on navigation to a destination (50) or a predetermined location for controlling the robot (100). However, as described above, the robot (100) may be a stationary robot rather than a mobile robot, or may include a virtual robot agent, and therefore, redundant descriptions in this regard may be omitted.

[0072]

[0073] FIG. 2 is a block diagram illustrating a computer system that performs a method for controlling a robot based on trajectory optimization, according to one embodiment.

[0074] The computer system (200) is a computing device that performs a method of controlling the robot (100) of the embodiment, and may be, for example, a computer device included in a robot control system (120). Alternatively, depending on the embodiment, the computer system (200) may be a computer device included in the robot (100). Alternatively, depending on the embodiment, the computer system (200) may be a server or other computer device for controlling a robot agent.

[0075] As illustrated in FIG. 2, the computer system (200) may include, as components, a memory (210), a processor (220), a communication interface (230), and an input / output interface (240). The input / output interface (240) may communicate with an input / output device (250) within the computer system (200) or separate from the computer system (200).

[0076] The memory (210) is a computer-readable recording medium, and may include a random access memory (RAM), a read only memory (ROM), and a permanent mass storage device such as a disk drive. Here, the ROM and the permanent mass storage device such as the disk drive may be included in the computer system (200) as a separate permanent storage device distinct from the memory (210). In addition, the memory (210) may store an operating system and at least one program code. These software components may be loaded into the memory (210) from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. In another embodiment, the software components may be loaded into the memory (210) through a communication interface (230) rather than a computer-readable recording medium. For example, software components may be loaded into the memory (210) of a computer system (200) based on a computer program that is installed by files received over a network (270).

[0077] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (220) via the memory (210) or the communication interface (230). For example, the processor (220) may be configured to execute instructions received according to program code stored in a storage device such as the memory (210).

[0078] The communication interface (230) may provide a function for the computer system (200) to communicate with other devices via a network (270). For example, requests, commands, data, files, etc. generated by the processor (220) of the computer system (200) according to program codes stored in a recording device such as a memory (210) may be transmitted to other devices via the network (270) under the control of the communication interface (230). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer system (200) via the communication interface (230) of the computer system (200) via the network (270). The signals, commands, data, etc. received via the communication interface (230) may be transmitted to the processor (220) or the memory (210), and the files, etc. may be stored in a storage medium (the aforementioned permanent storage device) that the computer system (200) may further include.

[0079] The communication method through the communication interface (230) is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, a wired Internet, a wireless Internet, a broadcasting network) that the network (270) may include, but also a short-range wired / wireless communication between devices. For example, the network (270) may include any one or more of a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), the Internet, and the like. In addition, the network (270) may include any one or more of a network topology including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree, or a hierarchical network.

[0080] The input / output interface (240) may be a means for interfacing with an input / output device (250). For example, the input device may include a device such as a microphone, keyboard, camera, or mouse, and the output device may include a device such as a display or speaker. As another example, the input / output interface (240) may be a means for interfacing with a device that integrates input and output functions, such as a touchscreen. The input / output device (250) may also be configured as a single device with the computer system (200).

[0081] Additionally, in other embodiments, the computer system (200) may include fewer or more components than those illustrated in FIG. 2. However, most conventional components need not be explicitly depicted. For example, the computer system (200) may be implemented to include at least some of the input / output devices (250) described above, or may further include other components such as a transceiver, a camera, various sensors, a database, and the like.

[0082] The processor (220) of the computer system (200) can identify a first state of the robot (100), as described below, obtain a plurality of actions that the robot (100) can perform in order to predict states after the first state, determine a cost function (or reward function) for determining a cost incurred or a reward obtained as the robot (100) performs the action in the first state, and can be configured to perform actions and steps for updating the obtained actions into actions that minimize (or maximize) the cost (or reward).

[0083] In the detailed description to be described later, for the convenience of explanation, the above steps are described as being performed by the computer system (200) without distinguishing between the robot (100) and the robot control system (120), and the operations performed by the processor (220) or other components of the computer system (200) or the operations performed by the application / program executed by the processor (220) may be described as operations performed by the computer system (200) for the convenience of explanation.

[0084] Above, the technical features described above with reference to Fig. 1 can also be applied to Fig. 2, so redundant descriptions are omitted.

[0085]

[0086] FIG. 3 is a block diagram illustrating a robot controlled according to trajectory optimization, according to one embodiment.

[0087] As described above, the robot (100) may be a mobile robot, such as a service robot used to provide services within a space (10), or may be a stationary robot.

[0088] The robot (100) may be a physical device and may include a control unit (104), a driving unit (108), a sensor unit (106), and a communication unit (102), as illustrated.

[0089] Alternatively, the robot (100) described in the embodiment may represent a robot agent, wherein the robot agent may not include the components (102 to 108) and may virtually include elements corresponding to the control unit (104), the drive unit (108), the sensor unit (106), and the communication unit (102).

[0090] The control unit (104) may be a physical processor built into the robot (100), and although not separately illustrated, may include a path planning processing module, a mapping processing module, a driving control module, a localization processing module, a data processing module, and a service processing module. In this case, the path planning processing module, the mapping processing module, and the localization processing module may be selectively included in the control unit (104) according to an embodiment to enable indoor autonomous driving of the robot (100) even when communication with the robot control system (120) is not established.

[0091] The communication unit (102) may be a configuration for the robot (100) to communicate with other devices (such as other robots, a computer system (200), or a robot control system (120)). In other words, the communication unit (102) may be a hardware module, such as an antenna, a data bus, a network interface card, a network interface chip, and a networking interface port of the robot (100), or a software module, such as a network device driver or a networking program, that transmits / receives data and / or information to / from other devices.

[0092] The driving unit (108) controls the movement of the robot (100) and may include equipment for performing the movement as a configuration that enables the movement. For example, the driving unit (108) may include a motor that is driven to move the robot (100) in a straight line or rotate it.

[0093] The sensor unit (106) may be configured to collect data required for the operation of the robot (100). The sensor unit (106) may not include expensive sensing equipment, but may only include sensors such as low-cost ultrasonic sensors and / or low-cost cameras. The sensor unit (106) may include sensors for identifying other robots or people in front and / or behind. For example, other robots, people, and other objects may be identified as obstacles (30) through the camera of the sensor unit (106). Alternatively, the sensor unit (106) may include an infrared sensor (or infrared camera). In addition to the camera, the sensor unit (106) may further include sensors for recognizing / identifying users, other robots, or objects in the vicinity. In this way, the sensor unit (106) may be configured to identify obstacles (30).

[0094] In the case where the robot (100) is a mobile robot, it may be distinguished from a mapping robot used to create an indoor map within a space (10). Since the robot (100) does not include expensive sensing equipment, it can process indoor autonomous driving using the output values ​​of sensors such as low-cost ultrasonic sensors and / or low-cost cameras. Meanwhile, if the robot (100) has previously processed indoor autonomous driving through communication with a robot control system (120), more accurate indoor autonomous driving may be possible using low-cost sensors by further utilizing mapping data included in the route data previously received from the robot control system (120). However, depending on the embodiment, the robot (100) may also function as the mapping robot.

[0095] Meanwhile, the driving unit (108) may further include not only equipment for moving the robot (100), but also equipment related to services provided by the robot (100) or equipment related to tasks performed by the robot (100). For example, the driving unit (108) may include at least one manipulator or robot arm, and may include at least one actuator for operating the manipulator or robot arm.

[0096] Meanwhile, if the robot (100) only provides sensing data for controlling the robot (100) to the robot control system (120), and the control of the robot (100) is performed through the robot control system (120), the robot (100) may be a brainless robot.

[0097] Alternatively, depending on the embodiment, a method of controlling the robot (100) of the embodiment may be performed from the robot (100) side, and at this time, the robot (100) may be configured to include the computer system (100) described above or at least include a configuration such as the processor (220) of the computer system (100).

[0098] The robot (100) may have different sizes and shapes depending on the type of machine, the service provided, the task used, etc.

[0099] The configuration and operation of the robot control system (120) that controls the robot (100) will be described in more detail with reference to FIGS. 4 and 5, which will be described later.

[0100] The description of the technical features described above with reference to FIGS. 1 and 2 can also be applied to FIG. 3, so any redundant description will be omitted.

[0101]

[0102] FIGS. 4A and 4B are block diagrams illustrating a robot control system for controlling a robot according to one embodiment.

[0103] The robot control system (120) may be a device that controls the movement, provision of services, and performance of tasks of the aforementioned robot (100).

[0104] The robot control system (120) may include at least one computing device and may be implemented as a server (i.e., a cloud server) located within the space (10) or outside the space (10). The robot control system (120) may include the computer system (200) described above.

[0105] The robot control system (120) may include a memory (330), a processor (320), a communication unit (310), and an input / output interface (340), as illustrated. The description of the memory (330), the processor (320), the communication unit (310), and the input / output interface (340) described above with reference to FIG. 2 may be similarly applied to the memory (210), the processor (220), the communication interface (230), and the input / output interface (240), and thus, redundant descriptions are omitted.

[0106] In other embodiments, the robot control system (120) may include more components than those illustrated.

[0107] The robot control system (120) can be configured to control a robot (100) which is a mobile robot, and in relation to this, the configurations (410 to 440) of the processor (320) are described in more detail with reference to FIG. 4b.

[0108] The processor (320) may include a map generation module (410), a localization processing module (420), a path planning processing module (430), and a service operation module (440), as illustrated. The components included in the processor (320) may be expressions of different functions performed by at least one processor included in the processor (320) according to control instructions based on the code of an operating system or the code of at least one computer program.

[0109] The map generation module (410) may be a component for generating an indoor map of a target facility using sensing data generated by a mapping robot (not shown) autonomously driving within the space (10) about the target facility (e.g., the interior of the space (10).

[0110] At this time, the localization processing module (420) can determine the location of the robot (100) inside the target facility by using sensing data received from the robot (100) through the network and the indoor map of the target facility generated through the map generation module (410).

[0111] The path planning processing module (430) can generate a control signal for controlling indoor autonomous driving of the robot (100) using the sensing data received from the robot (100) and the generated indoor map. For example, the path planning processing module (430) can generate a path (i.e., path data) of the robot (100). The generated path (path data) can be set for the robot (100) for driving the robot (100) along the path. The robot control system (120) can transmit information about the generated path to the robot (100) via a network. For example, the path information can include information indicating the current location of the robot (100), information for mapping the current location with the indoor map, and path planning information. The path information can include information about a path that the robot (100) should drive to reach a predetermined location within the space (10) or to provide a service to a predetermined user. The path planning processing module (430) can set a path (i.e., path data) for the robot (100). The robot control system (120) can control the movement of the robot (100) so that the robot (100) moves according to the set path (i.e., along the set path). The 'trajectory optimization' of the embodiment may be determining an optimal path for the robot (100) to move and determining optimal control parameter(s) of the robot (100) when driving along the optimal path.

[0112] The service operation module (440) may include a function for controlling the service provided by the robot (100) within the space (10). For example, the robot control system (120) or the service provider operating the space (10) may provide an IDE (Integrated Development Environment) for the service (e.g., cloud service) provided by the robot control system (120) to the user or manufacturer of the robot (100). In this case, the user or manufacturer of the robot (100) may create software for controlling the service provided by the robot (100) within the space (10) through the IDE and register the software in the robot control system (120). In this case, the service operation module (440) may control the service provided by the robot (100) using the software registered in connection with the robot (100).

[0113] The description of the technical features described above with reference to FIGS. 1 to 3 can also be applied to FIGS. 4a and 4b, so redundant descriptions are omitted.

[0114]

[0115] In the detailed description to be provided below, for convenience of explanation, steps and / or operations performed by the processor (220) and the like of the method for generating a learning model of the embodiment are described as operations performed by the computer system (200) including the processor (220).

[0116]

[0117] In addition, in the detailed description to be described later, for the convenience of explanation, steps and / or operations performed by a processor (220) or the like of a method for controlling a robot (100) of the embodiment are described as operations performed by a computer system (200), without distinguishing between the robot (100) and the robot control system (120).

[0118]

[0119] FIG. 5 is a flowchart illustrating a method for controlling a robot based on trajectory optimization according to optimization of compensation or cost, according to one embodiment.

[0120] Referring to FIG. 5, a robot control method performed by a computer system (200) including a robot (100) or a robot control system (120) that controls the robot (100) is described.

[0121] In step (510), the computer system (200) can identify a first state of the robot (100). The first state may be the current state of the robot (100). Alternatively, the first state may be any state of the robot (100) and may be an initial state for performing trajectory optimization through the computer system (200).

[0122] In step (520), the computer system (200) can acquire behaviors of the robot (100) for predicting the states of the robot (100) after the first state. The acquired behaviors may be a sequence of behaviors. The behaviors may be named actions. For example, the computer system (200) can acquire the behaviors of step (520) by acquiring a certain number of behaviors through random sampling. The behaviors represent actions that the robot (100) can take in a specific situation, and when the robot (100) performs the behavior in a specific state, the robot (100) can transition to the next state. These behaviors may be targets of updates through trajectory optimization of the embodiment, and may be updated with optimal behaviors for the robot (100) to achieve a goal (e.g., moving to a destination (50)) according to the trajectory optimization.

[0123] In step (530), the computer system (200) can predict the states of the robot (100) according to the acquired actions using a dynamic model. The dynamic model can be defined in a form such as f(s, a). Here, s can represent the state of the robot (100), and a can represent a specific action taken by the robot (100). f(s, a) can be a model for predicting the next state of the state s of the robot (100). In other words, f(s, a) can represent a state transition function when a specific action a is taken in a given state s, and accordingly, the next state can be determined based on the specific state and the taken action.

[0124] In step (540), the computer system (200) may determine a cost function (or reward function) to determine the cost incurred when the robot (100) takes the actions in the first state (or to determine the cost obtained when the robot (100) takes the actions in the first state) by using a cost model (or reward model). The cost model may be defined in a form such as c(s,a), for example. c(s,a) may be a model for predicting the cost incurred when the robot (100) takes action a in state s. Meanwhile, the reward model may be defined in a form such as r(s,a), which may be a model for predicting the reward obtained when the robot (100) takes action a in state s.

[0125] The computer system (200) can determine a cost function (or reward function) by synthesizing cost models (or reward models) corresponding to each state and action.

[0126] For example, the computer system (200) may determine a cost function (or reward function) for determining the cost incurred (or reward obtained) when the robot (100) performs the actions in the first state based on the first state, states subsequent to the first state, and actions (sequence of actions) for predicting the states. The reward function may be, for example, for determining the reward obtained by the robot (100) in reinforcement learning.

[0127] The determined cost function or reward function can be used for trajectory optimization. For example, the computer system (200) can determine actions (updated actions) that minimize the cost by updating the actions in a direction to minimize the cost determined by the cost function. Alternatively, the computer system (200) can determine actions (updated actions) that maximize the reward by updating the actions in a direction to maximize the reward determined by the reward function. For reference, in the detailed description to be described below, embodiments will be described based on 'cost' and 'cost function', and overlapping descriptions related to 'reward' and 'reward function' may be omitted.

[0128] At step (550), the computer system (200) may update actions to minimize (or maximize) a cost (or reward) based on a gradient value calculated for the cost function (or reward function) under constraints related to the control of the robot (100). For example, the finally determined updated actions may represent a sequence of actions that minimizes the cost of the robot (100) in the first state.

[0129] At step (560), the computer system (200) can control the robot (100) based on at least one of the updated actions. The computer system (200) can control the robot (100) to enter the next state by performing an action for transitioning from the first state to the next state among the updated actions.

[0130] The above constraints may be essential constraints (hard constraints) indicating constraints based on the hardware specifications of the robot (100), user needs, or safety constraints in the operating space (10). The robot (100) may be controlled subject to these constraints. For example, these constraints may include a limitation that the robot's movement speed cannot exceed a certain level, for example, 0.7 m / s, due to motor limitations or safety issues.

[0131] These constraints (hard constraints) are constraints that are essential for the safe and efficient operation of the robot (100), and may represent physical and mechanical limitations of the robot (100) or / and may additionally include rules and safety regulations specific to the working environment of the robot (100).

[0132] Among the constraints related to the mechanical limitations of the robot (100), there may be restrictions according to the mechanical design of the robot (100), such as the speed, acceleration, and turning radius of the robot (100). These may be constraints to prevent the robot (100) from becoming uncontrollable or being damaged during its operation. For example, the constraints may include at least one of the maximum speed of the robot (100) (maximum linear movement speed), the maximum acceleration of the robot (100), the maximum angular velocity (maximum rotation speed) of the robot (100), the maximum angular acceleration of the robot (100), and the maximum movement angle of a manipulator or robot arm included in the robot (100).

[0133] Among the constraints, those related to the energy consumption of the robot (100) may include the battery consumption of the robot (100) and the operating time according to the battery capacity of the robot (100).

[0134] Additionally, the constraints may include limitations related to the communication range of the robot (100).

[0135] Additionally, constraints may include restrictions on the path of the robot (100) to areas where access by the robot (100) is prohibited or restricted due to environmental, legal or task requirements.

[0136] The aforementioned cost function and dynamic model for predicting the state of the robot (100) based on its actions can be constructed to take these constraints into account in controlling the robot (100). At this time, the constructed cost function may include sections where differentiation (toward actions) is impossible due to the aforementioned constraints. In other words, the cost function may include discontinuous sections.

[0137] In other words, the cost function determined in step (540) may include discontinuous intervals, which are intervals where actions cannot be differentiated, due to the above constraints. The computer system (100) can determine the gradient of the cost function in the interval where the cost function cannot be differentiated as a predetermined value.

[0138] Below, we describe in more detail how to determine updated actions (i.e., how to update the actions obtained in step (520)) by computing the gradient of the cost function.

[0139] In step (522), the computer system (200) can calculate the gradient in the differentiable and non-differentiable sections of the cost function. The computer system (200) can calculate the gradient by partially differentiating the reward function with respect to the action (a). For the discontinuous section of the cost function, i.e., the non-differentiable section, the computer system (200) can determine the gradient of the non-differentiable section as the predetermined value using an STE (Straight Through Estimator)-based algorithm. For example, the gradient of the non-differentiable section of the cost function can be determined as the predetermined value 1.

[0140] Meanwhile, the computer system (200) can update actions to minimize (or maximize) the cost (or reward) based on the gradient of the cost function calculated in step (522) using a gradient descent algorithm (or gradient ascent algorithm).

[0141] At step (554), the computer system (200) can update the actions by changing the values ​​of the actions in the direction (or opposite direction) of the calculated gradient of the cost function.

[0142] In step (556), the computer system (200) can determine the final updated actions by repeating step (554) by updating the actions by changing the values ​​of the actions until the cost (or reward) calculated through the cost function converges. Convergence of the cost (or reward) may mean that the amount of change in the cost (or reward) falls below a certain level.

[0143] That is, the computer system (200) can use a gradient descent algorithm (or a gradient ascent algorithm) to determine actions that minimize cost (or maximize reward).

[0144] The robot (100) can be controlled based on these finally updated behaviors.

[0145] Below we describe in more detail how to perform trajectory optimization based on gradient descent (the gradient descent algorithm described above).

[0146] Gradient descent can be a method of determining the minimum value of a target function by using the gradient (i.e., slope) of the target function (in this embodiment, the cost function). The gradient of the target function can indicate the direction toward the local minimum of the target function, and by moving the value based on the direction of the gradient, the minimum value can be found, and the values ​​of the corresponding parameters can be determined. For example, the computer system (200) can select an arbitrary starting point for the parameters of the target function to be optimized, calculate the gradient of the function at the corresponding parameter using differentiation, and little by little update the parameters (e.g., in the opposite direction of the gradient) according to the direction of the calculated gradient. The computer system (200) can determine whether the target function converges (for example, whether the value is less than a certain value or the amount of change is less than a certain level) according to the parameter update, and repeat the parameter update until convergence to finally determine the updated parameters.

[0147] In the trajectory / path optimization of the robot (100) of the embodiment, as described above, the computer system (200) can define a cost function, calculate the gradient of this cost function, and then gradually adjust the path (i.e., sequence of actions) of the robot (100) to determine a path that minimizes the cost. The cost function can be a function for calculating the total cost for a path from the current state of the robot (100) (e.g., the current position or the first state described above) to a goal state (e.g., the destination (50)). This cost function can calculate the total cost according to the movement path of the robot (100), and can generally be defined by considering parameters such as the position of the robot (100), the movement direction, energy consumption according to the movement, and the distance from an obstacle. When the robot (100) must move while avoiding an obstacle, the cost function can be designed so that the cost increases according to the distance from the obstacle.

[0148] The computer system (200) can establish an initial path from the first state of the robot (100) to the alma mater state. This initial path may be a certain number of actions (i.e., a sequence of actions) acquired through the aforementioned random sampling. The computer system (200) can calculate the gradient of the cost function for these acquired actions using differentiation. The gradient may indicate how the cost changes in each state corresponding to the actions. The computer system (200) can update the actions little by little (e.g., in the opposite direction of the gradient) according to the direction of the calculated gradient. The computer system (200) can determine whether the target function converges (e.g., becomes a value below a certain value or the amount of change is below a certain level) based on the updated actions, and can repeat the update of the actions until convergence occurs to determine the final updated parameters. At this time, the computer system (200) can repeat the process of recalculating the gradient for the new path according to the updated actions and updating the actions. For example, after the step (554) described above, the gradient of the cost function for determining the cost (or reward) incurred by performing the updated actions in step (554) can be recalculated, and the actions can be further updated using this.

[0149] Meanwhile, in the embodiment, the gradient descent method (or gradient ascent method) can be applied to the cost function to determine updated actions, and for a section where the cost function is not differentiable, the gradient is determined to a predetermined value using an STE-based algorithm, so that the determination of updated actions through the gradient descent method can be made for the entire section of the cost function.

[0150] Below, we describe STE-based algorithms in more detail. A Straight Through Estimator (STE) can be a method for estimating gradients for non-differentiable functions and enabling backpropagation. For example, functions that are non-differentiable by STE (e.g., non-differentiable intervals of the cost function) can be assumed to have a gradient of 1.

[0151] The control of the robot (100) of the embodiment may be a navigation of the robot (100) to control the movement of the robot (100) from the current location of the robot (100) to the destination (50). The first state described above may represent the current location of the robot (100) or the control state of the robot (100) at the current location. The states after the first state described above may represent locations to which the robot (100) will move from the current location or the control state of the robot (100) at the locations to which it will move. The control state may include at least one of a speed state of the robot (100), an acceleration state of the speed, an angular velocity state of a driving unit for moving the robot (100), and an angular acceleration state of the angular velocity.

[0152] Each of the actions for predicting states after the first state described above may include at least one of the speed of the robot (100), the acceleration of the speed, the angular velocity of the driving unit for moving the robot (100), and the angular acceleration of the angular velocity, as a control parameter of the robot (100).

[0153] The trajectory optimization of the above-described embodiment can be similarly applied to the operation of the robot arm or manipulator of a stationary robot (100) in addition to the navigation of the robot (100).

[0154] The description of the technical features described above with reference to FIGS. 1 to 4 can also be applied to FIG. 5, so any duplicate description will be omitted.

[0155]

[0156] Figure 6 is a flowchart illustrating a method for controlling a robot based on updated actions (sequence of actions), according to an example.

[0157] At step (610), the computer system (200) can determine a second state following the first state of the robot (100) based on the actions updated by step (550).

[0158] At step (620), the computer system (200) can control the robot (100) from the first state to the second state. At this time, the robot (100) can perform an action (i.e., among the updated actions) to transition from the first state to the second state.

[0159] The computer system (200) can transition from a first state to subsequent states by sequentially performing each of the updated actions, thereby completing the goal at a minimum cost (or, to obtain a maximum reward).

[0160] Alternatively, as in step (630), the computer system (200) may perform the steps of setting the second state to the first state, re-acquiring the actions of the robot (100); determining a cost function; and updating the actions. That is, steps (510 to 550) described above with reference to FIG. 5 may be repeatedly performed for the second state. Through this repetition, a third state, which is a state subsequent to the second state, may be determined, and subsequent states may also be similarly determined. That is, steps (510 to 550) may be repeated each time the robot (100) transitions to each state.

[0161] The description of the technical features described above with reference to FIGS. 1 to 5 can also be applied to FIG. 6, so any duplicate description will be omitted.

[0162]

[0163] FIG. 7 illustrates a method for controlling a robot based on trajectory optimization by updating action parameters such as speed using gradient calculation and decision for a cost function, according to an example.

[0164] Figure 8 shows a model of the state of a robot that is the target of trajectory optimization according to an example.

[0165] A method of updating actions to minimize / maximize cost / reward based on a gradient value calculated for the corresponding cost function under constraints related to the control of a robot (100) in an embodiment and controlling the robot (100) based on at least one of the updated actions is described in more detail using mathematical formulas.

[0166] Referring to FIG. 5, the first state in the aforementioned step (510) is the state of the robot (100) at a specific point in time t, s t can be defined as. At this time, actions for predicting states after the first state can be defined as mathematical expression 1. In mathematical expression 1, a sequence A of actions of the robot (100) can be defined. Meanwhile, the first state s t The following states of the above robot can be defined by the model of mathematical expression 2.

[0167] [Mathematical Formula 1]

[0168]

[0169] [Equation 2]

[0170]

[0171] a may represent the behavior of the robot (100), and t may represent a specific point in time. f may be a dynamic model for predicting the next state of a specific state. The behaviors included in the behavior sequence A may be randomly sampled.

[0172] The cost function determined in the aforementioned step (540) can be defined by mathematical expression 3.

[0173] [Equation 3]

[0174]

[0175] c(s t , a t ) is the first state s t In the robot (100) action a t It may be a model for determining the cost incurred in performing the operation (i.e., the cost model described above).

[0176] Meanwhile, mathematical expression 3 can also be expressed as mathematical expression 3-1 below.

[0177] [Equation 3-1]

[0178]

[0179] Alternatively, if we build a cost function based on a score (c=r(s)) indicating whether a particular state s matches the action that the robot (100) should perform, then the cost function C = r(s t )+r(s t+1 )+ ... r(s t+h ) can be constructed in the same form.

[0180] An updated sequence A of the above actions A to minimize the cost determined by the above cost function. * can be defined by mathematical formula 4.

[0181] [Equation 4]

[0182]

[0183] In an embodiment, the computer system (200) can determine A* by calculating the gradient of C with respect to A and iteratively modifying (i.e., updating) A in the direction of the calculated gradient (or, in some cases, in the opposite direction of the gradient). With each iteration, A can be updated in a direction that increasingly lowers C.

[0184] The optimization process of updating A to optimize C as described above may correspond to the orbit optimization of the embodiment. Meanwhile, after orbit optimization, action a t+1 The robot (100) executes the next state s t+1 The process of performing optimization control by repeating the process can be called model predictive control.

[0185] In an embodiment, for example, the updated sequence A * In calculating , action a t The updated action of the above cost function is action a t Gradient differentiated by can be determined based on the gradient can be calculated by mathematical formula 5.

[0186] [Equation 5]

[0187] ,

[0188] Mathematical formula 5 can be calculated by mathematical formula 6.

[0189] [Equation 6]

[0190]

[0191] Meanwhile, for the convenience of explanation, the first term of the above mathematical expression 5 Although only the method of calculating is explained through mathematical formula 6, the remaining terms are also Since the relationship holds, it can be calculated in a similar way. That is, the remaining terms can also be calculated through repeated application of the same method by the chain rule.

[0192] In the embodiment, since constraints exist in the control of the robot (100), the cost function may include a non-differentiable section. In this case, according to the aforementioned STE, the gradient of this non-differentiable section may be assumed to be 1 as a predetermined value.

[0193] Meanwhile, according to the state modeling of the robot (100) described in Fig. 8, the first state of the robot (100) is s t can be defined as in mathematical formula 7.

[0194] [Equation 7]

[0195]

[0196] x t and y t is the position coordinate of the robot (100) at time t, is the position angle of the robot (100) at time t, v t is the (linear) velocity of the robot (100) at time t, w t can represent the angular velocity (i.e., rotational velocity) of the robot (100) at time point t.

[0197] At this time, the first state is s t The following state of the robot s t+1 can be defined by mathematical expression 8.

[0198] [Equation 8]

[0199]

[0200] The above dt may represent a timescale (or time step) between time points t and t+1. For example, dt may represent how many seconds actually elapsed between time points t and t+1. For example, if the goal is to have the robot (100) at a specific target coordinate at time point 5*dt, the cost function may be configured to measure how close the coordinate of the robot (100) at time point 5*dt is to the target coordinate. According to the trajectory optimization, the action sequences of the robot (100) that will be located at the target coordinate after 5*dt (or, closest to the target coordinate) may be determined.

[0201] above v t target is the target velocity of the robot (100) at time t, and w t target may be the target angular velocity of the robot (100) at time t. That is, the action a representing the velocity that the robot (100) should aim for at time t is [v t target , w t target ] can be defined as above a v is a constraint, which may be an acceleration constraint of the speed of the robot (100), and the above a w is a constraint, which may be a constraint of the angular acceleration of the angular velocity of the robot (100). The acceleration constraint may indicate an acceleration range or maximum acceleration in which the robot (100) can operate, and the angular acceleration constraint may indicate an angular acceleration range or maximum angular acceleration in which the robot (100) can operate. For example, the constraint may indicate a limit of the robot's (100) ability to accelerate / decelerate within a given time step dt, and the robot (100) may be a v Wow a w Acceleration / deceleration may be possible only under the constraints of .

[0202] Depending on the existence of constraints such as the above, v and w can be defined by min functions, and these min functions can make the cost function non-differentiable.

[0203] In the embodiment, in the cost function, the target speed v t target Go v t +a v *The gradient of the interval corresponding to the case greater than dt is, As cannot be calculated, it can be determined as a given value (e.g., 1). Similarly, the target angular velocity w t target Go w t +a w *The gradient of the interval corresponding to the case greater than dt can be determined as a given value (e.g., 1). That is, according to the application of the STE-based algorithm, the velocity v t target Go v t +a v *If it is greater than dt can be approximated as

[0204] Accordingly, an approximate gradient value that exhibits sufficiently excellent performance in trajectory optimization of a mobile robot (100) can be calculated.

[0205] The embodiment is a method for performing trajectory optimization for a robot (100) by taking into account constraints using Lagrange multipliers or Lagrange coefficients, in which a separate Lagrange multiplier is used. It is possible to determine the predicted future state values ​​(e.g., the position and velocity of the robot (100)) that comply with the constraints without requiring procedures such as tuning values. In other words, in the case of the method using Lagrange multipliers, there is a problem that it cannot be guaranteed that the A* determined as a result satisfies the actual constraints, for example, a separate Lagrange multiplier For small values, A* will likely ignore the constraints, and the Lagrange multipliers In case of large values, A* has a problem of focusing too much on complying with constraints, whereas the embodiment can solve this problem.

[0206] In this embodiment, in trajectory optimization where there is a hard constraint, by applying an STE-based algorithm, optimization of behavior using a gradient can be performed while ensuring that the behavior and state of the robot (100) derived in the optimization process do not violate the hard constraint.

[0207] In the embodiment, compared to the method using the aforementioned Lagrange coefficients, there is no need to tune the Lagrange coefficients and trajectory optimization that satisfies essential constraints can be performed with fewer iterations.

[0208] In this way, in the embodiment, when specifying a target speed for optimal control of a robot (100), which is a mobile robot, constraints regarding the maximum speed that the robot (100) can produce and the acceleration for the robot (100) to reach the target speed can be observed, and for this purpose, as described above, for example, optimization can be performed by applying an STE-based algorithm to a dynamic model that predicts the future speed of the robot (100) from a sequence of actions of the robot (100) (sequence of target speeds).

[0209] As another example, the maximum speed v of the robot (100) max and minimum speed v min We further describe an example of optimization using an STE-based algorithm in the presence of constraints.

[0210] In mathematical equation 9 below, v target is the target velocity of the robot (100) at time t, and v tis the velocity of the robot (100) at time t, a is the acceleration of the robot (100), and dt may be a time scale.

[0211] [Equation 9]

[0212]

[0213] Due to the constraints of maximum and minimum speed, the above mathematical expression 9 is v target Since differentiation is not possible, an STE-based algorithm can be applied to make this possible. Accordingly, the results of Equations 10 and 11 below can be derived.

[0214] [Equation 10]

[0215]

[0216] [Equation 11]

[0217]

[0218] The description of the technical features described above with reference to FIGS. 1 to 6 can also be applied to FIG. 7, so redundant descriptions are omitted.

[0219] The systems or devices described above may be implemented as hardware components, software components, or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0220] Software may include computer programs, codes, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0221] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0222] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0223] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. A robot control method performed by a robot or a robot control system that controls a robot, A step of identifying a first state of the robot; A step of obtaining the actions of the robot for predicting the states of the robot after the first state; and A step of determining a cost function for determining a cost incurred or a reward obtained as the robot performs the actions in the first state based on the first state, the states, and the actions; and A step of updating the actions to minimize or maximize the cost or the reward based on the gradient value calculated for the cost function under constraints related to the control of the robot. Including, A robot control method, wherein the above cost function includes an interval in which the actions cannot be differentiated due to the above constraints, and the gradient of the interval in which the actions cannot be differentiated is determined to a predetermined value.

2. In paragraph 1, A step of controlling the robot based on at least one of the updated above actions. A robot control method further comprising:

3. In paragraph 2, The control of the above robot is to control the movement of the robot from the current location of the robot to the destination, The first state represents the current position or the control state of the robot at the current position, and the states represent the positions to which the robot will move from the current position or the control state of the robot at the positions to which the robot will move. A robot control method, wherein each of the above actions includes at least one of a speed of the robot, an acceleration of the speed, an angular velocity of a driving unit for moving the robot, and an angular acceleration of the angular velocity.

4. In paragraph 1, The above constraints are, A robot control method comprising at least one of a maximum speed of the robot, a maximum acceleration of the robot, an angular velocity of the robot, a maximum angular acceleration of the robot, and a maximum movement angle of a manipulator or robot arm included in the robot.

5. In paragraph 1, The above updating steps are: A step of determining the gradient of the non-differentiable section as the predetermined value using an algorithm based on STE (Straight Through Estimator). Including, The above updating steps are: A robot control method that updates the actions to minimize or maximize the cost or reward based on the gradient of the cost function using a gradient descent algorithm or a gradient ascent algorithm.

6. In paragraph 5, A robot control method, wherein the gradient of the above-mentioned non-differentiable section is determined as the above-mentioned predetermined value of 1.

7. In paragraph 1, The step of acquiring the above robot behaviors is to acquire the above behaviors by acquiring a certain number of behaviors through random sampling, The above updating steps are: A step of updating the actions by changing the values ​​of the actions in the direction or opposite direction of the calculated gradient; A step of determining the final updated actions by repeating the step of updating the actions by changing the values ​​of the actions until the cost or reward calculated through the cost function converges. Including, A robot control method, wherein the robot is controlled based on the above finally updated actions.

8. In paragraph 1, The steps for controlling the above robot are: A step of determining a second state following the first state of the robot based on the updated actions; and A step of controlling the robot from the first state to the second state A robot control method comprising:

9. In paragraph 8, A step of obtaining the actions of the robot by setting the second state to the first state; a step of determining the cost function; and a step of repeating the step of updating the actions. A robot control method further comprising:

10. In paragraph 1, The above actions are a sequence A of the actions of the robot in mathematical expression 1, s is the first state t The following states of the above robot are defined by the model of mathematical expression 2, [Mathematical Formula 1] , [Equation 2] , a represents the behavior of the robot, t represents a specific point in time, The above cost function is defined by mathematical expression 3, [Equation 3] , c(s t , a t ) is the first state s t In the above robot acts a t It is a model for determining the cost incurred in performing An updated sequence A of the above actions to minimize the above cost * is defined by mathematical formula 4, [Equation 4] , robot control method.

11. In paragraph 10, The above updated sequence A * middle act as a t The updated action of the above cost function is action a t Gradient differentiated by is determined based on the gradient is calculated by mathematical formula 5, [Equation 5] , Mathematical formula 5 is calculated by mathematical formula 6, [Equation 6] , robot control method.

12. In paragraph 10, The first state above is s t is defined by mathematical formula 7, [Equation 7] , x t and y t is the position coordinate of the robot at time t, is the position angle of the robot at time t, v t is the velocity of the robot at time t, w t represents the angular velocity of the robot at time t, The first state above is s t The following state of the robot s t+1 is defined by mathematical formula 8, [Equation 8] , The above dt is the timescale between time points t and t+1, and the above v t target is the target velocity of the robot at time t, and w t target is the target angular velocity of the robot at time t, and a v and the above a w A robot control method, wherein the above constraints are, respectively, an acceleration constraint of the speed of the robot and an angular acceleration constraint of the angular velocity of the robot.

13. In paragraph 12, The above acceleration constraints represent the acceleration range or maximum acceleration within which the robot can operate, A robot control method wherein the above angular acceleration constraint condition indicates an angular acceleration range or maximum angular acceleration within which the robot can operate.

14. In paragraph 12, In the above cost function, The target speed above is v t +a v *The gradient of the section corresponding to the case greater than dt is determined by the above-mentioned value, The target angular velocity above is w t +a w *A robot control method in which the gradient of the section corresponding to the case greater than dt is determined by the above-mentioned value.

15. In paragraph 1, A robot control method, wherein the above-mentioned non-differentiable section of the above-mentioned cost function is a section in which the above-mentioned cost function is discontinuous.

16. A non-transitory computer-readable recording medium having recorded thereon a program for executing the method of paragraph 1 in the robot or robot control system, which is a computer system.

17. In a computer system for a robot moving in space, At least one processor implemented to execute computer-readable instructions Including, At least one processor, Identifying a first state of the robot, obtaining actions of the robot for predicting states of the robot after the first state, determining a cost function for determining a cost incurred or a reward obtained as the robot performs the actions in the first state based on the first state, the states, and the actions, and updating the actions to minimize or maximize the cost or the reward based on a gradient value calculated for the cost function under constraints related to the control of the robot. A computer system in which the above cost function includes an interval in which the actions cannot be differentiated due to the above constraints, and the gradient of the interval in which the actions cannot be differentiated is determined to a predetermined value.

Citation Information

Patent Citations

  • An apparatus for selecting action and method thereof, computer-readable storage medium

    KR1020180089769A

  • Spicule with controlled pore volume and method for manufacturing the same

    KR1020250014879A

  • Compound targeting fibroblast-activation protein and use thereof

    KR1020250113264A

  • Quantization aware training method for neural networks that supplements limitations of gradient-based learning by adding gradient-independent updates

    KR102389910B1

  • KR20230173840A