Control method, control device, and control program

A control method and device using deep and reinforcement learning predict a trajectory for a robot to open and close doors with complex handle force changes, enhancing the robot's ability to operate folding doors.

JP2025179659APending Publication Date: 2025-12-10TOKYO UNIVERSITY OF SCIENCE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024086557
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably open and close doors, particularly folding doors, where the direction and magnitude of force required to operate the handle change complexly, limiting the deployment of service robots.

Method used

A control method and device that utilize a prediction model trained through deep learning and reinforcement learning to predict a trajectory for the end effector to move from the handle change complex manner, the robot 50 can open and close the door, even in the case of a door in which the direction and magnitude of the force that must be applied to the handle to open and close the door change in a complex manner.

Benefits of technology

The method and device that utilize a prediction model trained through deep learning and reinforcement learning to predict a trajectory consisting of multiple waypoints that the end effector to move from the current position to a target position where the open/closed state of the door 30 can be increased by implementing the said and the said and the said and the said and the door 30 can be opened and closed with high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179659000001_ABST
    Figure 2025179659000001_ABST
Patent Text Reader

Abstract

To increase a probability that a robot opens or closes a door in a case of targeting the door for which a direction and magnitude of force to be applied to a handle for opening and closing change in a complex manner.SOLUTION: A control method causes a computer to: acquire an image obtained by capturing a door to be opened / closed by an end effector provided on a manipulator; detect a handle position of the door from the image; calculate a current position of the end effector using kinematics; input the handle position and the current position into a prediction model to predict a trajectory composed of a plurality of via points until the end effector moves from the current position to a target position where an open / close state of the door is changed; control the manipulator on the basis of the trajectory; and repeatedly perform the image acquisition, the handle position detection, the current position calculation, the trajectory prediction, and the manipulator control.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a control method, a control device, and a control program. [Background technology]

[0002] Due to the global labor shortage, the number of service robots being installed is steadily increasing. Service robots are required to have two functions: the ability to move autonomously and the ability to grasp objects. Service robots use these functions to perform tasks such as security and transportation. However, doors separating rooms present obstacles, limiting the environments in which service robots can be deployed. Doors come in various types, including sliding doors, swing doors, and folding doors, each with its own opening and closing mechanism. A variety of research has been conducted on the opening operation of sliding and swing doors.

[0003] For example, Patent Document 1 states that "the information processing device (10) is an information processing device (10) that controls a robot that can grasp a door handle, and has a recognition means (103) that recognizes the shape of the door handle from an image of the door handle captured by the imaging means, and a control means (106) that controls the robot to grasp the door handle by an operation according to the shape of the handle recognized by the recognition means (103)."

[0004] Furthermore, Patent Document 2 states that "the remote control system comprises an imaging unit that captures an image of the environment in which an operated object equipped with an end effector exists; a recognition unit that recognizes a graspable part that can be grasped by the end effector based on the image of the environment captured by the imaging unit; an operation terminal that displays the captured image and accepts input of handwritten input information for the displayed captured image; and an estimation unit that estimates a graspable object from among the graspable parts that is required to be grasped by the end effector based on the graspable part recognized by the recognition unit and the handwritten input information input for the captured image, and estimates how the end effector is required to perform the grasping operation on the graspable object."

[0005] Patent Document 3 also states that "The robot 100 includes an imaging device 110d, a manipulator 110, and a controller 153 that controls the manipulator 110. The controller 153 identifies a location to which the end effector 110b installed on the manipulator 110 should be aligned based on at least a portion of the sensor data acquired from the imaging device 110d, and controls the manipulator 110 to move the end effector 110b along the surface of the location identified based on at least a portion of the sensor data."

[0006] Patent document 4 also states, "The system is configured to receive first image information generated when the camera has a first camera pose when an object is or was within the camera field of view of the camera. The system is further configured to determine a first estimate of the object structure based on the first image information, and identify object corners based on the first estimate of the object structure or based on the first image information. The system is further configured to cause the end effector device to move the camera to a second camera pose and receive second image information representing the structure of the object. The system is configured to determine a second estimate of the object structure based on the second image information, and generate a motion plan based on at least the second estimate." [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-139872 [Patent Document 2] Patent Publication No. 2021-94605 [Patent Document 3] Japanese Patent Publication No. 2022-50147 [Patent Document 4] Japanese Patent Publication No. 2022-44830 Summary of the Invention [Problem to be solved by the invention]

[0008] However, with the technologies of Patent Documents 1 and 2, when the target door is a folding door or other door in which the direction and magnitude of the force that must be applied to the handle to open and close it changes in a complex manner, the robot may not be able to open or close the door. Also, the technologies of Patent Documents 3 and 4 are not related to opening and closing doors.

[0009] The present disclosure has been made in consideration of the above circumstances, and aims to provide a method, device, and program that can increase the probability that a robot can open and close a door, even in the case of a door in which the direction and magnitude of the force that must be applied to the handle to open and close the door changes in a complex manner. [Means for solving the problem]

[0010] A control method according to a first aspect of the present disclosure includes a computer acquiring an image of a door to be opened or closed by an end effector attached to a manipulator, detecting a handle position of the door from the image, calculating a current position of the end effector using kinematics, inputting the handle position and the current position into a prediction model to predict a trajectory consisting of multiple waypoints that the end effector will take to move from the current position to a target position that will change the open / closed state of the door, controlling the manipulator based on the trajectory, and repeatedly performing the steps of acquiring the image, detecting the handle position, calculating the current position, predicting the trajectory, and controlling the manipulator.

[0011] A control method according to a second aspect of the present disclosure is the control method according to the first aspect, further comprising the computer learning the prediction model by deep learning.

[0012] A control method according to a third aspect of the present disclosure is the control method according to the second aspect, wherein learning by deep learning includes supervised learning using the trajectory geometrically created in advance as training data.

[0013] A control method according to a fourth aspect of the present disclosure is the control method according to the second or third aspect, further comprising the computer additionally learning the predictive model learned by deep learning by reinforcement learning.

[0014] A control method according to a fifth aspect of the present disclosure is the control method according to the fourth aspect, wherein the additional learning through reinforcement learning includes setting a reward according to at least one of the manipulability and joint torque of the manipulator.

[0015] A control method according to a sixth aspect of the present disclosure is a control method according to any one of the first to fifth aspects, in which detecting the steering wheel position includes detecting the steering wheel position from the image using an object detection algorithm.

[0016] A control method according to a seventh aspect of the present disclosure is the control method according to the sixth aspect, wherein detecting the steering wheel position includes calculating coordinates of the steering wheel position from coordinates of a partial area extracted by the object detection algorithm.

[0017] A control method according to an eighth aspect of the present disclosure is a control method according to any one of the first to seventh aspects, in which acquiring the image includes acquiring multiple images of the door, each taken at a different position.

[0018] A control method according to a ninth aspect of the present disclosure is the control method according to the eighth aspect, wherein acquiring the images includes acquiring the multiple images from a stereo camera provided on a robot having the manipulator.

[0019] A control method according to a tenth aspect of the present disclosure is the control method according to any one of the first to ninth aspects, wherein the door is a folding door.

[0020] A control device according to an eleventh aspect of the present disclosure includes a processor, which acquires an image of a door to be opened or closed by an end effector attached to a manipulator, detects the handle position of the door from the image, calculates the current position of the end effector using kinematics, inputs the handle position and the current position into a prediction model to predict a trajectory consisting of multiple waypoints for the end effector to move from the current position to a target position that changes the open / closed state of the door, controls the manipulator based on the trajectory, and repeatedly executes the processes of acquiring the image, detecting the handle position, calculating the current position, predicting the trajectory, and controlling the manipulator.

[0021] A control program according to a twelfth aspect of the present disclosure causes a computer to acquire an image of a door to be opened or closed by an end effector attached to a manipulator, detect the handle position of the door from the image, calculate the current position of the end effector using kinematics, input the handle position and the current position into a prediction model to predict a trajectory consisting of multiple waypoints for the end effector to move from the current position to a target position that changes the open / closed state of the door, control the manipulator based on the trajectory, and repeatedly execute the steps of acquiring the image, detecting the handle position, calculating the current position, predicting the trajectory, and controlling the manipulator.

[0022] According to the control method, control device, and control program disclosed herein, the probability that a robot can open and close a door can be increased even when the door being opened has complex changes in the direction and magnitude of the force that must be applied to the handle to open and close the door. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is a diagram showing an example of a schematic configuration of a door opening and closing system 10 according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an example of a hardware configuration of a control device 100 according to the present embodiment. [Figure 3] 1 is a diagram illustrating an example of a functional configuration of a control device 100 according to the present embodiment. [Figure 4] FIG. 2 is a diagram showing an example of a block diagram of a control device 100 according to the present embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of detecting a steering wheel position using an object detection algorithm. [Figure 6] FIG. 2 is a diagram showing an example of a control flow executed by the control device 100 according to the present embodiment. [Figure 7] FIG. 10 is a diagram showing an example of a simulation result in a reaching task. [Figure 8] FIG. 10 is a diagram showing an example of a simulation result in an opening task. DETAILED DESCRIPTION OF THE INVENTION

[0024] An example of an embodiment of the present disclosure will be described below with reference to the drawings. In each drawing, the same or equivalent components and parts are designated by the same reference numerals. Furthermore, the dimensional proportions in the drawings are exaggerated for the sake of explanation and may differ from the actual proportions.

[0025] 1 is a diagram showing an example of a schematic configuration of a door opening and closing system 10 according to this embodiment. The door opening and closing system 10 includes a door 30, a robot 50, and a control device 100.

[0026] The door 30 is a door that can be opened and closed. In this figure, the door 30 is shown as a folding door as an example. A folding door is a door that opens by folding sections.

[0027] There are various types of doors, such as sliding doors, swing doors, and folding doors, each with a different opening and closing mechanism. When opening and closing a sliding door, the handle moves in a substantially linear path. When opening and closing a swing door, the handle moves in a substantially arc-shaped path. When attempting to open or close these doors, the direction and magnitude of the force to be applied to the handle can be expressed by a relatively simple function.

[0028] In contrast, when opening and closing a folding door, the trajectory of the handle is a substantially elliptical arc. When attempting to open or close such a door, the direction and magnitude of the force to be applied to the handle cannot be expressed by a simple function and change in a complex manner over time. The technology of the present disclosure may be used to open and close doors in which the direction and magnitude of the force to be applied to the handle to open and close the door change in a complex manner. It goes without saying that the technology of the present disclosure can be applied to sliding doors, swing doors, etc., and can be similarly opened and closed.

[0029] The robot 50 opens and closes the door 30 under the control of the control device 100. In this figure, the robot 50 is shown as an autonomously moving robot with an autonomous movement function. The robot 50 has an automatic guided vehicle (AGV) 52, a manipulator 54, an end effector 56, and a stereo camera 58.

[0030] The automated guided vehicle 52 is a vehicle that travels automatically. The automated guided vehicle 52 may function as a platform on which the manipulator 54 and the stereo camera 58 that constitute the robot 50 are placed, and the position of the robot 50 may be freely moved.

[0031] The manipulator 54 is the part that operates the robot 50. In this figure, the manipulator 54 is shown as an example of a robot arm that operates using an articulated structure and a servo motor.

[0032] The end effector 56 is a device that is attached to the manipulator 54 and acts on a workpiece. In this figure, the end effector 56 is shown as a hook attached to the tip of the manipulator 54 as an example.

[0033] The stereo camera 58 is a camera that can capture images of a space from a plurality of different positions. By using images captured simultaneously at different positions using the stereo camera 58, it becomes possible to grasp the space three-dimensionally, including the depth direction. In this figure, as an example, the stereo camera 58 includes two cameras, a first camera 58_1 and a second camera 58_2.

[0034] The control device 100 is a device that controls the robot 50 to open and close the door 30. The control device 100 may be a terminal (such as a personal computer or tablet) separate from the robot 50, or may be integrated into the robot 50.

[0035] 2 is a diagram showing an example of the hardware configuration of the control device 100 according to this embodiment. The control device 100 includes a processor 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a storage 104, a communication interface 105, and a user interface 106. These components are connected via a bus 109 so as to be able to communicate with each other.

[0036] The processor 101 executes various programs and controls each component. Here, the processor 101 is assumed to be a CPU (Central Processing Unit). The ROM 102 stores various programs and various data. The RAM 103 temporarily stores programs or data as a working area. The storage 104 is configured with an HDD (Hard Disk Drive) or an SSD (Solid State Drive), and stores various programs including an operating system and various data.

[0037] In the control device 100 according to this embodiment, a control program is stored in the ROM 102 or the storage 104. The processor 101 reads the control program from the ROM 102 or the storage 104 and executes it using the RAM 103 as a working area, thereby controlling each component and performing various arithmetic processing in accordance with the control program.

[0038] The communication interface 105 is an interface through which the control device 100 communicates with other devices. The user interface 106 is an input / output interface through which the control device 100 exchanges information with a user. The user interface 106 may include input devices such as a mouse, keyboard, touch panel, and microphone, and output devices such as a monitor and speaker.

[0039] 3 is a diagram showing an example of the functional configuration of the control device 100 according to this embodiment. The control device 100 includes an acquisition unit 110, a detection unit 120, a calculation unit 130, a prediction unit 140, and a control unit 150. The control device 100 may further include a learning unit 160. These functional configurations may be realized by the processor 101 reading a control program from the ROM 102 or the storage 104, expanding the program in the RAM 103, and executing the program.

[0040] The acquisition unit 110 acquires an image of the door 30 that is to be opened and closed by the end effector 56 provided on the manipulator 54. The acquisition unit 110 also acquires sensor data that detects the operation of the manipulator 54.

[0041] The detection unit 120 detects the handle position of the door 30 from the image.

[0042] The calculation unit 130 calculates the current position of the end effector 56 using kinematics.

[0043] The prediction unit 140 inputs the handle position and the current position into a prediction model and predicts a trajectory consisting of multiple waypoints that the end effector 56 will take to move from the current position to a target position where the open / closed state of the door 30 will be changed.

[0044] The control unit 150 controls the manipulator 54 based on the trajectory.

[0045] The learning unit 160 learns a prediction model by deep learning. Furthermore, the learning unit 160 additionally learns the prediction model learned by deep learning by reinforcement learning.

[0046] 4 is a diagram showing an example of a block diagram of the control device 100 according to this embodiment. The operation of each functional unit will be described in detail using this diagram.

[0047] The acquisition unit 110 may acquire a first RGB image of the door 30 captured by the first camera 58_1. Similarly, the acquisition unit 110 may acquire a second RGB image of the door 30 captured by the second camera 58_2. That is, the acquisition unit 110 may acquire a plurality of images of the door 30 captured at different positions. In this way, for example, the acquisition unit 110 can acquire a plurality of images from the stereo camera 58 provided on the robot 50 having the manipulator 54.

[0048] Furthermore, the acquiring unit 110 may acquire sensor data that detects the movement of the manipulator 54. More specifically, the acquiring unit 110 may acquire the values ​​of the joint angles of the manipulator 54 from an encoder that detects the movement of the manipulator 54.

[0049] The detection unit 120 may detect the steering wheel position using an object detection algorithm. In this case, from the viewpoint of detection speed and accuracy, YOLO (You Only Look Once) v7 may be adopted as the object detection algorithm. Here, FIG. 5 is referred to.

[0050] 5 is a diagram illustrating an example of detecting a steering wheel position using an object detection algorithm. The detection unit 120 may input an image acquired by the acquisition unit 110 at time step k (where k is an element of the set {1, 2, ..., K}) to YOLO, which is one of the object detection models. In response to this, YOLO may output data of a bounding box 40.

[0051] Here, the bounding box 40 refers to a partial region obtained by enclosing the object region with the smallest rectangle relative to the external region and separating it by a boundary. Note that while the diagram shows an example in which there is one bounding box 40, multiple bounding boxes 40 may be output. In this case, a reliability may be calculated for each bounding box 40. The reliability indicates the probability that an object actually exists within the bounding box 40; the closer the value is to 1, the higher the possibility that an object exists, and the closer the value is to 0, the lower the possibility that an object exists. If multiple bounding boxes 40 are output, the detection unit 120 may select the bounding box with the highest reliability.

[0052] Here, the bounding box 40 may be defined by two corner points of a rectangle. min ,y min ) and point (x max ,y max ) is used as an example. In this case, the detection unit 120 may define the bounding box 40 by, for example, ((x min +x max ) / 2,(y min +y max ) / 2), the center coordinate c(x center ,y center ) may be calculated.

[0053] The detection unit 120 may calculate the first center coordinate c1 and the second center coordinate c2 by performing, for example, such processing on the first RGB image and the second RGB image, respectively. In this case, the handle position h at time step k is k can be expressed by the following formula:

number

[0054] In the above description, the case where the center coordinate c of the bounding box 40 is detected as the handle position h has been described as an example, but the present invention is not limited to this. The detection unit 120 may detect the coordinates of a predetermined position in a partial region as the handle position h depending on the type of door 30 and the shape of the handle 35. Returning to the description of FIG. 4.

[0055] The calculation unit 130 may calculate the current position of the end effector 56 using kinematics. In this case, the calculation unit 130 may calculate the current position of the end effector 56 by calculating forward kinematics using the values ​​of the joint angles acquired by the acquisition unit 110 at time step k. More specifically, the calculation unit 130 may calculate the three-dimensional coordinates (p x ,p y ,p z ) may be calculated. In this case, the current position p of the end effector 56 at time step k is k can be expressed by the following formula:

number

[0056] The prediction unit 140 predicts the steering wheel position h k and current position p k is input to a model prediction network (hereinafter also referred to as a "prediction model") to determine whether the end effector 56 is at the current position p kA trajectory consisting of multiple waypoints from the target position where the door 30 changes its open / closed state may be predicted.

[0057] Such a predictive model may be trained in advance in a learning phase. In this case, the processor 101 may train the predictive model in two phases: a phase in which the predictive model is trained in advance by deep learning, and a phase in which the predictive model trained in advance by deep learning is additionally trained by reinforcement learning. This type of learning method is referred to as MPRL (Model Predictive Reinforcement Learning).

[0058] The purpose of the pre-learning is to roughly predict the trajectory of the end effector 56 along a geometrically calculated trajectory. In the pre-learning, the processor 101 may perform supervised learning of the network using a geometrically prepared trajectory as training data.

[0059] Here, the handle position h at time step k k and the current position p of the end effector 56 k are expressed as (Equation 1) and (Equation 2), respectively, the input x k can be defined as follows:

number

[0060] In the above explanation, the handle position h k and current position p k Although the case where only the above information is input to the prediction model is shown as an example, the present invention is not limited to this. In addition to the above information, information regarding the width and opening / closing angle of the door 30 may also be input to the prediction model. This information can be obtained, for example, by analyzing images captured by the stereo camera 58.

[0061] Then, if the trajectory predicted by the prediction model is represented by n waypoints, the output of the prediction model is y k can be defined as follows:

number

[0062] The purpose of the additional learning is to correct the trajectory of the end effector 56. For example, the additional learning may use PPO (Proximal Policy Optimization), which is a reinforcement learning scheme. In this case, the processor 101 may set a reward according to at least one of the manipulability and the joint torque of the manipulator 54.

[0063] More specifically, state s at time step k k may be defined as follows:

number

[0064] Also, action a at time step k k may be defined as the amount of change in the position of the end effector 56 by the following equation: where p1 indicates the waypoint at the next time.

number

[0065] And the reward r at time step k k may be defined by the following equation using manipulability o, joint angular velocity q, and joint torque τ:

number

[0066] where d represents the Euclidean distance between the end effector 56 and the handle 35, and θ represents the opening / closing angle of the door 30. Also, α, β, γ, δ, and ε are hyperparameters for adjusting the priority of each term.

[0067] The processor 101 may use the prediction model learned in this way to predict a trajectory consisting of a plurality of waypoints. Then, the waypoint p1 at the next time point may be supplied to the control unit 150.

[0068] The control unit 150 controls the manipulator 54 based on the trajectory predicted by the prediction unit 140. More specifically, the control unit 150 may calculate the values ​​of the joint angles of the manipulator 54 by calculating inverse kinematics using the coordinates of the waypoint p1 at the next time as target values. Then, the control unit 150 may control the manipulator 54 based on the calculated values ​​of the joint angles.

[0069] The control device 100 performs each task by repeatedly acquiring images, detecting the handle position, calculating the current position, predicting the trajectory, and controlling the manipulator 54. Such tasks include a reaching task and an opening task in the opening operation, and a reaching task and a closing task in the closing operation.

[0070] Here, the reaching task is a task from moving the end effector 56 from an initial position away from the handle 35 to hooking it onto the handle 35. The opening task is a task from hooking the end effector 56 onto the handle 35 to opening the door 30. Similarly, the closing task is a task from hooking the end effector 56 onto the handle 35 to closing the door 30.

[0071] 6 is a diagram showing an example of a control flow executed by the control device 100 according to this embodiment. This flow may be started by the processor 101 reading a control program from the ROM 102 or the storage 104, expanding it in the RAM 103, and executing it.

[0072] In step S210, the processor 101, as the acquisition unit 110, acquires an image of the door 30 that is to be opened and closed by the end effector 56 provided on the manipulator 54. At this time, the processor 101 may acquire a plurality of images that are each captured at different positions of the door 30, as described above. More specifically, the processor 101 may acquire a plurality of images (e.g., a first RGB image and a second RGB image) from the stereo camera 58 provided on the robot 50 having the manipulator 54.

[0073] In step S215, the processor 101, functioning as the acquisition unit 110, acquires sensor data that detects the movement of the manipulator 54. At this time, the processor 101 may acquire the values ​​of the joint angles of the manipulator 54 from the encoder that detects the movement of the manipulator 54, as described above.

[0074] In step S220, the processor 101, functioning as the detection unit 120, determines the handle position h of the door 30 from the image acquired in step S210. k In this case, the processor 101 detects the steering wheel position h from the image using an object detection algorithm (for example, YOLO) as described above. k More specifically, the processor 101 may calculate the coordinates of the handle position (for example, the center coordinate c) from the coordinates of the partial region (for example, a bounding box) extracted by the object detection algorithm.

[0075] In step S230, the processor 101, functioning as the calculation unit 130, calculates the current position p of the end effector 56 using kinematics. k In this case, the processor 101 may calculate the forward kinematics using the values ​​of the joint angles, as described above.

[0076] In step S240, the processor 101, functioning as the prediction unit 140, calculates the steering wheel position h detected in step S220. k , and the current position p calculated in step S230 k is input into the prediction model to predict a trajectory consisting of a plurality of waypoints until the end effector 56 moves from the current position to a target position where the open / closed state of the door 30 is changed. In this case, the processor 101 may learn the prediction model by deep learning as described above. In this case, the processor 101 may perform supervised learning using a geometrically prepared trajectory in advance as training data. Furthermore, the processor 101 may additionally learn the prediction model learned by deep learning by reinforcement learning. In this case, the processor 101 calculates a reward r according to at least one of the manipulability о and the joint torque τ of the manipulator. k may be set.

[0077] In step S250, the processor 101, as the control unit 150, controls the manipulator 54 based on the trajectory predicted in step S240. At this time, as described above, the processor 101 may calculate the values ​​of the joint angles of the manipulator 54 by calculating inverse kinematics using the coordinates of the way point p1 at the next time as target values. Then, the processor 101 may control the manipulator 54 based on the calculated values ​​of the joint angles.

[0078] In step S260, the processor 101 may determine whether there has been a change in the open / closed state of the door 30. If it is determined that the open / closed state has changed (YES), the processor 101 ends the process.

[0079] On the other hand, if it is determined that the open / closed state has not changed (No), the processor 101 returns the process to steps S210 and S215 and continues the flow. That is, the processor 101 repeatedly acquires an image, detects the handle position, calculates the current position, predicts the trajectory, and controls the manipulator 54 until the open / closed state of the door 30 changes.

[0080] In this embodiment, verification tests were conducted on two tasks, a reaching task and an opening task, using an opening operation as an example. In the verification tests, the target was opening and closing a folding door installed in a closet. Furthermore, for the two tasks, a prediction model was pre-trained under two conditions: when the position of the closet was fixed and when it was randomly placed, and then a verification test of trajectory prediction by the prediction model was conducted.

[0081] In the randomly placed closet condition, the coordinates of the closet handle position were randomly determined within a predetermined interval. In the reaching task, the coordinates of the initial position of the end effector 56 were randomly determined within a predetermined interval. The simulation environment was created using Isaac Gym, a physics simulator for reinforcement learning.

[0082] To evaluate the prediction accuracy for each task, we calculated the accuracy rate by defining a correct answer as one where the Euclidean distance between the geometrically calculated trajectory and the predicted waypoint was within 1 cm. The accuracy rates for the reaching and opening tasks when the closet location was fixed were 0.87 and 1.00, respectively. On the other hand, the accuracy rates for the reaching and opening tasks when the closet location was random were 0.75 and 0.87, respectively.

[0083] Next, we show the results of using a predictive model that was pre-trained under the condition that the position of the closet is fixed, and moving the end effector 56 along a trajectory consisting of 15 waypoints predicted by the predictive model when it is in its initial position.

[0084] Fig. 7 is a diagram showing an example of a simulation result in a reaching task. Fig. 8 is a diagram showing an example of a simulation result in an opening task. Here, the time when the end effector 56 is at the initial position is defined as t0, and the time when the end effector 56 reaches the n-th waypoint is defined as tn It states that:

[0085] This confirmed that it is possible to predict with high accuracy the reaching and opening tasks in opening operations.

[0086] Although the explanation so far has focused on the opening operation, it is also possible to predict with high accuracy the reaching and closing tasks in the closing operation.

[0087] As described above, in this embodiment, the computer detects the handle position of the door 30 from an image of the door 30, calculates the current position of the end effector 56, and inputs the handle position and the current position into a prediction model to repeatedly execute a series of processes to control the manipulator 54 based on a trajectory consisting of multiple waypoints predicted.

[0088] As a result, according to this embodiment, compared to controlling the manipulator 54 based only on the result of a one-time prediction, the manipulator 54 is controlled based on the result of prediction each time taking into account changes in the environment, so the probability that the robot 50 can open and close the door 30 can be increased even when the door being targeted has complex changes in the direction and magnitude of the force that must be applied to the handle 35 to open and close the door.

[0089] Various studies have been conducted on the opening operation of sliding doors and swing doors. However, no research has been conducted on folding doors. This is thought to be because the trajectory of the handle 35 is a roughly elliptical arc, and the direction and magnitude of the force to be applied to the handle 35 cannot be expressed by a simple function, making it very difficult. However, in this embodiment, not only sliding doors and swing doors but also folding doors can be opened and closed. As a result, according to this embodiment, even if the door 30 separating the space is a folding door, the robot 50 can autonomously open and close the door 30, thereby expanding the environments in which the robot 50 can provide services.

[0090] In this embodiment, the computer pre-learns the prediction model by deep learning and then performs additional learning by reinforcement learning. As a result, according to this embodiment, learning is performed in two stages, pre-learning and additional learning, which can speed up the convergence of rewards in reinforcement learning.

[0091] In this case, when performing pre-learning using deep learning, the computer performs supervised learning using a geometrically pre-created trajectory as training data. As a result, according to this embodiment, it is possible to roughly predict the trajectory of the end effector 56 along a geometrically calculated trajectory using a prediction model.

[0092] Furthermore, when performing additional learning through reinforcement learning, the computer sets a reward according to at least one of the manipulator manipulator 54's manipulability and joint torque. As a result, according to this embodiment, the prediction model can predict a trajectory taking into consideration postures that the manipulator 54 can actually take. Therefore, according to this embodiment, it is possible to reduce, for example, the probability that the manipulator 54 will assume a singular posture or that the manipulator 54 will be subjected to a high load.

[0093] In this embodiment, the computer uses an object detection algorithm to detect the steering wheel position. In this case, the computer detects the coordinates of the steering wheel position from the coordinates of the partial area extracted by the object detection algorithm. As a result, this embodiment allows the steering wheel position to be detected quickly and accurately, and coordinate calculations to be performed using simple operations.

[0094] Furthermore, in this embodiment, the computer acquires multiple images of the door 30 taken from different positions. At this time, the computer acquires the multiple images from the stereo camera 58 provided on the robot 50 having the manipulator 54. As a result, according to this embodiment, the space including the door 30 can be grasped three-dimensionally, including the depth direction. At this time, even if the robot 50 moves, images taken in an environment where the relative relationship of the photographing position with respect to the manipulator 54 is always constant are used, further increasing the probability of opening or closing the door 30.

[0095] The above-described processing can also be realized by dedicated hardware circuits. In this case, the processing may be performed by a single piece of hardware or by multiple pieces of hardware.

[0096] In addition, in the above description, processor 101 refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).

[0097] Furthermore, the operations of processor 101 in the above description may not only be performed by a single processor, but may also be performed by multiple processors located at physically separate locations working together. Furthermore, the order of the operations of processor 101 is not limited to the above-described order, and may be changed as appropriate.

[0098] The above-mentioned program may be provided by a computer-readable non-transitory recording medium such as a USB (Universal Serial Bus) memory, a flexible disk, or a CD-ROM (Compact Disc Read Only Memory), or may be provided online via a network such as the Internet. In this case, the program recorded on the computer-readable non-transitory recording medium is typically transferred and stored in a memory or storage device. The program may be provided as standalone application software, or may be incorporated into the software of each device as a function of the device.

[0099] The above-described program can be provided as a program product. The program product includes any product for providing the program. For example, the program product includes a program provided over a network such as the Internet, and a non-transitory computer-readable recording medium such as a CD-ROM or DVD on which the program is stored.

[0100] The present disclosure is not limited to the above, and it goes without saying that various modifications can be made without departing from the spirit of the present disclosure. [Explanation of symbols]

[0101] 10 Door Opening System 30 doors 35 Handle 50 Robot 52 Automated Guided Vehicle 54 Manipulator 56 End Effector 58 Stereo Camera 100 control device 101 processors 102 ROM 103 RAM 104 Storage 105 Communication Interface 106 User Interface 109 Bus 110 Acquisition Department 120 Detector 130 Calculation Unit 140 Prediction Department 150 control section 160 Learning Department

Claims

1. The computer Acquiring an image of a door to be opened or closed by an end effector provided on a manipulator; Detecting the door handle position from the image; calculating a current position of the end effector using kinematics; inputting the handle position and the current position into a prediction model to predict a trajectory of the end effector, which includes a plurality of waypoints, from the current position to a target position at which the door is opened or closed; controlling the manipulator based on the trajectory; repeatedly performing the acquisition of the image, the detection of the handle position, the calculation of the current position, the prediction of the trajectory, and the control of the manipulator; Control method.

2. The computer further comprising training the predictive model by deep learning. The control method according to claim 1 .

3. The learning by deep learning includes supervised learning using the trajectory geometrically created in advance as training data. The control method according to claim 2 .

4. The computer Further comprising additionally learning the prediction model learned by the deep learning through reinforcement learning. The control method according to claim 2 .

5. the additional learning by reinforcement learning includes setting a reward according to at least one of manipulability and joint torque of the manipulator. The control method according to claim 4.

6. Detecting the steering wheel position includes detecting the steering wheel position from the image using an object detection algorithm. A control method according to any one of claims 1 to 5.

7. Detecting the steering wheel position includes calculating coordinates of the steering wheel position from coordinates of the partial region extracted by the object detection algorithm. The control method according to claim 6.

8. acquiring the images includes acquiring a plurality of images of the door taken at different positions; A control method according to any one of claims 1 to 5.

9. acquiring the images includes acquiring the plurality of images from a stereo camera provided on a robot having the manipulator. The control method according to claim 8.

10. The door is a folding door. A control method according to any one of claims 1 to 5.

11. a processor, the processor comprising: An image of a door to be opened or closed is acquired by an end effector provided on a manipulator; Detecting the door handle position from the image; Calculating a current position of the end effector using kinematics; inputting the handle position and the current position into a prediction model to predict a trajectory of the end effector, which is composed of a plurality of via points, from the current position to a target position at which the door is changed in open or closed state; controlling the manipulator based on the trajectory; repeatedly performing the acquisition of the image, the detection of the handle position, the calculation of the current position, the prediction of the trajectory, and the control of the manipulator; Control device.

12. On the computer, An image of a door to be opened or closed is acquired by an end effector provided on the manipulator; Detecting the door handle position from the image; calculating a current position of the end effector using kinematics; inputting the handle position and the current position into a prediction model to predict a trajectory of the end effector, which is composed of a plurality of via points, from the current position to a target position at which the door is changed in open / closed state; Controlling the manipulator based on the trajectory; repeatedly executing the acquisition of the image, the detection of the handle position, the calculation of the current position, the prediction of the trajectory, and the control of the manipulator; Control program.

Citation Information

Patent Citations

  • Information processor, robot, and program

    JP2015139872A

  • Remote operation system and remote operation method

    JP2021094605A

  • Method and computing system for performing motion planning based on image information generated by a camera

    JP2022044830A

  • Robot

    JP2022050147A