Robot arm control device, control method, and control program

A robot arm control system generates composite depth images from door shapes for reinforcement learning, enabling a single algorithm to open multiple swing doors, improving indoor robot mobility.

JP2025179654APending Publication Date: 2025-12-10TOKYO UNIVERSITY OF SCIENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024086550
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Existing methods for opening swing doors using robot manipulators require multiple algorithms for different types of swing doors, lacking a unified solution for various door types.

Method used

A robot arm control system that utilizes depth images from a camera to generate composite depth images emphasizing uneven door shapes, combined with reinforcement learning to create a single algorithm capable of opening multiple types of swing doors.

Benefits of technology

Enables seamless door opening operations for various swing doors using a single algorithm, enhancing the versatility of service robots in indoor environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179654000001_ABST
    Figure 2025179654000001_ABST
Patent Text Reader

Abstract

To open multiple kinds of swing doors by a single algorithm.SOLUTION: A control device 10 comprises: an acquisition part 11A that acquires a depth image, which photographs a lever type door handle and the periphery of the lever type door handle, from a camera on a robot arm; a generation part 11B that generates a first depth image which represents a distance between the lever type door handle and an end effector on the robot arm, and a second depth image, in which an uneven shape of the lever type door handle is emphasized, from the depth image, and generated a synthesis depth image from the first depth image and the second depth image; and a learning part 11C that performs reinforcement learning by using the synthesis depth image and positional information of the end effector so that a predetermined reward increases at each predetermined time step, thereby inputting the synthesis depth image and the positional information of the end effector and outputting control data of the robot arm corresponding to an opening operation of a swing door.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a control device, a control method, and a control program for a robot arm. [Background technology]

[0002] In recent years, labor shortages have accelerated the demand for service robots as a substitute for labor. Mobile manipulators, in particular, are becoming increasingly popular in indoor environments because they can perform a variety of tasks. In indoor environments, doors often prevent seamless movement of service robots, so door opening is required.

[0003] From the viewpoint of cost and versatility, door-opening functions are being added to service robots. There are many different types of doors, and especially in the case of swing doors, there are a wide variety of types, such as right-opening, left-opening, inward-opening, and outward-opening.

[0004] Taking into account the diversity of doors, various methods have been proposed for the purpose of opening a swing door using a manipulator, including reinforcement learning (see, for example, Non-Patent Document 1), teaching-based deep predictive learning (see, for example, Non-Patent Document 2), and position-force control (see, for example, Non-Patent Document 3). [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Urakami, Y., Hodgkinson, A., Carlin, C., Leu, R., Rigazio,L. and Abbeel, P., “DoorGym: A Scalable Door Opening Environment And Baseline Agent,” arXiv, arXiv:1908.01887v4, 2022. [Non-patent document 2] Ito, H., Yamamoto, K., Mori, H. and Ogata, T., “Efficient predictive multitask learning with an embodied model for door opening and entry with whole-body control,” Science Robotics, vol.7, no.65, eaax8177, 2022. [Non-patent document 3] Kang, G., Seong, H., Lee, D. and Shim, D., H, “A Versatile Door Opening System with Mobile Manipulator through Adaptive Position-Force Control and Reinforcement Learning,”arXiv, arXiv:2307.04422, 2023. Summary of the Invention [Problem to be solved by the invention]

[0006] However, the methods described in the above Non-Patent Documents 1 to 3 cannot handle a wide variety of swing door opening operations with a single algorithm, so in order to open multiple types of swing doors, multiple algorithms must be prepared for each type of swing door.

[0007] In view of the above-mentioned problems, the present disclosure aims to provide a control device, a control method, and a control program for a robot arm that can open multiple types of swing doors using a single algorithm. [Means for solving the problem]

[0008] In order to achieve the above object, a robot arm control device according to one embodiment of the present disclosure is a control device for a robot arm that performs an opening operation of a swing door having a lever-type door handle, and includes: an acquisition unit that acquires depth images of the lever-type door handle and the area around the lever-type door handle from a camera carried by the robot arm; a generation unit that generates from the depth images a first depth image representing the distance between the lever-type door handle and an end effector carried by the robot arm and a second depth image in which the uneven shape of the lever-type door handle is emphasized, and generates a composite depth image by combining the first depth image and the second depth image; and a learning unit that uses the composite depth image and position information of the end effector to perform reinforcement learning so as to increase a predetermined reward for each predetermined time step, thereby generating a trained model that takes the composite depth image and position information of the end effector as input and outputs control data for the robot arm corresponding to the opening operation of the swing door.

[0009] The generation unit may perform threshold processing on each pixel value of the depth image using a threshold range based on the working distance of the camera, replace pixel values ​​outside the threshold range with a predetermined value that does not belong to the threshold range, normalize the replaced depth image obtained by the replacement within the threshold range to generate the first depth image, and normalize the first depth image between a minimum and maximum pixel value to generate the second depth image.

[0010] The synthetic depth image may be an image obtained by averaging the first depth image and the second depth image.

[0011] The predetermined reward may be represented by the distance between the lever-type door handle and the end effector, the rotation angle of the lever-type door handle, the rotation angle of the door hinge of the swing door, and the norm of the output of the trained model.

[0012] The robot may further include an operation control unit that controls the robot arm in accordance with the control data output from the trained model.

[0013] The trained model may be a model including a convolutional neural network and a multi-layer perceptron.

[0014] A control method for a robot arm according to one aspect of the present disclosure is a control method for a robot arm that performs an opening operation of a swing door having a lever-type door handle, in which a computer executes a process of acquiring depth images of the lever-type door handle and the area around the lever-type door handle from a camera carried by the robot arm, generating from the depth images a first depth image representing the distance between the lever-type door handle and an end effector carried by the robot arm and a second depth image in which the uneven shape of the lever-type door handle is emphasized, generating a composite depth image by combining the first depth image and the second depth image, and using the composite depth image and position information of the end effector, performing reinforcement learning so as to increase a predetermined reward for each predetermined time step, thereby generating a trained model that takes the composite depth image and the position information of the end effector as input and outputs control data for the robot arm corresponding to the opening operation of the swing door.

[0015] A control program for a robot arm according to one embodiment of the present disclosure is a control program for a robot arm that performs an opening operation on a swing door having a lever-type door handle, and causes a computer to execute the following process: acquire depth images of the lever-type door handle and the area around the lever-type door handle from a camera possessed by the robot arm; generate from the depth images a first depth image representing the distance between the lever-type door handle and an end effector possessed by the robot arm and a second depth image that emphasizes the uneven shape of the lever-type door handle; generate a composite depth image by combining the first depth image and the second depth image; and use the composite depth image and position information of the end effector to perform reinforcement learning so as to increase a predetermined reward for each predetermined time step, thereby generating a trained model that takes the composite depth image and the position information of the end effector as input and outputs control data for the robot arm corresponding to the opening operation of the swing door. [Effects of the Invention]

[0016] According to the robot arm control device, control method, and control program of the present disclosure, multiple types of swing doors can be opened using a single algorithm. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a diagram illustrating an example of the appearance of a robot arm according to an embodiment. [Figure 2] FIG. 10 is a diagram showing an enlarged view of the end effector and depth camera of the robot arm. [Figure 3] FIG. 2 is a block diagram showing an example of the electrical configuration of a control device for a robot arm. [Figure 4] 1A and 1B are diagrams showing an example of a plurality of types of swing doors with different opening directions. [Figure 5] FIG. 2 is a block diagram illustrating an example of a functional configuration of a control device according to the embodiment. [Figure 6] 1A and 1B are diagrams illustrating an example of a depth image before pre-processing and a depth image after pre-processing. [Figure 7] FIG. 10 is a diagram illustrating states, actions, and rewards related to a swing door opening operation. [Figure 8] (A) and (B) are diagrams explaining the states, actions, and rewards related to the swing door opening operation. (C) and (D) are diagrams explaining the numerical simulation of the swing door opening operation. [Figure 9] 10A and 10B are diagrams illustrating a numerical simulation of a swing door opening operation. [Figure 10] 10A and 10B are diagrams illustrating a numerical simulation of a swing door opening operation. [Figure 11] 10 is a flowchart illustrating an example of a processing flow according to a control program according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, each embodiment for carrying out the present disclosure will be described with reference to the drawings. Note that the scope necessary for the explanation to achieve the object of the present disclosure will be schematically shown below, and the scope necessary for explaining the relevant parts of the present disclosure will be mainly explained, and the parts for which explanation is omitted will be considered to be publicly known technologies.

[0019] Fig. 1 is a diagram showing an example of the appearance of a robot arm 100 according to this embodiment. Fig. 2 is a diagram showing an enlarged view of an end effector 40 and a depth camera 50 provided to the robot arm 100.

[0020] As shown in FIG. 1, the robot arm 100 according to this embodiment includes a control device 10, a support base 20, a manipulator 30, an end effector 40, a depth camera 50, and an encoder 55.

[0021] A plurality of wheels 21 are provided on the bottom of the support base 20, and the wheels 21 enable the support base 20 to move on the floor surface. A manipulator 30 is placed on the top of the support base 20, and a control device 10 is provided inside the manipulator 30. The control device 10 may also be provided outside the support base 20.

[0022] The manipulator 30 includes multiple joints and moves three-dimensionally under the control of the control device 10. Here, the direction perpendicular to the wall surface 70 on which the swing door 60 is provided is defined as the X-axis direction, the direction along the horizontal direction (left-right direction) of the wall surface 70 is defined as the Y-axis direction, and the direction along the vertical direction (up-down direction) of the wall surface 70 is defined as the Z-axis direction.

[0023] The manipulator 30 has a built-in encoder 55. The encoder 55 is a device that detects the amount of mechanical movement, direction, joint angle, etc. of the manipulator 30 using a sensor, and outputs the detected information as an electrical signal (encoder data).

[0024] An end effector 40 is detachably attached to the tip of the manipulator 30. As shown in Figs. 1 and 2, the end effector 40 is, for example, a hand member that can open a swing door 60 by rotating a lever-type door handle 61 attached to the swing door 60. It is desirable to prepare a plurality of different types of end effectors 40. This allows the end effector 40 to be replaced and used according to the shape of the lever-type door handle 61, etc.

[0025] The end effector 40 is provided with a depth camera 50. The depth camera 50 is an example of a camera that captures depth images. The depth camera 50 captures the lever-type door handle 61 and the area around the lever-type door handle 61, and visualizes distance information to the lever-type door handle 61 as a depth image.

[0026] FIG. 3 is a block diagram showing an example of the electrical configuration of the control device 10 of the robot arm 100.

[0027] As shown in FIG. 3, the control device 10 according to this embodiment includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, an input / output interface (I / O) 14, a memory unit 15, and a connection unit 16.

[0028] The CPU 11, ROM 12, RAM 13, and I / O 14 are connected to each other via a bus. The I / O 14 is connected to various functional units including a storage unit 15 and a connection unit 16. These functional units can communicate with the CPU 11 via the I / O 14.

[0029] The control unit is configured with the CPU 11, ROM 12, RAM 13, and I / O 14. The control unit may be configured as a sub-control unit that controls part of the operation of the control device 10, or may be configured as part of the main control unit that controls the overall operation of the control device 10. For example, an integrated circuit such as an LSI (Large Scale Integration) or an IC (Integrated Circuit) chip set is used for part or all of the blocks of the control unit. Individual circuits may be used for each of the above blocks, or a circuit in which part or all of the blocks are integrated may be used. The above blocks may be provided integrally, or some of the blocks may be provided separately. Furthermore, part of each of the above blocks may be provided separately. The integration of the control unit is not limited to LSI, and a dedicated circuit or a general-purpose processor may also be used.

[0030] For example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, or the like is used as the storage unit 15. A control program 15A according to this embodiment is stored in the storage unit 15. Note that this control program 15A may be stored in the ROM 12.

[0031] The control program 15A may be pre-installed in the control device 10, for example. The control program 15A may be realized by storing it in a non-volatile storage medium or distributing it via a network and installing it appropriately in the control device 10. Examples of non-volatile storage media include a CD-ROM (Compact Disc Read Only Memory), a magneto-optical disk, a HDD, a DVD-ROM (Digital Versatile Disc Read Only Memory), a flash memory, a memory card, etc.

[0032] The connection unit 16 is connected to each of the manipulator 30, the depth camera 50, and the encoder 55, and connects each of these manipulator 30, the depth camera 50, and the encoder 55 to the CPU 11 so that they can communicate with each other.

[0033] 4 is a diagram showing examples of multiple types of swing doors 60A to 60D that open in different directions. The swing door 60A is a right-opening / inward-opening type, the swing door 60B is a right-opening / outward-opening type, the swing door 60C is a left-opening / inward-opening type, and the swing door 60D is a left-opening / outward-opening type.

[0034] The control device 10 according to this embodiment controls the robot arm 100 to perform an opening operation on multiple types of swing doors 60A to 60D using a single algorithm. Hereinafter, in this embodiment, the multiple types of swing doors 60A to 60D are referred to as targets of the opening operation, but the targets of the opening operation are not limited to these swing doors 60A to 60D. Note that when it is not necessary to distinguish between the multiple types of swing doors 60A to 60D in the description, they will be collectively referred to as the swing door 60.

[0035] Specifically, the CPU 11 of the control device 10 writes a control program 15A stored in the storage unit 15 into the RAM 13 and executes the program, thereby functioning as each unit shown in FIG.

[0036] FIG. 5 is a block diagram showing an example of the functional configuration of the control device 10 according to this embodiment.

[0037] As shown in FIG. 5, the CPU 11 of the control device 10 functions as an acquisition unit 11A, a generation unit 11B, a learning unit 11C, and an operation control unit 11D.

[0038] The acquisition unit 11A acquires a depth image of the lever-type door handle 61 and the area around the lever-type door handle 61 from the depth camera 50 of the robot arm 100. The acquisition unit 11A also acquires encoder data related to the manipulator 30 from the encoder 55 of the robot arm 100.

[0039] The generating unit 11B generates a first depth image and a second depth image from the depth images acquired by the acquiring unit 11A, and generates a composite depth image by combining the first depth image and the second depth image. The first depth image is a depth image that represents the distance between the lever-type door handle 61 and the end effector 40 of the robot arm 100. On the other hand, the second depth image is a depth image in which the uneven shape of the lever-type door handle 61 is emphasized.

[0040] Specifically, the generation unit 11B performs threshold processing on each pixel value of the depth image using a threshold range based on the working distance of the depth camera 50, and replaces pixel values ​​outside the threshold range with a predetermined value that does not belong to the threshold range. Then, the generation unit 11B normalizes the replaced depth image obtained by the replacement within the threshold range to generate a first depth image, and normalizes the first depth image between the minimum and maximum pixel values ​​to generate a second depth image.

[0041] Here, a method for generating a composite depth image from the first depth image and the second depth image will be specifically described. Hereinafter, R is a real number. The depth camera 50 that acquires the depth image has a working distance [d min ,d max ](d min ,d max ∈R). The working distance is the distance from the tip of the lens to the subject.

[0042] First, a depth image D∈R with height h and width w acquired from the depth camera 50 is h×w Each pixel value D i,j For (i=1, 2, , h, j=1, 2, , w), the threshold range [d min ,d max ] is used for thresholding, and the threshold range [d min ,d max ] are set as pixel values ​​outside the threshold range [d min ,d max ] a given value d that does not belong to out ∈R to obtain a depth image D' after substitution. Each pixel value D' of the depth image D' after substitution i,j is expressed as follows. Here, as an example, the threshold range is set to the working distance [d min ,d max ], but the working distance [d min ,d max An appropriate threshold range may be set within the range of .

[0043] JPEG2025179654000002.jpg1794(1)

[0044] Next, the replaced depth image D' is calculated using the threshold range [d min ,d max ] and normalize it to the first depth image D - 1 (where - is directly above D, and so on). The first depth image D - Each pixel value D - 1i,j is expressed as follows:

[0045] JPEG2025179654000003.jpg1753(2)

[0046] First depth image D - Each pixel value D - 1i,j Define a function Normalize to normalize between two values, and use it to find the first depth image D- 1 is expressed as follows:

[0047] JPEG2025179654000004.jpg984(3)

[0048] Furthermore, the first depth image D - 1 to the pixel value D - 1i,j The minimum value of d - 1min and the maximum value d - 1max The second depth image D is obtained by normalizing the image between the two and emphasizing the uneven shape. - 2. The second depth image D - 2 is expressed as follows:

[0049] JPEG2025179654000005.jpg1090(4)

[0050] Finally, the first depth image D - 1 and the second depth image D - 2 and 3 are combined to create a composite depth image D - n2 The synthetic depth image D - n2 is, for example, the first depth image D - 1 and the second depth image D - 2 and 3 are averaged images.

[0051] 6 is a diagram illustrating an example of a depth image before pre-processing and a depth image after pre-processing. The depth image before pre-processing is represented as a depth image D, and the depth image after pre-processing is represented as a composite depth image D. - n2 It is expressed as:

[0052] As shown in Figure 6, the pre-processed synthetic depth image D - n2 It can be seen that the concave and convex shape of the lever-type door handle 61 is emphasized compared to the depth image D before pre-processing. - n2According to this, the uneven shape of the lever-type door handle 61 is emphasized, making it easy to recognize whether the swing door 60 opens to the left or right.

[0053] Furthermore, the generation unit 11B calculates the position information of the end effector 40 by forward kinematics from the encoder data acquired by the acquisition unit 11A.

[0054] The learning unit 11C generates a synthetic depth image D - n2 and the position information of the end effector 40, reinforcement learning is performed so that a predetermined reward increases at each predetermined time step, thereby generating a synthetic depth image D - n2 and position information of the end effector are input, and a trained model 151 is generated that outputs control data for the robot arm 100 corresponding to the opening operation of the swing door 60. The trained model 151 is stored in the storage unit 15, for example.

[0055] The trained model 151 is a model including, for example, a Convolutional Neural Network (hereinafter referred to as "CNN") and a Multilayer Perceptron (hereinafter referred to as "MLP").

[0056] In addition, the predetermined reward is represented, for example, by the distance between the lever-type door handle 61 and the end effector 40, the rotation angle of the lever-type door handle 61, the rotation angle of the door hinge of the swing door 60, and the norm of the output of the trained model 151.

[0057] The operation control unit 11D controls the robot arm 100 (that is, the manipulator 30) in accordance with the control data output from the trained model 151.

[0058] Next, with reference to FIGS. 7, 8(A) to 8(D), 9, and 10, a swing door opening operation using the robot arm 100 according to this embodiment will be specifically described.

[0059] State, Behavior, and Reward 7, 8(A), and 8(B) are diagrams for explaining the states, actions, and rewards related to the swing door opening operation.

[0060] Here, it is assumed that in the initial time step, the lever-type door handle 61 can be measured using a sensor (not shown) and the manipulator 30 is installed in a position where the manipulator 30 can open the swing door 60. Note that the movement of the robot arm 100 is not taken into consideration, and the manipulator 30 is fixed in the same position in order to focus on the swing door opening operation by the manipulator 30.

[0061] First, the state is explained. The state S input to the MLP 113 at a predetermined time step k∈{1, 2, . . . , K} is k is defined as follows: The CNN 112 and the MLP 113 form a policy network.

[0062] JPEG2025179654000006.jpg1282(5)

[0063] where f k ∈R n is a composite depth image D obtained by performing depth image processing 110 on a depth image D acquired by the depth camera 50. - n2 is input to the CNN 112, and represents the n-dimensional feature quantity obtained as the output of the CNN 112. As described above, the depth image processing 110 is a process for emphasizing the uneven shape of the lever-type door handle 61. Note that T represents a transposed matrix. Also, Δp k (:=p k -p k-1 ) is the joint angle θ of the manipulator 30 acquired as encoder data from the encoder 55. k The position and orientation p of the end effector 40 calculated by forward kinematics 111 from k ∈R6 and the position and orientation of the previous time step p k-1 It represents the difference between the two. k is an example of the position information of the end effector 40.

[0064] Next, we will explain the action. The action a k is defined as follows: Action a k is an example of control data for the manipulator 30.

[0065] JPEG2025179654000007.jpg1260(6)

[0066] where p * k+1 (:=p k +Δp * k+1 ) is the target position and posture of the end effector 40 at the next time step k+1. k From inverse kinematics (=f -1 ) 116, the target joint angle θ of the manipulator 30 at time step k+1 * k+1 is calculated as follows and input to the manipulator 30:

[0067] JPEG2025179654000008.jpg9100(7)

[0068] Here, c∈R represents the action scale, ΔT represents the sampling period of the depth camera 50, and DOF represents the degrees of freedom of the manipulator 30.

[0069] Next, the reward will be explained. The reward r obtained by performing the reward process 114 is k As shown in FIG. 8(A), the hinge rotation angle α k , the rotation angle β of the lever-type door handle 61 k As shown in FIG. 8B, the distance d between the center P1 of the lever-type door handle 61 and the tip P2 of the end effector 40 is k, and action a k Penalty for |a k Using |2, it is defined as follows: In this embodiment, the vector a∈R m Let us define the Euclidean norm for |a|2.

[0070] JPEG2025179654000009.jpg14127(8)

[0071] where w i (i∈{1, ,4}) represents the weight for each reward. Also, the penalty |a k |2 is an example of the norm of the output of the MLP 113.

[0072] The policy optimizer 115 then calculates the reward r k Enter the reward k The CNN 112 and the MLP 113 are updated so that

[0073] <Verification by numerical simulation> 8(C), 8(D), 9, and 10 are diagrams for explaining a numerical simulation of a swing door opening operation.

[0074] To verify this numerical simulation, we used the simulation environment based on "Isaac Gym" described in Reference 1 and the reinforcement learning library "skrl" described in Reference 2.

[0075] (Reference 1) Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., and State, G., “Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning,” arXiv,arXiv:2108.10470v2, 2021.

[0076] (Reference 2) Serrano-Munoz, A., Chrysostomou, D., Bogh, S. and Arana-Arexolaleiba, N., “skrl: Modular and Flexible Library for Reinforcement Learning,” Journal of Machine Learning Research, vol.24, no.254, pp.1-9, 2023.

[0077] In this numerical simulation, the movement of the end effector 40 was learned in 512 parallel environments, with one episode consisting of 300 time steps. An episode is the period from the start to the end of a task to be solved by reinforcement learning. In each episode, the initial position and orientation of the end effector 40 was set to the position and orientation shown in Figure 8(C), and a uniformly distributed random displacement given in the range [-0.25 [m], 0.25 [m]] was added to each of the x-axis, y-axis, and z-axis, and a uniformly distributed random angle given in the range [-0.125 [rad], 0.125 [rad]] was added around each axis.

[0078] As a result of learning only the movement of the end effector 40 through this numerical simulation, it was confirmed that opening operations could be performed for the four types of swing doors 60A to 60D shown in FIG. 8(D).

[0079] Figure 9 shows the frequency distribution of the maximum door hinge rotation angle (= maximum door rotation angle) for left-opening swing doors in each episode. In Figure 9, the vertical axis shows frequency (unit: %), and the horizontal axis shows the maximum door rotation angle (unit: degrees). According to the frequency distribution in Figure 9, out of a total of 2,500 episodes, the door rotation angle was 70 degrees or more in more than 90% of the episodes, confirming that the swing door opening operation was successful.

[0080] Figure 10 shows the swing door opening operation for a left-opening swing door. The example in Figure 10 shows the swing door opening operation performed by the manipulator 30 connected to the end effector 40 in 0.5-second intervals (A) to (H). As shown in Figure 10, it was confirmed that the swing door opening operation was successful when the end effector 40 was connected to the manipulator 30.

[0081] Next, the operation of the robot arm 100 according to this embodiment will be described with reference to FIG.

[0082] FIG. 11 is a flowchart showing an example of the flow of processing by the control program 15A according to this embodiment.

[0083] First, when the control device 10 receives an instruction to start the swing door opening operation, the CPU 11 reads and executes the control program 15A.

[0084] In step S101 of FIG. 11, the CPU 11 acquires a depth image D from the depth camera 50, as shown in FIG. 7 above, for example.

[0085] In step S102, the CPU 11 generates a composite depth image D in which the concave and convex shapes of the lever-type door handle 61 are emphasized from the depth image D, as shown in FIG. 7, for example. - n2 Generate.

[0086] In step S103, the CPU 11 generates a synthetic depth image D - n2 Enter.

[0087] On the other hand, in step S104, the CPU 11 receives the joint angle θ of the manipulator 30 as encoder data from the encoder 55, as shown in FIG. k Get.

[0088] In step S105, the CPU 11 calculates the joint angle θ of the manipulator 30 as shown in FIG. 7, for example. k The position and orientation p of the end effector 40 are calculated by forward kinematic calculation. k The generated position and orientation p of the end effector 40 is calculated by inverse kinematics calculation. k and the position and orientation of the previous time step p k-1 The difference Δp k Here, steps S101 to S103 and steps S104 and S105 are processed in parallel, but the timing of the processing does not necessarily have to coincide.

[0089] In step S106, the CPU 11 calculates the feature f k and the difference Δp k and are input into the MLP 113.

[0090] In step S107, the CPU 11 performs the action a output from the MLP 113 as shown in FIG. 7, for example. k from the target joint angle θ of the manipulator 30 * k+1 is calculated and input to the manipulator 30.

[0091] In step S108, the CPU 11 determines whether or not it is time for learning the CNN 112 and the MLP 113. If it is determined that it is time for learning the CNN 112 and the MLP 113 (if the determination is positive), the process proceeds to step S109, and if it is determined that it is not time for learning the CNN 112 and the MLP 113 (if the determination is negative), the process proceeds to step S111.

[0092] In step S109, the CPU 11 calculates the hinge rotation angle α of the swing door 60 as shown in FIG. 7, for example. k , the rotation angle β of the lever-type door handle 61 k , the distance d between the lever-type door handle 61 and the end effector 40 k , and action a k Penalty for |a k Using |2, the reward r k Calculate.

[0093] In step S110, the CPU 11 calculates the reward r k The CNN 112 and the MLP 113 are updated so that

[0094] In step S111, the CPU 11 determines whether the end timing has arrived, such as whether the swing door opening operation has been successful. If it is determined that the end timing has not arrived (in the case of a negative determination), the process returns to steps S101 and S104 and is repeated. If it is determined that the end timing has arrived (in the case of a positive determination), the series of processes by this control program 15A is terminated.

[0095] According to this embodiment, by carrying out reinforcement learning using a synthetic depth image in which the uneven shape of a lever-type door handle is emphasized, multiple types of swing doors can be opened using a single algorithm.

[0096] In addition, since it can accommodate multiple types of swing doors, it can enable the service robot to move seamlessly in indoor environments.

[0097] While one form of the robot arm 100 has been described above using the embodiment, the disclosed form of the robot arm 100 is merely an example, and the form of the robot arm 100 is not limited to the scope described in the embodiment. Various changes and improvements can be made to the embodiment without departing from the gist of the present disclosure, and forms incorporating such changes or improvements are also included in the technical scope of the disclosure.

[0098] In the above embodiment, as an example, a configuration has been described in which the control processing of the robot arm 100 is realized by software processing. However, the control processing of the robot arm 100 may be processed by hardware. In this case, the processing speed can be increased compared to when it is realized by software processing.

[0099] In the above embodiment, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU 11) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).

[0100] Furthermore, the operations of the processors in the above embodiments may not only be performed by a single processor, but may also be performed by multiple processors located at physically separate locations working together. Furthermore, the order of the operations of the processors is not limited to the order described in the above embodiments, and may be changed as appropriate.

[0101] In the above embodiment, an example has been described in which the control program 15A is stored in the storage unit 15 or the ROM 12. However, the storage destination of the control program 15A is not limited to the storage unit 15 or the ROM 12. The control program 15A of the present disclosure can also be provided in a form stored in a computer-readable storage medium.

[0102] For example, the control program 15A may be provided in a form stored on an optical disk such as a CD-ROM, a DVD-ROM, or a Blu-ray disc. The control program 15A may also be provided in a form stored on a portable semiconductor memory such as a USB (Universal Serial Bus) memory or a memory card. These CD-ROM, DVD-ROM, Blu-ray disc, USB, and memory card are examples of non-transitory storage media.

[0103] The following supplementary notes are provided regarding the above-described embodiments.

[0104] (Appendix 1) A control device for a robot arm that performs an opening operation of a swing door having a lever-type door handle, an acquisition unit that acquires a depth image of the lever-type door handle and the area around the lever-type door handle from a camera included in the robot arm; a generating unit that generates, from the depth image, a first depth image representing the distance between the lever-type door handle and the end effector of the robot arm, and a second depth image in which the uneven shape of the lever-type door handle is emphasized, and generates a composite depth image in which the first depth image and the second depth image are combined; a learning unit that uses the synthetic depth image and the position information of the end effector to perform reinforcement learning so that a predetermined reward increases at each predetermined time step, thereby generating a trained model that receives the synthetic depth image and the position information of the end effector as inputs and outputs control data for the robot arm corresponding to the opening operation of the swing door; A control device for a robot arm comprising: (Appendix 2) The generation unit performs threshold processing on each pixel value of the depth image using a threshold range based on the working distance of the camera, replaces pixel values ​​outside the threshold range with a predetermined value that does not belong to the threshold range, normalizes the replaced depth image obtained by the replacement within the threshold range to generate the first depth image, and normalizes the first depth image between a minimum value and a maximum value of pixel values ​​to generate the second depth image. 2. A control device for a robot arm according to claim 1. (Appendix 3) The synthetic depth image is an image obtained by averaging the first depth image and the second depth image. 10. A control device for a robot arm according to claim 1 or 2. (Appendix 4) The predetermined reward is represented by a distance between the lever-type door handle and the end effector, a rotation angle of the lever-type door handle, a rotation angle of a door hinge of the swing door, and a norm of an output of the trained model. A control device for a robot arm according to any one of Supplementary Note 1 to Supplementary Note 3. (Appendix 5) Further comprising an operation control unit that controls the robot arm according to the control data output from the trained model, A control device for a robot arm according to any one of Supplementary Note 1 to Supplementary Note 4. (Appendix 6) The trained model is a model including a convolutional neural network and a multilayer perceptron. A control device for a robot arm according to any one of Supplementary Note 1 to Supplementary Note 5. (Appendix 7) A method for controlling a robot arm that performs an opening operation of a swing door having a lever-type door handle, comprising: acquiring a depth image of the lever-type door handle and the area around the lever-type door handle from a camera provided on the robot arm; From the depth images, a first depth image representing the distance between the lever-type door handle and the end effector of the robot arm and a second depth image in which the uneven shape of the lever-type door handle is emphasized are generated, and a composite depth image is generated by combining the first depth image and the second depth image; a process of generating a trained model that uses the synthetic depth image and the position information of the end effector as inputs and outputs control data for the robot arm corresponding to the opening operation of the swing door by performing reinforcement learning using the synthetic depth image and the position information of the end effector so that a predetermined reward increases at each predetermined time step; A computer-implemented method for controlling a robotic arm. (Appendix 8) A control program for a robot arm that performs an opening operation of a swing door having a lever-type door handle, acquiring a depth image of the lever-type door handle and the area around the lever-type door handle from a camera provided on the robot arm; From the depth images, a first depth image representing the distance between the lever-type door handle and the end effector of the robot arm and a second depth image in which the uneven shape of the lever-type door handle is emphasized are generated, and a composite depth image is generated by combining the first depth image and the second depth image; a process of generating a trained model that uses the synthetic depth image and the position information of the end effector as inputs and outputs control data for the robot arm corresponding to the opening operation of the swing door by performing reinforcement learning using the synthetic depth image and the position information of the end effector so that a predetermined reward increases at each predetermined time step; A control program for the robot arm to be executed by a computer. [Explanation of symbols]

[0105] 10 Control device 11 CPU 12 ROM 13 RAM 14 I / O 15 Storage section 15A Control Program 16 Connection 20 Support stand 21 wheels 30 Manipulator 40 End Effector 50 Depth Camera 55 Encoder 60, 60A~60D Swing Door 61 Lever door handle 100 Robot Arm

Claims

1. A control device for a robot arm that performs an opening operation of a swing door having a lever-type door handle, an acquisition unit that acquires a depth image of the lever-type door handle and the area around the lever-type door handle from a camera included in the robot arm; a generating unit that generates, from the depth image, a first depth image representing the distance between the lever-type door handle and the end effector of the robot arm, and a second depth image in which the uneven shape of the lever-type door handle is emphasized, and generates a composite depth image in which the first depth image and the second depth image are combined; a learning unit that uses the synthetic depth image and the position information of the end effector to perform reinforcement learning so that a predetermined reward increases at each predetermined time step, thereby generating a trained model that receives the synthetic depth image and the position information of the end effector as inputs and outputs control data for the robot arm corresponding to the opening operation of the swing door; A control device for a robot arm comprising:

2. The generation unit performs threshold processing on each pixel value of the depth image using a threshold range based on the working distance of the camera, replaces pixel values ​​outside the threshold range with a predetermined value that does not belong to the threshold range, normalizes the replaced depth image obtained by the replacement within the threshold range to generate the first depth image, and normalizes the first depth image between a minimum value and a maximum value of pixel values ​​to generate the second depth image. The robot arm control device according to claim 1 .

3. The synthetic depth image is an image obtained by averaging the first depth image and the second depth image. The robot arm control device according to claim 1 .

4. The predetermined reward is represented by a distance between the lever-type door handle and the end effector, a rotation angle of the lever-type door handle, a rotation angle of a door hinge of the swing door, and a norm of an output of the trained model. The robot arm control device according to claim 1 .

5. Further comprising an operation control unit that controls the robot arm according to the control data output from the trained model, The robot arm control device according to claim 1 .

6. The trained model is a model including a convolutional neural network and a multilayer perceptron. The robot arm control device according to claim 1 .

7. A method for controlling a robot arm that performs an opening operation of a swing door having a lever-type door handle, comprising: acquiring a depth image of the lever-type door handle and the area around the lever-type door handle from a camera provided on the robot arm; From the depth images, a first depth image representing the distance between the lever-type door handle and the end effector of the robot arm and a second depth image in which the uneven shape of the lever-type door handle is emphasized are generated, and a composite depth image is generated by combining the first depth image and the second depth image; a process of generating a trained model that uses the synthetic depth image and the position information of the end effector as inputs and outputs control data for the robot arm corresponding to the opening operation of the swing door by performing reinforcement learning using the synthetic depth image and the position information of the end effector so that a predetermined reward increases at each predetermined time step; A computer-implemented method for controlling a robotic arm.

8. A control program for a robot arm that performs an opening operation of a swing door having a lever-type door handle, acquiring a depth image of the lever-type door handle and the area around the lever-type door handle from a camera provided on the robot arm; From the depth images, a first depth image representing the distance between the lever-type door handle and the end effector of the robot arm and a second depth image in which the uneven shape of the lever-type door handle is emphasized are generated, and a composite depth image is generated by combining the first depth image and the second depth image; a process of generating a trained model that uses the synthetic depth image and the position information of the end effector as inputs and outputs control data for the robot arm corresponding to the opening operation of the swing door by performing reinforcement learning using the synthetic depth image and the position information of the end effector so that a predetermined reward increases at each predetermined time step; A control program for the robot arm to be executed by a computer.