Remote operation control device, remote operation control method, and program

The remote operation control system addresses the challenge of operating multiple objects by estimating operator and object movements and postures, generating control commands to enhance operability through relationship-based constraints.

JP2025102036APending Publication Date: 2025-07-08HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023219221
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing technologies face difficulties in remotely operating multiple objects due to the lack of effective methods for handling the relationships between objects, leading to challenges in determining appropriate operation constraints.

Method used

A remote operation control system that utilizes environmental and operator sensors to estimate the movement and posture of both the operator and objects, acquires relationships between them, and generates control commands based on these estimations to manage multiple objects effectively.

Benefits of technology

The system enhances operability by applying operation constraints considering the relationships between objects, enabling the handling of multiple objects with improved precision and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102036000001_ABST
    Figure 2025102036000001_ABST
Patent Text Reader

Abstract

To provide a remote operation control device, a remote operation control method, and program which can handle a plurality of objects.SOLUTION: A remote operation control device comprises: an intention estimation part which estimates behavior of an operator based on a first sensor value obtained by an environment sensor acquiring information of a robot or a peripheral environment and a second sensor value representing behavior of the operator obtained by an operator sensor; a relationship acquisition part which acquires relationship between a first operation object and a second operation object; and a control command generation part which generates a control command based on the estimated behavior of the operator and the information obtained by the relationship acquisition part.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote operation control device, a remote operation control method, and a program.

Background Art

[0002] In order to remotely operate a robot having a multi-fingered hand (including two fingers) using a non-exoskeleton type operation interface, the operability is improved by restricting the robot trajectory in accordance with an operation determined for each object.

[0003] For example, in the technique described in Patent Document 1, in a robot remote operation in which the movement of an operator is recognized and the robot is operated by transmitting the movement of the operator to the robot, an intention estimation unit is configured to estimate the movement of the operator based on a robot environment sensor value obtained by an environment sensor installed in the robot or the surrounding environment of the robot and an operator sensor value that is the movement of the operator obtained by an operator sensor, and a control command generation unit generates an appropriate control command for some degrees of freedom of the operator's movement based on the estimated movement of the operator, thereby reducing the degrees of freedom of the operator's movement and generating a control command.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the prior art has a problem that it is difficult to remotely operate multiple objects.

[0006] The present invention has been made in view of the above problems, and an object thereof is to provide a remote operation control device, a remote operation control method, and a program capable of handling multiple objects.

Means for Solving the Problems

[0007] (1) To achieve the above object, in a robot remote operation that recognizes the movement of an operator and transmits the movement of the operator to a robot to operate the robot, an intention estimation unit that estimates the operation of the operator based on a first sensor value obtained by an environmental sensor that acquires information on the robot or the surrounding environment of the robot and a second sensor value that is the movement of the operator obtained by an operator sensor, a first operation target object, a relationship acquisition unit that acquires the relationship between the first operation target object and the second target object, and a control command generation unit that generates a control command based on the estimated operation of the operator and the information acquired by the relationship acquisition unit.

[0008] (2) In the remote operation control device according to an aspect of the present invention described in (1) above, a robot whole-body joint angle estimation unit that estimates the joint angles of the whole body of the robot based on the finger joint angles and wrist postures of the operator included in the second sensor value is further provided, and the control command generation unit may generate a control command using the estimated joint angles of the whole body of the robot as well.

[0009] (3) In the remote operation control device according to an aspect of the present invention described in (1) or (2) above, an object posture estimation unit that estimates the postures of the first operation target object and the second target object based on the first sensor value is further provided, and the control command generation unit may generate a control command using the estimated postures of the first operation target object and the second target object as well. It may be like this.

[0010] (4) In the remote operation control device according to an aspect of the present invention described in any one of (1) to (3) above, the control command generation unit may reduce the degrees of freedom of the operator's operation by generating an appropriate control command for some of the degrees of freedom of the operator's operation and generate a control command.

[0011] (5) In the remote operation control device according to any one of the above (1) to (4) aspects of the present invention, the second target object may be an object related to the first operation target object in the operation.

[0012] (6) In the remote operation control device according to any one of the above (1) to (5) aspects of the present invention, when there are a plurality of operation target objects, the relationship acquisition unit uses the information obtained by identifying the target object from the first sensor value from a database in which the relationships between the operation target objects or between the operation target object and the surrounding environment are described in advance to acquire relationship information.

[0013] (7) In the remote operation control device according to any one of the above (1) to (6) aspects of the present invention, the relationship acquisition unit extracts the appearance feature amounts and geometric feature amounts of each of the operation target objects using the identification information regarding the object detected using the image included in the first sensor value, the position and size of the region of interest, and the luminance within the region of interest, and uses the extracted appearance feature amounts and geometric feature amounts of each of the operation target objects to acquire the relationship between the first operation target object and the second target object and output the relationship as text.

[0014] (8) In the remote operation control device according to one aspect of the present invention described in (1) above, based on the finger joint angles and wrist postures of the operator included in the second sensor value, a robot whole body joint angle estimation unit that estimates the joint angles of the whole body of the robot, and based on the first sensor value, an object posture estimation unit that estimates the postures of the first operation target object and the second target object respectively. The relationship acquisition unit uses the identification information regarding the object detected using the image included in the first sensor value, the position and size of the region of interest, and the luminance within the region of interest to extract the appearance feature amounts and geometric feature amounts of each of the operation target objects. Using the extracted appearance feature amounts and geometric feature amounts of each of the operation target objects, the relationship between the first operation target object and the second target object is acquired and the relationship is output in text. The control command generation unit encodes a set of the estimated motion of the operator, the estimated joint angles of the whole body of the robot, and the postures of the first operation target object and the second target object respectively to generate a first feature vector, encodes the text output by the relationship acquisition unit to generate a second feature vector, associates the first feature vector and the second feature vector, and encodes the associated data to generate a control command that is a joint angle trajectory sequence of the robot.

[0015] (9) To achieve the above object, a remote operation control method according to one aspect of the present invention is a remote operation control method for performing a robot remote operation that recognizes the motion of an operator and transmits the motion of the operator to a robot to operate the robot. An intention estimation unit estimates the operation of the operator based on a first sensor value obtained by an environment sensor that acquires information on the robot or the surrounding environment of the robot and a second sensor value that is the motion of the operator obtained by an operator sensor. A relationship acquisition unit acquires the relationship between a first operation target object and a second target object. A control command generation unit generates a control command based on the estimated operation of the operator and the information acquired by the relationship acquisition unit.

[0016] (10) To achieve the above object, a program according to an aspect of the present invention causes a computer of a remote operation control device that performs robot remote operation for recognizing the movement of an operator and transmitting the movement of the operator to a robot to operate the robot, to estimate the operation of the operator based on a first sensor value obtained by an environment sensor that acquires information on the robot or the surrounding environment of the robot, and a second sensor value that is the movement of the operator obtained by an operator sensor, acquire the relationship between a first operation target object and a second target object, and generate a control command based on the estimated operation of the operator, the acquired first operation target object, and the information on the relationship between the second target objects.

Advantages of the Invention

[0017] According to the above (1) to (10), since the related information between objects is used, a plurality of objects can be handled. According to the above (1) to (10), since operation constraints considering the relationship between objects are applied, the operability is improved.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0019] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each member is appropriately changed in order to make each member recognizable in size. In all the drawings for explaining the embodiments, those having the same function are denoted by the same reference numerals, and repeated explanations are omitted. In addition, “based on XX” as used in the present application means “at least based on XX”, and includes cases where it is based on another element in addition to XX. Further, “based on XX” is not limited to the case where XX is directly used, and includes cases where it is based on something obtained by performing arithmetic operations or processing on XX. “XX” is an arbitrary element (for example, arbitrary information).

[0020] [Overview of Remote Operation and Object to be Operated] First, an overview of the remote operation and the object to be operated will be described. FIG. 1 is a diagram for explaining the remote operation of the robot and the overview of the object to be operated. As shown in FIG. 1, in the remote operation space, an operator Us wears, for example, an HMD (head-mounted display) 4 on the head and wears operation units 5 (5L, 5R) such as data gloves on the hands. An environmental sensor 3 is installed in the robot work space. Note that the environmental sensor 3 may be attached to the robot 2. The robot 2 includes a control device 6, end effectors 21 (first end effector 21L, second end effector 21R), and hands 211 (211L, 211R).

[0021] The target object obj is composed of a plurality of objects. The target object obj is, for example, a plastic bottle or a bottle, and includes a body and a cap. The operator Us operates the robot 2 remotely to operate the target object obj, for example, by moving the hand or finger wearing the operation unit 5 while looking at the image displayed on the HMD 4. An example of the operation content is, as shown in FIG. 2, attaching and closing the cap to the body, or opening the cap from the body, etc. FIG. 2 is a diagram showing an example of the operation content.

[0022] However, when operating a plurality of objects, the operation method cannot be determined without considering the relationship between the objects. As a result, a constraint method is not required. Here, an example of the relationship between a plurality of objects when handling a plurality of objects will be described. For example, when closing the cap of a plastic bottle, the restraint condition on the receptacle side is obtained only from the plastic bottle, but the restraint condition on the connector side is obtained only from the cap. Only by superimposing these two constraint conditions can a meaningful operation constraint be obtained.

[0023] [Configuration of the Remote Operation Control System] Next, a configuration example of the remote operation control system 1 will be described. FIG. 3 is a diagram showing a configuration example of the remote operation control system according to the present embodiment. As shown in FIG. 3, the remote operation control system 1 includes, for example, a robot 2, an environmental sensor 3, an HMD 4, an operation unit 5, a control device 6 (remote operation control device), and a DB 8.

[0024] The robot 2 includes, for example, a first end effector 21L and a second end effector 21R. The first end effector 21L includes, for example, a hand 211L, an actuator 212L, and a sensor 213L. The second end effector 21R includes, for example, a hand 211R, an actuator 212R, and a sensor 213R. Note that the robot 2 may include a communication unit, a power supply unit, a body, legs, a head, etc. not shown in the figure. In the following description, when the first end effector 21L and the second end effector 21R are not distinguished, they are also referred to as the end effector 21. Similarly, when the hand 211L and the hand 211R are not distinguished, they are also referred to as the hand 211, when the actuator 212L and the actuator 212R are not distinguished, they are also referred to as the actuator 212, and when the sensor 213L and the sensor 213R are not distinguished, they are also referred to as the sensor 213. The robot 2 transmits and receives various information to and from the control device 6 via a wired or wireless network NW.

[0025] The HMD 4 includes, for example, an image display unit 41 and a line-of-sight detection unit 42. Note that the HMD 4 also includes a communication unit, a power supply unit, etc., which are not shown in the figure. The HMD 4 transmits and receives various information to and from the control device 6 via a wired or wireless network NW.

[0026] The operation unit 5 includes, for example, a sensor 51. Note that the operation unit 5 includes a communication unit, which is not shown in the figure. The operation unit 5 transmits information to the control device 6 via a wired or wireless network NW.

[0027] The control device 6 includes, for example, an acquisition unit 61, a detection unit 62, an object relationship acquisition unit 63, an affordance estimation unit 64, an intention estimation unit 65, an object posture estimation unit 66, a robot whole-body joint angle estimation unit 67, a control command generation unit 68, a drive circuit 69, and an image generation unit 70. Note that the control device 6 also includes a communication unit, a power supply unit, etc., which are not shown in the figure. The control device 6 transmits and receives various information to and from the robot 2 and the HDM 4 via a wired or wireless network NW. The control device 6 receives various information from the environment sensor 3 and the operation unit 5 via a wired or wireless network NW.

[0028] Note that the configuration example shown in FIG. 3 is an example and is not limited thereto.

[0029] [Functions of Each Device in the Remote Operation Control System] Next, the functions of each device in the remote operation control system will be described with reference to FIG. 3. (Robot 2) The hand 211 includes, for example, a plurality of finger portions. Each finger portion includes a joint. Note that the hand 211 may be a gripper or the like. The actuator 212 is attached to each joint. The sensor 213 is, for example, a six-axis sensor attached to a joint, a tactile sensor attached to a finger portion, or the like. The six-axis sensor detects forces in three axes (x, y, z) and moments in three axes (α, β, γ). Note that the robot 2 may include a drive circuit that drives the actuator 212.

[0030] (Environmental sensor 3) As shown in FIG. 1, the environmental sensor 3 is installed, for example, in a robot working space. The environmental sensor 3 is, for example, an RGB-D camera that acquires RGB (red, green, blue) information and depth information. Note that the information is acquired, for example, at predetermined time intervals.

[0031] (HMD 4) The image display unit 41 displays an image output by the control device 6. The gaze detection unit 42 detects the gaze direction and movement of the operator Us.

[0032] (Operation unit 5) The operation unit 5 is, for example, a data glove. The operation unit 5 detects the finger joint angles and wrist joint angles of the operator Us's hand.

[0033] (Control device 6) The acquisition unit 61 acquires the first sensor value detected by the sensor 213 from the robot 2. The acquisition unit 61 acquires the first sensor value detected by the environmental sensor 3. The acquisition unit 61 acquires the second sensor value detected by the sensor 51 of the operation unit 5. The acquisition unit 61 acquires the second sensor value detected by the gaze detection unit 42 of the HMD 4.

[0034] The detection unit 62 detects an object using the first sensor value acquired by the acquisition unit 61, and for each target object, detects point cloud data, identification information ID for identifying the object, a region of interest (ROI), and a luminance value. Note that the configuration example and processing example of the detection unit 62 will be described in detail later.

[0035] The object relationship acquisition unit 63 uses the identification information ID, region of interest, and luminance value for identifying the object detected by the detection unit 62 to extract feature amounts for each target object, estimate the relationships between the plurality of target objects, and generate a generated caption. Note that the generated caption is information indicating the relationship between objects, for example, in natural language (text). Note that the configuration example and processing example of the object relationship acquisition unit 63 will be described in detail later.

[0036] The affordance estimation unit 64 estimates an affordance using the point cloud data detected by the detection unit 62. Note that the configuration example and processing example of the affordance estimation unit 64 will be described in detail later.

[0037] The intention estimation unit 65 estimates the operation intention of the operator Us using the information indicating the affordance estimated by the affordance estimation unit 64, the gaze information (second sensor value) detected by the HMD 4, and the finger joint angle and wrist joint angle detected by the operation unit 5, and generates a point cloud with affordance.

[0038] The object pose estimation unit 66 estimates the pose of each operating object using the point cloud data detected by the detection unit 62. Note that the configuration example and processing example of the object pose estimation unit 66 will be described in detail later.

[0039] The robot whole body joint angle estimation unit 67 estimates the whole body joint angles of the robot 2 using the finger joint angle and wrist joint angle detected by the operation unit 5, and generates a joint angle trajectory sequence. Note that the whole body joint angles are, for example, the joint angles of each end effector 21, for example, the joint angles of the arm, the joint angle of the wrist, the joint angles of the finger parts, etc. Note that the configuration example and processing example of the robot whole body joint angle estimation unit 67 will be described in detail later.

[0040] Based on the generated caption output by the object relationship acquisition unit 63, the point cloud data with affordance output by the intention estimation unit 65, the joint angle trajectory sequence output by the robot whole-body joint angle estimation unit 67, and the information indicating the object posture of each target object output by the object posture estimation unit 66, the control command generation unit 68 generates a control command. Note that the configuration example and processing example of the control command generation unit 68 will be described in detail later.

[0041] Based on the control command generated by the control command generation unit 68, the drive circuit 69 generates a drive signal for controlling the robot 2. When the robot 2 is equipped with the drive circuit 69, the control device 6 may not be equipped with the drive circuit 69. Alternatively, the control device 6 and the robot 2 may each be equipped with a part of the drive circuit 69.

[0042] Based on the image captured by the environment sensor 3 and the like, the image generation unit 70 generates an image to be provided to the HMD 4. Note that, for the generation method of the image displayed on the HMD 4 and examples of the image, for example, the method described in the pamphlet of Japanese Patent Application No. 2022-156322 is used.

[0043] The DB 8 is a database that stores the result estimated by the affordance estimation unit 64.

[0044] [Processing configuration block] Next, the processing configuration block of the control device 6 will be described. FIG. 4 is a diagram showing the processing configuration block of the control device according to the present embodiment. FIGS. 5 to 10 are block diagrams of each configuration. FIG. 5 is a diagram showing a configuration example of the detection unit of the present embodiment. FIG. 6 is a diagram showing a configuration example of the object relationship estimation unit of the present embodiment. FIG. 7 is a diagram showing a configuration example of the affordance estimation unit of the present embodiment. FIG. 8 is a diagram showing a configuration example of the object relationship estimation unit of the present embodiment. FIG. 9 is a diagram showing a configuration example of the robot whole-body joint angle estimation unit of the present embodiment. FIG. 10 is a diagram showing a configuration example of the control command generation unit of the present embodiment.

[0045] (Detection unit 62) First, the detection unit 62 will be described with reference to FIGS. 4 and 5. As shown in FIG. 5, the detection unit 62 includes, for example, an RGB-D sensor 621, an object detection unit 622, a self-position estimation unit 623, an addition unit 624, a three-dimensional reconstruction unit 625, and a data integration unit 626.

[0046] The RGB-D sensor 621 corresponds to, for example, the sensors included in the environmental sensor 3. The RGB-D sensor 621 acquires an RGB image and a depth image. Note that the images acquired by the RGB-D sensor 621 include, for example, a plurality of target objects, the finger parts of the robot, etc. The RGB-D sensor 621 outputs the images to the object detection unit 622 and the self-position estimation unit 623.

[0047] The object detection unit 622 performs well-known image processing (e.g., binarization, image enhancement, feature amount extraction, contour extraction, clustering processing, etc.) on the images output by the RGB-D sensor 621 to detect each object. For each detected object, the object detection unit 622 outputs, in association, the identification information ID, the position and size of the region of interest ROI, and the luminance value in the ROI. Also, for each detected object, the object detection unit 622 outputs, in association with the identification information ID, a mask image to the addition unit 624. Note that the object detection unit 622 performs, for each detected object, detection of the region of interest ROI, generation of a mask image, etc. by image processing. Note that the object detection unit 622 may input the image and the teacher data (a set of the identification information ID, the region of interest ROI, and the luminance value for each object, a set of the identification information ID and the mask image for each object) into a model, output a set of the identification information ID, the region of interest ROI, and the luminance value for each object, and a set of the identification information ID and the mask image for each object, and perform object detection using a model such as R-CNN (Region CNN) that is learned until the difference between the teacher data and the output is within a predetermined value. Also, for the detection of the region of interest ROI, for example, the Selective Search method may be used.

[0048] The self-position estimation unit 623 estimates the location of the robot coordinates for each object image output by the RGB-D sensor 621, for example, with reference to a pre-determined map coordinate system.

[0049] The addition unit 624 combines, for each object, the identification information ID and the mask image output by the object detection unit 622 and the depth image output by the RGB-D sensor 621, and outputs the combined data to the three-dimensional reconstruction unit 625.

[0050] The three-dimensional reconstruction unit 625 reconstructs by converting from the depth image to a three-dimensional point cloud using the output of the addition unit 624 and the position information for each object output by the self-position estimation unit 623. The reference coordinate system of the point cloud at this time is the robot coordinate system (it is assumed that the conversion from the sensor coordinate system to the robot coordinate system is internally performed). Note that the point cloud data includes position information. Note that the point cloud output by the three-dimensional reconstruction unit 625 is based on the robot coordinates at time t. Therefore, the three-dimensional reconstruction unit 625 performs coordinate transformation as a point cloud in the map coordinate system using the "robot coordinates at time t in the map coordinate system" output by the self-position estimation unit 623. In this embodiment, by performing these processes, a point cloud in the map coordinate system can be obtained regardless of the time, and the data can be integrated simply by adding.

[0051] The data integration unit 626 integrates, for each object, the three-dimensional reconstructed point cloud data output by the three-dimensional reconstruction unit 625 in time series based on the position of each object output by the self-position estimation unit 623. The data integration unit 626 outputs the integrated point cloud data for each object.

[0052] (Object relationship acquisition unit 63) Next, the object relationship acquisition unit 63 will be described with reference to FIGS. 4 and 6. As shown in FIG. 6, the object relationship acquisition unit 63 includes, for example, an object detection unit 622, a feature extraction unit 631, and a relationship estimation unit 632. As shown in FIG. 6, the object relationship acquisition unit 63 also uses the object detection unit 622 as the detection unit 62.

[0053] The feature extraction unit 631 extracts the appearance feature amount and geometric feature amount for each object by a well-known method using the data of the identification information ID for each object, the region of interest ROI, and the luminance value output by the object detection unit 622.

[0054] The relationship estimation unit 632 includes, for example, an encoder 633 and a decoder 634. The relationship estimation unit 632 inputs the appearance feature amount and geometric feature amount for each object output by the feature extraction unit 631 to the learned encoder 633, and inputs the output of the encoder 633 to the learned decoder 634 to output a generated caption.

[0055] In the encoder-decoder model (or network), the encoder performs, for example, a bottom-up encoding process on the input in order and outputs a low-dimensional latent representation. The decoder performs, for example, a top-down decoding process on the low-dimensional latent representation output by the encoder. Also, the encoder and decoder are composed of, for example, networks such as an RNN (Recurrent Neural Network) or a CNN (Convolutional neural network). Further, the encoder 633 and the decoder 634 are learned, for example, by repeatedly adding teacher data to each input and inputting it until the difference between the output and the teacher data is within a predetermined value. Note that the relationship estimation unit 632 is not limited to an encoder-decoder, and may be, for example, a Transformer which is a network architecture in which an encoder and a decoder are connected by an Attention model.

[0056] (Affordance Estimation Block 640) Next, the affordance estimation block 640 will be described with reference to FIGS. 4 and 7. As shown in FIG. 7, the affordance estimation block 640 includes, for example, an affordance estimation unit 64 and a DB8.

[0057] The affordance estimation unit 64 includes, for example, a point cloud encoder 641 and a point cloud decoder 642. Note that the encoder and decoder may be network-based or may be transformers. Note that the point cloud encoder 641 and the point cloud decoder 642 are, for example, inputted by adding teacher data to each input in advance, and learning is repeated until the difference between the output and the teacher data is within a predetermined value.

[0058] The learned point cloud encoder 641 of the affordance estimation unit 64 is inputted with the point cloud data for each object outputted by the detection unit 62. The learned point cloud encoder 641 encodes the inputted point cloud data and outputs a low-dimensional latent representation to the learned point cloud decoder 642. The learned point cloud decoder 642 decodes the inputted low-dimensional latent representation and outputs the result of estimating the affordance. The output information is associated with the information indicating the estimated affordance for each of the inputted point cloud data. Note that the information regarding the affordance is, for example, information obtained from the shape of the target object or the category of the target object, such as the operation content of attaching and turning the lid to close the mouth of the body in the case of the body and lid of a plastic bottle. Then, the affordance estimation unit 64 stores the estimation result in the DB 8 and further outputs it to the intention estimation unit 65.

[0059] Note that the affordance estimation unit 64 may stop the process after estimating the affordance in a series of operations.

[0060] (Intention Estimation Unit 65) Next, the intention estimation unit 65 will be described with reference to FIG. 4. The intention estimation unit 65 receives information indicating the affordance output by the affordance estimation block 640, the gaze information detected by the HMD 4, and the information on the knuckle angles and wrist postures output by the operation unit 5. Note that the information indicating the affordance also includes information on the object to be operated. The intention estimation unit 65 estimates the operation intention of the operator Us using the input information. That is, the intention estimation unit 65 estimates what kind of operation is intended for each target object. Note that the intention estimation may be performed using a taxonomy (see, for example, Reference 1), or based on the method described in, for example, Japanese Patent Application No. 2022-008829, and further using the information indicating the affordance. In the present embodiment, since there are a plurality of target objects, the operation intention for the plurality of target objects can be appropriately estimated by also using the information indicating the affordance.

[0061] Note that the intention estimation unit 65 may be configured by a learned encoder and decoder. The encoder and decoder may be network-based or transformers. Note that the data output by the intention estimation unit 65 is point cloud data with affordances including the estimated operation intention. Note that the point cloud data with affordances is data including affordance information, for example, when represented as a heat map, the temperature of the pressed or grasped part is displayed high, or color-coded for each affordance (see, for example, Patent Documents 2 and 3). For example, when the target objects are a plastic bottle and a lid, it is an image in which affordances are shown for the main body part to be grasped, the drinking mouth part where the lid is attached, and the lid respectively.

[0062] Reference 1: Thomas Feix, Javier Romero, et al., “The GRASP Taxonomy of Human Grasp Types” IEEE Transactions on Human-Machine Systems (Volume: 46, Issue: 1, Feb. 2016), IEEE, p66 - 77 Reference 2: Yuanzhi Liang, Xiaohan Wang, et al., “MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects”, ICCV 2023, 2023, p217-227 Reference 3: Hideichi Akizuki, Yoshimitsu Aoki, “6-Degree-of-Freedom Pose Estimation of Similar Shaped Objects Focusing on the Spatial Arrangement of Functional Attributes”, Transactions of the JSPE, Vol. 85, No. 1, JSPE, 2019

[0063] (Object pose estimation unit 66) Next, the object pose estimation unit 66 will be described with reference to FIGS. 4 and 8. As shown in FIG. 8, the object pose estimation unit 66 includes, for example, a point cloud encoder 661 and a point cloud decoder 662. Note that the encoder and decoder may be network-based or transformers. The point cloud encoder 661 and the point cloud decoder 662 are, for example, trained by repeatedly adding teacher data to each input and inputting it until the difference between the output and the teacher data is within a predetermined value.

[0064] The point cloud data for each object output by the detection unit 62 is input to the trained point cloud encoder 661, and the input point cloud data is encoded to output, for example, a low-dimensional latent representation to the trained point cloud decoder 662. The trained point cloud decoder 662 decodes the low-dimensional latent representation to estimate the pose of each object, and outputs information indicating the estimated pose of each object to the control command generation unit 68. Note that the estimated object pose also includes the time-series trajectory of the pose change of each object.

[0065] (Robot whole-body joint angle estimation unit 67) Next, the robot whole-body joint angle estimation unit 67 will be described with reference to FIGS. 4 and 9. As shown in Fig. 9, the robot's whole-body joint angle estimator 67 includes, for example, a point cloud encoder 671 and a point cloud decoder 672. Note that the encoder and decoder may be network-based or transformers. The point cloud encoder 671 and the point cloud decoder 672 are, for example, input with teacher data added to each input in advance and repeatedly trained until the difference between the output and the teacher data is within a predetermined value.

[0066] Note that the point cloud encoder 671 and the point cloud decoder 672 estimate, for example, the posture of the operator Us from the first-person perspective and further absorb the dimensional differences between the operator Us and the robot 2 to estimate the whole-body joint angles of the robot. Information indicating the finger joint angles and the wrist posture is input from the operation unit 5 to the trained point cloud encoder 671. The trained point cloud encoder 671 encodes the input information and outputs, for example, a low-dimensional latent representation to the trained point cloud decoder 672. Note that the information indicating the finger joint angles and the wrist posture includes information for each of the left and right hands of the operator Us. The trained point cloud decoder 672 decodes the low-dimensional latent representation to estimate a joint angle trajectory sequence and outputs the estimated joint angle trajectory sequence data to the control command generation unit 68.

[0067] Note that in the prior art, although it has been proposed to estimate a person's joint angles using the first-person perspective of the person, it has not been possible to estimate the whole-body joint angles of the robot 2. In contrast, in the present embodiment, the whole-body joint angles of the robot 2 are estimated using the movements of the fingers and wrists of the operator Us.

[0068] (Control command generation unit 68) Next, the control command generation unit 68 will be described with reference to Figs. 4 and 10. As shown in Fig. 10, the control command generation unit 68 includes, for example, a motion encoder 681, a text encoder 682, and a motion decoder 683. The motion encoder 681 receives the affordance point cloud data, the joint angle trajectory sequence, and the information indicating the object pose. The motion encoder 681 encodes the input data and outputs, for example, a low-dimensional latent representation. The text encoder 682 receives the generated caption. The text encoder 682 encodes the input generated caption and outputs, for example, a low-dimensional latent representation.

[0069] Note that the motion encoder 681, the text encoder 682, and the motion decoder 683 perform learning by, for example, adding teacher data to each input in advance and repeating the process until the difference between the output and the teacher data is within a predetermined value.

[0070] The control command generation unit 68 may be based on, for example, OpenAI's CLIP (Contrastive Language-Image Pre-training). A matrix 684 is generated using the feature vectors that are the low-dimensional latent representations encoded by the motion encoder 681 and the text encoder 682, respectively. The motion decoder 683 estimates and outputs the joint trajectory sequence generated by decoding the feature amounts of the motion encoder 681 and the text encoder 682, respectively. In this way, in the present embodiment, by associating the generated caption with the motion information (affordance, joint angle trajectory sequence, object pose), a joint angle command that can appropriately control a plurality of objects can be generated.

[0071] With this configuration, the control device 6 can generate, as joint angle commands for the robot 2, how to operate on a certain object. The control device 6 may use the generated joint angle commands as constraints, and can control the operation of the robot 2 using the generated joint angle commands. Note that when the control command generation unit 68 uses the generated joint angle trajectory example as a constraint, for example, by generating appropriate control commands for some degrees of freedom of the operation of the operator Us, the degrees of freedom of the operation of the operator may be reduced to generate control commands.

[0072] [Example of processing procedure] Next, an example of the processing procedure performed by the remote operation control system 1 will be described. FIG. 11 is a flowchart of an example of the processing procedure performed by the remote operation control system according to the present embodiment.

[0073] (Step S1) The acquisition unit 61 acquires the first sensor value detected by the sensor 213. The acquisition unit 61 acquires the first sensor value detected by the environment sensor 3. The acquisition unit 61 acquires the second sensor value detected by the sensor 51. The acquisition unit 61 acquires the second sensor value detected by the gaze detection unit 42.

[0074] (Step S2) The detection unit 62 detects an object using the first sensor value acquired by the acquisition unit 61, and for each target object, detects identification information ID, the position and size of the region of interest ROI, and the luminance value in the region of interest ROI.

[0075] (Step S3) The detection unit 62 detects an object using the first sensor value acquired by the acquisition unit 61, and generates three-dimensional point cloud data for each target object in time series.

[0076] (Step S4) The object relationship acquisition unit 63 uses the identification information ID, region of interest, and luminance value for identifying the object detected by the detection unit 62 to extract feature amounts for each target object, estimate the relationships between the plurality of target objects, and generate a generated caption.

[0077] (Step S5) The affordance estimation unit 64 estimates the affordance using the point cloud data detected by the detection unit 62.

[0078] (Step S6) The intention estimation unit 65 estimates the operation intention of the operator Us using the information indicating the affordance, the gaze information (second sensor value), the finger joint angles, and the wrist joint angles, and generates a point cloud with affordance.

[0079] (Step S7) The object pose estimation unit 66 estimates the pose of each of the manipulated objects using the three-dimensional point cloud data.

[0080] (Step S8) The robot whole body joint angle estimation unit 67 estimates the whole body joint angles of the robot 2 using the finger joint angles and the wrist joint angles detected by the operation unit 5, and generates a joint angle trajectory sequence.

[0081] (Step S9) The control command generation unit 68 generates a control command using the generated caption, the point cloud data with affordance, the joint angle trajectory sequence, and the information indicating the object pose of each target object.

[0082] (Step S10) The control command generation unit 68 determines whether or not the remote operation by the operator Us has ended. Note that the operator Us may indicate the end of the remote operation by performing a predetermined action manually, by performing a predetermined eye movement such as a blink, or by turning off the power switch provided in the HMD 4. When the remote operation has ended (Step S10; YES), the control command generation unit 68 ends the process. When the remote operation has not ended (Step S10; NO), the control command generation unit 68 returns the process to Step S1.

[0083]

[0084] The processing procedure described with reference to FIG. 11 is an example and is not limited thereto. For example, the control device 6 may perform some processes simultaneously or in parallel.In the above example, an example of estimating the operation intention was described, but it is not limited to this. For example, when there is only one correspondence between the relationship and the operation method and there are no multiple correspondences, it may not be necessary to estimate the operation intention. In this case, the output of the affordance estimation block 640 may be output to the control command generation unit 68.

[0085] As described above, in the present embodiment, in addition to the affordance specific to the object, the relationship between objects or between the object and the environment is described in advance or the relationship between objects is obtained from an image via a caption using a large language model or the like. Since the relationship between objects and the operation method often correspond one-to-one, the operation method = constraint conditions are obtained by searching using the relationship as a key. Further, in the present embodiment, when there are multiple correspondences between the relationship and the operation method, it is determined using the operation intention estimation result of the operator Us. And in the present embodiment, within the range that satisfies these constraint conditions, the human operation input is converted into a robot trajectory sequence and output.

[0086] As a result, according to the present embodiment, operation constraints considering the relationship between objects are applied, so the operability is improved.

[0087] In the above example, as an example of a plurality of objects, a plastic bottle and a lid were described as an example, but it is not limited to this. The plurality of objects may be, for example, "a screw and a member to which the screw is attached", "a case with a complicated shape and an object to be housed in the case", etc. Also, the number of the plurality of objects may be three or more. That is, in the embodiment, the "plurality of objects" is not limited to the target object, and may be the target object and the surrounding environment of the target object (for example, a table, a stand, a case, a floor, a wall, a ground, a ceiling, etc.).

[0088] Note that a program for realizing all or part of the functions of the control device 6 in the present invention may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to perform all or part of the processing performed by the control device 6. Here, the "computer system" is assumed to include hardware such as an OS and peripheral devices. Further, the "computer system" is also assumed to include a WWW system equipped with a homepage providing environment (or display environment). Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" also includes a volatile memory (RAM) inside a computer system that becomes a server or a client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and holds the program for a certain period of time.

[0089] Also, the above program may be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by a transmission wave in the transmission medium. Here, the "transmission medium" for transmitting the program refers to a medium having a function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication wire) such as a telephone line. Further, the above program may be for realizing a part of the functions described above. Furthermore, it may be a so-called difference file (difference program) that can realize the functions described above in combination with a program already recorded in a computer system.

[0090] As described above, the embodiments for carrying out the present invention have been described using the embodiments, but the present invention is not limited to such embodiments, and various modifications and substitutions can be made without departing from the gist of the present invention.

Explanation of Reference Numerals

[0091] 1…Remote operation control system, 2…Robot, 3…Environmental sensor, 4…HMD, 5…Operation unit, 6…Control device, 8…DB, 21…End effector, 21L…First end effector, 21R…Second end effector, 211, 211L, 211R…Hand, 212, 212L, 212R…Actuator, 213, 213, 213L, 213R…Sensor, 41…Image display unit, 42…Gaze detection unit, 51…Sensor, 61…Acquisition unit, 62…Detection unit, 63…Object relationship acquisition unit, 64…Affordance estimation unit, 65…Intention estimation unit, 66…Object pose estimation unit, 67…Robot whole body joint angle estimation unit, 68…Control command generation unit, 69…Drive circuit, 70…Image generation unit, NW…Network, 621…RGB-D sensor, 622…Object detection unit, 623…Self-position estimation unit, 624…Addition unit, 625…Three-dimensional reconstruction unit, 626…Data integration unit, 631…Feature extraction unit, 632…Relationship estimation unit 632

Claims

1. In a robot remote operation that recognizes the movements of an operator and transmits the movements of the operator to a robot to operate the robot, an intention estimation unit that estimates the actions of the operator based on a first sensor value obtained by an environmental sensor that acquires information on the robot or the surrounding environment of the robot and a second sensor value that is the movement of the operator obtained by an operator sensor; a relationship acquisition unit that acquires the relationship between a first operation target object and a second target object; a control command generation unit that generates a control command based on the estimated action of the operator and the information acquired by the relationship acquisition unit; A remote operation control device comprising:

2. Further comprising a robot whole-body joint angle estimation unit that estimates the joint angles of the whole body of the robot based on the finger joint angles and wrist postures of the operator included in the second sensor value, The control command generation unit generates a control command using the estimated joint angles of the whole body of the robot as well. The remote operation control device according to claim 1.

3. Further comprising an object posture estimation unit that estimates the postures of the first operation target object and the second target object respectively based on the first sensor value, The control command generation unit generates a control command using the estimated postures of the first operation target object and the second target object respectively as well. The remote operation control device according to claim 1 or claim 2.

4. The control command generation unit Reduces the degrees of freedom of the operator's movement by generating appropriate control commands for some of the degrees of freedom of the operator's movement and generates a control command. The remote operation control device according to claim 1 or claim 2.

5. The second target object is an object related to the first operation target object in the operation. The remote operation control device according to claim 1 or claim 2.

6. The relationship acquisition unit When there are a plurality of operation target objects, the relationship information is acquired using the information obtained by specifying the target object from the first sensor value from a database in which the relationships between the operation target objects or between the operation target object and the surrounding environment are described in advance. The remote operation control device according to claim 1 or claim 2.

7. The relationship acquisition unit Using the identification information regarding the object detected using the image included in the first sensor value, the position and size of the region of interest, and the luminance within the region of interest, extract the appearance feature amounts and geometric feature amounts of each of the objects to be operated on, and using the extracted appearance feature amounts and geometric feature amounts of each of the objects to be operated on, obtain the relationship between the first object to be operated on and the second object and output the relationship as text. The remote operation control device according to claim 1 or claim 2.

8. A robot whole-body joint angle estimation unit that estimates the joint angles of the whole body of the robot based on the finger joint angles and wrist postures of the operator included in the second sensor value. Further comprising an object posture estimation unit that estimates the postures of the first object to be operated on and the second object based on the first sensor value. The relationship acquisition unit Using the identification information regarding the object detected using the image included in the first sensor value, the position and size of the region of interest, and the luminance within the region of interest, extract the appearance feature amounts and geometric feature amounts of each of the objects to be operated on, and using the extracted appearance feature amounts and geometric feature amounts of each of the objects to be operated on, obtain the relationship between the first object to be operated on and the second object and output the relationship as text. The control command generation unit Encode a set of the estimated motion of the operator, the estimated joint angles of the whole body of the robot, and the estimated postures of the first object to be operated on and the second object to generate a first feature vector. Encode the text output by the relationship acquisition unit to generate a second feature vector. Associate the first feature vector and the second feature vector, and generate a control command that is a joint angle trajectory sequence of the robot by encoding the associated data. The remote operation control device according to claim 1.

9. A remote operation control method for performing a robot remote operation that recognizes the movement of an operator and transmits the movement of the operator to the robot to operate the robot, An intention estimation unit estimates the operation of the operator based on a first sensor value obtained by an environment sensor that acquires information on the robot or the surrounding environment of the robot and a second sensor value that is the movement of the operator obtained by an operator sensor. A relationship acquisition unit acquires the relationship between a first object to be operated on and a second object. The control instruction generation unit generates a control instruction based on the estimated operation of the operator and the information acquired by the relationship acquisition unit. Remote operation control method.

10. In a computer of a remote operation control device that performs a robot remote operation for recognizing the movement of an operator and transmitting the movement of the operator to a robot to operate the robot, estimating the operation of the operator based on a first sensor value obtained by an environmental sensor that acquires information on the robot or the surrounding environment of the robot and a second sensor value that is the movement of the operator obtained by an operator sensor; acquiring the relationship between a first operation target object and a second target object; generating a control instruction based on the estimated operation of the operator and the acquired information on the relationship between the first operation target object and the second target object; Program.

Citation Information

Patent Citations

  • Robot remote operation control device, robot remote operation control system, robot remote operation control method and program

    JP2022157101A