Remote operation assistance system, remote operation assistance method, and program
The remote operation assistance system addresses the limitation of handling only registered objects by approximating shapes to geometric primitives and determining optimal grasping points, enabling flexible manipulation of diverse objects.
Patent Information
- Application Number
- JP2022136628
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2042-08-30
AI Technical Summary
Existing technologies can only handle objects that have been registered in the system in advance, limiting their ability to adapt to unfamiliar shapes and weights.
A remote operation assistance system that includes an action acquisition unit, intention understanding unit, environmental situation determination unit, and operation amount determination unit, which approximates object shapes to geometric primitives, compares with registered information, and determines optimal grasping points and paths using environmental data and operator intentions.
Enables grasping and manipulation of objects not previously registered, enhancing adaptability and efficiency in handling diverse objects.
Smart Images

Figure 0007809033000001 
Figure 0007809033000002 
Figure 0007809033000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a remote operation assistance system, a remote operation assistance method, and a program. [Background technology]
[0002] A control device for making a robot grasp an object has been proposed. As such a control device, a method has been proposed in which a grasp point that can be physically contacted is tentatively determined, a grasp force is determined taking into consideration the balance of forces, the quality of the grasp force is evaluated as the volume of the envelope of the grasp force, and a grasp point with higher quality is searched for (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6476358 Summary of the Invention [Problem to be solved by the invention]
[0004] In the technology described in Patent Document 1, the shape and gripping point of each object are registered in advance, the position and orientation of the object are estimated using the object shape, and then a gripping trajectory is generated from the relationship between the robot's current orientation and the gripping point. For this reason, the technology described in Patent Document 1 has the problem that it can only handle objects (shapes and weights) that have been registered in the system in advance.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a remote operation assistance system, a remote operation assistance method, and a program that can perform work on objects that have not been registered in the system in advance. [Means for solving the problem]
[0006] (1) In order to achieve the above object, a remote operation assistance system according to one aspect of the present invention is a remote operation assistance system for remotely operating at least an end effector, and includes: an action acquisition unit that acquires information related to the action of an operator operating the end effector; an intention understanding unit that uses the information acquired by the action acquisition unit to estimate a target object that is to be operated by the end effector and a task that is a method of operating the target object; an environmental situation determination unit that acquires environmental information about an environment in which the end effector is operated; and an operation amount determination unit that acquires information from the intention understanding unit and the environmental situation determination unit and determines an operation amount of the end effector from the acquired information, wherein the operation amount determination unit includes an approximation result reflection unit that approximates an object shape of the target object to geometric primitives and compares the approximation result with registered information of a pre-registered operation strategy for the end effector and reflects the result of the approximation.
[0007] (2) In the remote operation assistance system according to (1) of the present invention, the approximation result reflecting unit performs shape fitting of the geometric primitives for each of the instance IDs (identifiers) of the target object, where the environment information is shape information about the target object including an instance ID. Note that in this embodiment, the shape of the target object is voxel information or point cloud information.
[0008] (3) In the remote operation assistance system according to (1) or (2) of the aspect of the present invention, the geometric primitive is one or more, or a region of interest.
[0009] (4) Furthermore, in the remote operation assistance system according to (3) of one aspect of the present invention, when the geometric basic element is the attention area, the approximation result reflecting unit transforms an average shape of the target object into the geometric basic element to calculate a mapping function between areas, obtains a grasping probability map based on the motion and taxonomy estimated by the intention understanding unit and the class ID separated by the environmental situation determining unit, maps the grasping probability map to the area of the geometric basic element based on the calculated mapping function of correspondence between areas, and solves an optimization problem to determine a grasping point that has the shortest path and the shortest route with the highest probability while satisfying geometric constraints from the current wrist posture.
[0010] (5) In addition, in the remote operation assistance system according to any one of (1) to (4) of an aspect of the present invention, the environmental situation assessment unit includes an RGB sensor that acquires RGB data, a depth sensor that acquires a depth image, a mask generation unit that uses an instance segmentation technique to generate a mask image based on an instance ID of an object in an image using a color image output by the RGB sensor, an adder that generates a depth image for each instance ID using the depth image and the mask image, a self-location estimation unit that performs image processing on the color image output by the RGB sensor to detect a position of the environmental situation assessment unit, a three-dimensional reconstruction unit that uses posture information of the environmental situation assessment unit output by the self-location estimation unit and a depth image for each instance ID for each class output by the adder to obtain a point cloud for each instance ID for each class by three-dimensional reconstruction, and a data integration unit that acquires a point cloud for each instance ID for each key frame, acquires and manages current data at time t and time series data of past frames t-1, t-2, ..., and integrates these data.
[0011] (6) In order to achieve the above object, a remote operation assistance method according to one aspect of the present invention is a remote operation assistance method in a remote operation assistance system that remotely operates at least an end effector, the remote operation assistance method including: a motion acquisition step in which a motion acquisition unit acquires information related to at least a motion of an operator operating the end effector; an intention understanding step in which an intention understanding unit uses the information acquired in the motion acquisition step to infer a target object that is to be operated by the end effector and a task that is a method of operating the target object; an environmental situation determination step in which an environmental situation determination unit acquires environmental information of an environment in which the end effector is operated; and a manipulation amount determination step in which an operation amount determination unit acquires the information acquired in the intention understanding step and the information acquired in the environmental situation determination step and determines a manipulation amount of the end effector from the acquired information, wherein in the manipulation amount determination step, the manipulation amount determination unit approximates an object shape of the target object to geometric primitives and compares the approximation result with registered information of a pre-registered manipulation strategy for the end effector to reflect the approximated result.
[0012] (7) In order to achieve the above object, one aspect of the present invention provides a program for a computer in a remote operation assistance system that remotely operates at least an end effector, the program executing the following: acquiring motion information regarding the motion of an operator operating the end effector; using the motion information, estimating a target object that is to be operated by the end effector and a task that is a method of operating the target object; acquiring environmental information regarding an environment in which the end effector is operated; acquiring information on the estimated target object and task and the environmental information; determining an operation amount of the end effector from the acquired information; approximating a shape of the target object to geometric primitives in determining the operation amount of the end effector; and comparing the approximation result with registered information of a pre-registered operation strategy for the end effector to reflect the result. [Effects of the Invention]
[0013] According to (1) to (7), even objects that have not been registered in the system in advance can be worked on. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating a configuration example of a remote operation assistance system according to an embodiment; [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of an environmental situation determination unit according to the embodiment. [Figure 3] FIG. 1 is a diagram illustrating an overview of remote control according to an embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of a geometric primitive. [Figure 5] 10 is a flowchart of an example of a processing procedure performed by the remote operation assistance system according to the embodiment. [Figure 6] 10 is a flowchart of a processing procedure for a single primitive and multiple primitives according to an embodiment. [Figure 7] 1 is a diagram for explaining an outline of a configuration and an outline of a process of a remote operation assistance system according to an embodiment; [Figure 8] FIG. 10 is a diagram for explaining a comparative example. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).
[0016] [Example of remote operation support system configuration] Fig. 1 is a diagram showing an example of the configuration of a remote operation assistance system according to this embodiment. As shown in Fig. 1, the remote operation assistance system 1 includes an environmental situation determination unit 2 (2-1, 2-2, 2-3, ...), an HMD 3, an instruction detection unit 4, an action acquisition unit 5, an intention understanding unit 6, an operation amount determination unit 7, a robot 15, and a control unit 16. The HMD 3 includes, for example, a gaze detection unit 31. The operation amount determination unit 7 includes, for example, a model 8, a data integration unit 9, an approximation result reflection unit 17, a motion estimation unit 11, a grip DB 12, and a database 14. The approximation result reflecting unit 17 includes a shape approximation unit 10 and a grasp planning unit 13 . The robot 15 includes, for example, an end effector 151 , an arm 152 , and a sensor 153 .
[0017] The environmental situation determination unit 2 includes an RGB sensor and a depth sensor, as will be described later. The environmental situation determination unit 2 acquires environmental information about the environment in which the end effector 151 is operated. The environmental situation determination unit 2 outputs shape information about the target object. In this embodiment, the shape about the target object is, for example, voxel information or point cloud information. Furthermore, a plurality of environmental situation determination units 2 are installed in the operating environment, for example. An example configuration of the environmental situation determination unit 2 will be described later.
[0018] The HMD 3 is, for example, a head-mounted display device. The HMD 3 may be, for example, a glasses-type device, and may be for one eye or both eyes. The HMD 3 also includes a display device, and may acquire and display an image captured by the RGB sensor 211 (FIG. 2) of the environmental situation determination unit 2, or may display a virtual image generated using the acquired image data.
[0019] The line of sight detection unit 31, for example, captures an image of the operator's eyes and performs image processing on the captured image to detect the line of sight.
[0020] The instruction detection unit 4 detects the joint angles of the operator's fingers and the posture of the wrist. The instruction detection unit 4 is, for example, a data glove, a joint angle detection sensor, or the like.
[0021] The action acquisition unit 5 acquires information about at least the action of the operator operating the end effector 151. The action acquisition unit 5 acquires, for example, the detection result detected by the gaze detection unit 31 and the detection result detected by the instruction detection unit 4.
[0022] The intention understanding unit 6 uses the information acquired by the action acquisition unit 5 to estimate the target object to be operated by the end effector 151 and the task (operation intention) that is the method of operating the target object.
[0023] The operation amount determination unit 7 acquires the information output by the intention understanding unit 6 and the information output by the environmental situation determination unit 2, and determines the operation amount of the end effector 151 from the acquired information.
[0024] The model 8 is a trained model that the motion estimation unit 11 uses when estimating motion.
[0025] The data integration unit 9 generates integrated information of the three-dimensional shape using voxel information included in the shape information about the target object output by the multiple environmental situation assessment units 2. The data integration unit 9 may generate the three-dimensional information by integrating point cloud information included in the shape information about the target object output by the multiple environmental situation assessment units 2. The integrated information has an instance ID, and the data integration unit 9 outputs the integrated information of the three-dimensional shape integrated for each instance ID.
[0026] The approximation result reflecting unit 17 compares the approximation result with the registered information of the operation strategy registered in advance for the end effector 151 and reflects the result.
[0027] The shape approximation unit 10 acquires the integrated information of the integrated three-dimensional shape output by the data integration unit 9. The shape approximation unit 10 generates geometric primitive information by performing a shape fitting process on geometric primitives for each instance ID. As will be described later, the shape approximation unit 10 may perform shape approximation using a single primitive, multiple primitives, or a region of interest. In the case of multiple primitives, geometric information is used as information for classifying shapes. When performing shape approximation using multiple primitives, the shape approximation unit 10 acquires the affordance of each region by referring to the database 14 after fitting each primitive. The shape approximation unit 10 classifies the objects into classes (e.g., plastic bottle, wrench, box, etc.), for example.
[0028] The motion estimation unit 11 acquires operation intention information from the intention understanding unit 6, joint angle measurements from the robot 15, and instance IDs (identification information) and geometric primitives of visible objects (detected by the environmental situation determination unit 2) from the shape approximation unit 10. The motion estimation unit 11 references the model 8 to estimate a motion, which is a set of operation intentions, based on the operation intention, robot state, and possible operations associated with the target object. For example, the motion of switching between two hands for a wrench consists of operation intentions for right-hand grip, two-hand support, and left-hand grip. For example, the motion estimation unit 11 estimates the next motion the robot 15 should perform. The motion estimation unit 11 outputs estimated motion information for the robot 15's hands. The motion information also includes taxonomy information related to the task (see, for example, Reference 1). To reduce the amount of data, the model 8 does not store a wide variety of plastic bottle models, but instead stores a representative model of a plastic bottle. Therefore, the motion estimation unit 11 can estimate the motion by using the class of the target object and the information stored in the model 8.
[0029] Reference 1; Thomas Feix, Javier Romero, et al., “The GRASP Taxonomy of Human GraspTypes” IEEE Transactions on Human-Machine Systems (Volume: 46, Issue: 1, Feb.2016), IEEE, p66-77
[0030] The grasping DB 12 is a database that stores shape models with probabilities, and stores them for each class ID (identifier). Shape refers to the average shape of the class. Probability refers to the grasping probability for each operation intention, and is stored for each polygon, for example. This shape is transformed to fit the geometric primitive shape to obtain the grasping probability for the obtained instance shape, but this may also be approximated using a neural network. The grasping DB 12 stores the general weight (per unit volume) of the object class. The grasping DB 12 stores the grasping plan for the average shape for each taxonomy and class.
[0031] The grasp planning unit 13 acquires geometric primitive shape information for each instance ID from the shape approximation unit 10 and motion information estimated from the motion estimation unit 11. The grasp planning unit 13 determines a grasp point (grasp position) based on the geometric primitive for each instance ID included in the acquired information, the future operation intention and taxonomy obtained by the motion, and the grasp probability for each operation intention. For example, the grasp plan keeps the fingers closed until they come into contact after the grasp point is determined. The grasp planning unit 13 outputs hand trajectory sequence information and finger trajectory sequence information based on the grasp point. The processing performed by the grasp planning unit 13 will be described later.
[0032] The database 14 stores, for example, the ahordances provided by the primitives. In the case of multiple primitives, the database 14 stores the ahordances provided by each primitive. The database 14 stores instance IDs in association with the class IDs to which they belong. The database 14 stores average shape models for each class. For primitives in the region of interest, the database 14 stores positional information within the region that is easy to grasp in association with the average shape.
[0033] The robot 15 is, for example, one of a single-arm robot, a double-arm robot, and a multi-arm robot. The end effector 151 has at least three fingers. The number of fingers that the end effector 151 has may be any number that can realize the taxonomy. The arm 152 has an end effector 151 connected to its tip.
[0034] The control unit 16 generates joint angle command values for the end effector 151 and the arm 152 based on the hand trajectory information and finger trajectory information output by the grasp planning unit 13 and the joint angle measurement values output by the sensor 153 of the robot 15, and controls the operation of the robot 15.
[0035] [Configuration example of environmental situation assessment unit] Next, an example of the configuration of the environmental situation determination unit 2 will be described. 2 is a diagram showing an example of the configuration of the environmental situation determination unit according to this embodiment. As shown in Fig. 2, the environmental situation determination unit 2 includes, for example, an environmental sensor 21, a self-position estimation unit 22, a mask generation unit 23, an addition unit 24, a 3D reconstruction unit 25, and a data integration unit 26. The environment sensor 21 includes, for example, an RGB sensor 211 and a depth sensor 212 .
[0036] The RGB sensor 211 is an imaging device such as a CCD (Charge Coupled Device) imaging device or a CMOS (Complementary Metal Oxide Semiconductor) imaging device, and outputs a RGB (Red, Green, Blue) color image. The RGB sensor 211 also includes, for example, a fisheye lens.
[0037] The depth sensor 212 acquires a depth image including depth information and outputs the acquired depth image. Note that the depth image includes brightness information as seen from the depth sensor 212.
[0038] The adder 24 generates a depth image for each instance ID using the depth image and the mask image, and outputs the generated depth image for each instance ID. The output image contains the depth for each instance ID. This makes it possible to extract the location of areas with the same instance ID, and therefore the distance for each instance ID.
[0039] The mask generation unit 23 uses, for example, an instance segmentation technique to generate and output a mask image based on the instance ID (identification information) of an object in the image using the color image output by the RGB sensor 211. Note that instance segmentation is, for example, a problem of estimating the foreground region mask of an object instance appearing in an image or an RGB-D image while distinguishing each object instance from other objects. The foreground region mask represents "foreground object or background" using, for example, two values [1 or 0].
[0040] The self-position estimation unit 22 performs image processing on the color image output by the RGB sensor 211, detects the position of the environmental situation determination unit 2, and outputs attitude information of the image capturing device.
[0041] The three-dimensional reconstruction unit 25 obtains a point cloud for each instance ID for each class by three-dimensional reconstruction using the posture information of the image capture device output by the self-position estimation unit 22 and the depth image for each instance ID for each class output by the addition unit 24. In other words, the three-dimensional reconstruction unit 25 classifies the object into instance IDs and constructs the shape of the object using, for example, voxels or a point cloud. Note that the point cloud is point cloud data, which is three-dimensional data having information such as basic position information of X, Y, and Z, and color.
[0042] The data integration unit 26 acquires a point cloud for each instance ID for each key frame. The data integration unit 26 also acquires and manages current data at time t and time series data of past frames such as t-1 and t-2, and integrates this data. The data integration unit 26 outputs shape information about the target object, including the current frame and past frames, to the operation amount determination unit 7. For example, if two objects are placed on a table, an instance ID is associated with each of the two classes (first object, second object).
[0043] [Remote control overview] Fig. 3 is a diagram showing an overview of remote control according to this embodiment. Note that, although the robot 15 shown in Fig. 3 is an example having a pair of arms, a head, and a body, the configuration and shape of the robot 15 are not limited to this. As shown in Fig. 3, the operator Us wears, for example, an HMD (head-mounted display) 3 and instruction detection units 4a and 4b. An environmental sensor 21a is installed on the robot 15 or around the robot 15, and an environmental sensor 21b is also installed in the working environment. The environmental sensor 21 may be attached to the robot 15. The robot 15 also includes an end effector 151 (151a, 151b) and an arm 152. The operator Us remotely controls the robot 15 by moving the hand or fingers wearing the instruction detection units 4a and 4b while viewing the image displayed on the HMD 3.
[0044] [Example of geometric primitives] Here, an example of a geometric primitive will be described using an example of a task in which the robot 15 is made to hold a wrench and tighten a screw. 4 is a diagram for explaining an example of a geometric primitive. In the case of a task in which the robot 15 is made to hold a wrench and tighten a screw, the object of the operator's attention is the wrench.
[0045] In the example of image g100, the target object is a single primitive with one frame (g101).
[0046] In the example of image g110, the target object is a multiple primitive with two frames (g111, g112).
[0047] In the image g120, in the case of screw tightening work, the area of interest is within the frame g112. Thus, shape fitting in geometric primitives may be performed in one frame, in two or more frames, or using regions of interest.
[0048] In this embodiment, the target object is approximated, for example, on a plane by geometric primitives. Note that in this embodiment, the geometric primitives are basic shapes in Euclidean geometry, such as a triangle, a square, a rectangle, a trapezoid, a polygon, a circle, and an ellipse.
[0049] [Example of processing procedure] Next, a description will be given of an example of a processing procedure performed by the remote operation assistance system 1. Fig. 5 is a flowchart of an example of a processing procedure performed by the remote operation assistance system 1 according to this embodiment.
[0050] (Step S1) The environmental situation determination unit 2 acquires environmental information of the environment in which the end effector 151 is operated. The environmental situation determination unit 2 outputs shape information on the target object, which is the acquired environmental information, to the operation amount determination unit .
[0051] (Step S2) The data integration unit 9 integrates three-dimensional shapes for each instance ID using voxel information included in the shape information about the target object acquired from the multiple environmental situation determination units 2. The shape approximation unit 10 generates geometric primitive information based on the integrated information of the integrated three-dimensional shapes for each instance ID. The grasp planning unit 13 acquires the geometric primitive information for each instance ID from the shape approximation unit 10.
[0052] (Step S3) The operation amount determination unit 7 acquires information about the movement of the operator who operates the end effector 151 from the HMD 3 and the instruction detection unit 4.
[0053] (Step S4) The intention understanding unit 6 uses the information output by the action acquisition unit 5 to estimate a target object to be operated by the end effector 151 and a task that is a method of operating the target object.
[0054] (Step S5) The motion estimation unit 11 estimates a motion, which is a set of operational intentions, from the operational intentions, the robot state, and possible operations linked to the target object. The grasp planning unit 13 acquires motion information from the motion estimation unit 11.
[0055] (Step S6) The grasp planning unit 13 uses the geometric primitive information and estimated motion information for each acquired instance ID to search the database 14 for the class ID to which the instance ID belongs. For example, if there are instances called Spanner A and Spanner B in the Spanner class, the grasp planning unit 13 searches the database 14 for these class IDs.
[0056] (Step S7) The grasp planning unit 13 acquires an average shape model associated with the class of the class ID from the database 14. For example, if the class is a wrench, the average shape model is a model of the average shape of a wrench. Note that such an average shape model for each class is stored in the database 14.
[0057] (Step S8) The grasp planning unit 13 transforms the average shape into a geometric primitive to obtain a mapping function between regions. The grasp planning unit 13 can obtain a mapping function between the average shape and the primitive shape, for example, by fitting the geometric primitive to shape information about the target object. Note that the target object is not limited to a rigid body, and may be a non-rigid body as long as its deformation can be defined. The grasp planning unit 13 also obtains the general weight (per unit volume) of the object class from the database 14.
[0058] (Step S9) The grasping planning unit 13 acquires a grasping probability map from the grasping DB 12 based on the motion, the taxonomy, and the class ID.
[0059] (Step S10) The grasp planning unit 13 maps the grasp probability map to the region of the geometric primitive based on the obtained mapping function of the correspondence between the regions. This allows obtaining probability heat map-like information according to the geometric primitive shape obtained online.
[0060] (Step S11) The grasp planning unit 13 solves an optimization problem to find a grasp point that satisfies the geometric constraints and has the shortest path with the highest probability, based on the current wrist posture.
[0061] (Step S12) After the gripping point is selected, the grip planning unit 13 calculates a trajectory for approaching the gripping point and outputs it to the robot 15. The grip planning unit 13 adaptively controls the change in weight from the general weight (per unit volume) of the object class.
[0062] In the above-described embodiment, grasping is described as an example of the work, but the work content is not limited to this.
[0063] [Processing for single and multiple primitives] Here, the processing in the case of a single primitive and multiple primitives will be further explained. Fig. 6 is a flowchart of the processing procedure in the case of a single primitive and multiple primitives according to this embodiment.
[0064] (Step S101) The environmental situation determination unit 2 acquires environmental information of the environment in which the end effector 151 is operated. The environmental situation determination unit 2 outputs shape information on the target object, which is the acquired environmental information, to the operation amount determination unit .
[0065] (Step S102) The data integration unit 9 integrates three-dimensional shapes for each instance ID using voxel information included in the shape information about the target object acquired from the multiple environmental situation determination units 2. The shape approximation unit 10 generates geometric primitive information based on the integrated information of the integrated three-dimensional shapes for each instance ID. The grasp planning unit 13 acquires the geometric primitive information for each instance ID from the shape approximation unit 10.
[0066] (Step S103) The operation amount determination unit 7 acquires information about the movement of the operator who operates the end effector 151 from the HMD 3 and the instruction detection unit 4.
[0067] (Step S104) The intention understanding unit 6 uses the information output by the action acquisition unit 5 to estimate a target object to be operated by the end effector and a task that is a method for operating the target object.
[0068] (Step S105) The motion estimation unit 11 estimates a motion, which is a set of operational intentions, from the operational intention, the robot state, and possible operations linked to the target object. The grasp planning unit 13 acquires motion information from the motion estimation unit 11.
[0069] (Step S106) The grasp planning unit 13 uses the acquired geometric primitive information and the estimated motion information to search the database 14 for the class ID to which the instance ID belongs. For example, if there are instances called Spanner A and Spanner B in the Spanner class, the database 14 is searched for these class IDs.
[0070] (Step S107) In the case of multiple primitives, the grasp planning unit 13 determines the selectable taxonomy of each primitive based on the operator's operation input, the class ID, and the primitive. Note that in the case of a single primitive, the grasp planning unit 13 determines the taxonomy based on the operator's operation input.
[0071] (Step S108) The grasp planning unit 13 determines the initial weight of the object from the density corresponding to the class ID and the estimated volume.
[0072] (Step S109) The grasp planning unit 13 solves an optimization problem to select a grasp point that satisfies the geometric constraints and reaches the object surface of the target object via the shortest path, based on the current wrist posture.
[0073] (Step S110) After a gripping point is selected, the grip planning unit 13 calculates a trajectory for approaching the gripping point and outputs the trajectory to the robot 15. The grip planning unit 13 adaptively controls the change in weight from the general weight (per unit volume) of the object class.
[0074] [Processing and configuration overview] Here, an outline of the configuration and processing of the remote operation assistance system 1 will be described. FIG. 7 is a diagram for explaining an outline of the configuration and processing of the remote operation assistance system according to this embodiment. 7, the mask generation unit 23 generates and outputs a mask image based on the instance ID of an object in an image in RGB data. Note that a prerequisite here is that the objects in the image can be classified into classes. The three-dimensional reconstruction unit 25 obtains a point cloud for each instance ID for each class by three-dimensional reconstruction based on the mask image and depth image. The three-dimensional reconstruction unit 25 clusters and outputs voxels or point clouds, which are shape information about the target object. The three-dimensional reconstruction unit 25 performs active sensing using the RGB sensor 211 and depth sensor 212, and integrates the sensed information into three-dimensional points. The shape approximation unit 10 performs primitive fitting using shape information on the clustered target object to approximate the object shape with geometric primitive elements (for example, a rectangular parallelepiped, a cylinder, a sphere, etc.). The grasp planning unit 13 realizes object grasping according to the operation intention by scaling a pre-registered grasp strategy.
[0075] [Comparative Example] Here, a comparative example will be described with reference to Fig. 8. Fig. 8 is a diagram for explaining the comparative example. As shown in Figure 8, conventional technology could only handle objects whose information was registered in advance in a database. For example, even if the object was a wrench, if the shape and weight were different, it had to be registered. Furthermore, conventional technology required that the three-dimensional shape and weight of the target object be known. Furthermore, with congestion technology, the gripping point and gripping posture of the target object also had to be known.
[0076] In contrast, in this embodiment, only the object class (e.g., a wrench) and a grasping strategy corresponding to the operation intention are given in advance as prerequisites. Then, in this embodiment, the object shape for which the grasping strategy corresponding to the object class and operation intention is given in advance is approximated by a geometric primitive (e.g., a rectangular parallelepiped), and the pre-registered grasping strategy is scaled to realize object grasping corresponding to the operation intention. Furthermore, in this embodiment, the weight is adaptively controlled by the amount of change from the general weight (per unit volume) of the object class.
[0077] As a result, according to this embodiment, the amount of information to be registered in advance is reduced, making it possible to perform grasping planning for various instances.
[0078] A program for implementing all or part of the functions of the manipulated variable determination unit 7 of the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the manipulated variable determination unit 7. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a website provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line.
[0079] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0080] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0081] 1...Teleoperation assistance system, 2,2-1,2-2,2-3,......Environmental situation judgment unit, 3...HMD, 4...Instruction detection unit, 5...Movement acquisition unit, 6...Intention understanding unit, 7...Operation amount determination unit, 8...Model, 9...Data integration unit, 10...Shape approximation unit, 11...Movement estimation unit, 12...Grasp DB, 13...Grasp planning unit, 14...Database, 15...Robot, 16...Control unit, 17...Approximation result reflection unit, 31...Gaze detection unit, 151...End effector, 152...Arm, 153...Sensor, 21...Environment sensor, 22...Self-position estimation unit, 23...Mask generation unit, 24...Adder, 25...3D reconstruction unit, 26...Data integration unit, 211...RGB sensor, 212...Depth sensor
Claims
1. A teleoperation assistance system for remotely operating at least an end effector, a motion acquisition unit that acquires motion information relating to a motion of an operator operating at least the end effector; an intention understanding unit that uses the operation information to estimate a target object to be operated by the end effector and a task that is a method of operating the target object intended by the operator; an environmental situation determination unit that acquires environmental information of an environment in which the target object exists and the end effector is operated with respect to the target object; a motion estimation unit that estimates a motion based on the task, a state of the robot equipped with the end effector, and an object shape of the target object generated based on the environmental information; an operation amount determination unit that determines an operation amount of the end effector based on the estimated motion and the acquired environmental information, The operation amount determination unit a shape approximation unit that approximates the object shape of the target object to geometric primitives; a grasp planning unit that generates a grasp plan by comparing the geometric basic elements with registered information of an operation strategy as a strategy for an operation to grasp an object pre-registered for the end effector, The grasping planning unit fitting the object shape of the target object to the geometric primitives to calculate a mapping function between the area of the object shape of the target object and the area of the geometric primitives; determining a taxonomy capable of performing said task; acquiring a grasp probability map indicating a grasp probability for each operation intention based on the motion information, the determined taxonomy, and a class ID to which an instance ID assigned to the object shape of the target object belongs; mapping the grasp probability map to the region of the geometric basic element based on a mapping function of correspondence between the calculated regions; and generating the grasp plan based on a current wrist posture. Remote control assistance system.
2. The shape approximation unit approximates the object shape of the target object to a geometric primitive for each of the instance IDs. The remote operation assistance system according to claim 1 .
3. The geometric primitive may be one or more. The remote operation assistance system according to claim 1 .
4. The geometric primitives include a shape of a region of interest that is an area of the target object that is easy to grasp. The remote operation assistance system according to claim 1 .
5. The environmental situation determination unit an RGB sensor for acquiring RGB data; a depth sensor for acquiring a depth image; a mask generation unit that uses an instance segmentation technique to generate a mask image based on an instance ID of an object in the image using the color image output by the RGB sensor; an adder that generates and outputs a depth image for each instance ID using the depth image and the mask image; a self-position estimation unit that performs image processing on a color image output by the RGB sensor to detect a position and orientation of the RGB sensor, and generates and outputs orientation information based on the position and orientation; a three-dimensional reconstruction unit that obtains a point cloud for each instance ID for each class by three-dimensional reconstruction using the orientation information of the RGB sensor output by the self-position estimation unit and the depth image for each instance ID output by the addition unit; A data integration unit that acquires a point cloud of each instance ID for each key frame, acquires and manages current data at time t, and time series data of past frames t-1, t-2, ..., and integrates these data; 3. The remote operation assistance system according to claim 1, further comprising:
6. a computer of a teleoperation assistance system that teleoperates at least the end effector, acquiring motion information relating to the motion of an operator who operates at least the end effector; Using the motion information, a target object to be operated by the end effector and a task that is a method of operating the target object intended by the operator are estimated; acquiring environmental information of an environment in which the target object exists and in which the end effector is operated relative to the target object; Estimating a motion based on the task, a state of the robot equipped with the end effector, and an object shape of the target object generated by the environmental information; determining an operation amount of the end effector based on the estimated motion and the acquired environmental information; The computer approximating the object shape of the target object to geometric primitives; generating a grasp plan by comparing the geometric primitives with registered information of an operation strategy as a strategy for grasping a pre-registered object for the end effector; The computer fitting the object shape of the target object to the geometric primitives to calculate a mapping function between the area of the object shape of the target object and the area of the geometric primitives; determining a taxonomy capable of performing said task; acquiring a grasp probability map indicating a grasp probability for each operation intention based on the motion information, the determined taxonomy, and a class ID to which an instance ID assigned to the object shape of the target object belongs; mapping the grasp probability map to the region of the geometric basic element based on a mapping function of correspondence between the calculated regions; and generating the grasp plan based on a current wrist posture. Remote control assistance method.
7. A computer of a teleoperation assistance system that teleoperates at least the end effector, acquiring motion information relating to at least the motion of an operator operating the end effector; using the motion information, a target object to be operated by the end effector and a task that is a method of operating the target object intended by the operator are estimated; acquiring environmental information of an environment in which the target object exists and in which the end effector is operated relative to the target object; A motion is estimated based on the task, a state of the robot equipped with the end effector, and an object shape of the target object generated by the environmental information. determining an operation amount of the end effector based on the estimated motion and the acquired environmental information; The computer, approximating the object shape of the target object to geometric primitives; generating a grasp plan by comparing the geometric primitives with registered information of an operation strategy as a strategy for grasping an object previously registered for the end effector; The computer, fitting the object shape of the target object to the geometric primitives to calculate a mapping function between a region of the object shape of the target object and a region of the geometric primitives; determining a taxonomy capable of performing said task; a grasping probability map indicating a grasping probability for each operation intention is obtained based on the motion information, the determined taxonomy, and a class ID to which an instance ID assigned to the object shape of the target object belongs; the grasping probability map is mapped to the region of the geometric basic element based on a mapping function of correspondence between the calculated regions; and the grasping plan is generated based on a current wrist posture. program.
Citation Information
Patent Citations
METHOD FOR FOLLOWING AN OPERATOR'S INTENTIONS TO MOVE A ROBOT SYSTEM
DE102013204789A1
Qualitative inference processing method
JP1989076358A
Operation program composing method and composing device for robot
JP1998128684A
Object handling estimating method and object handling estimating device
JP2004188533A
End effector control method
JP2015136769A