Posture determines method

The method determines wrist posture for unknown objects by measuring and calculating object centers, using a pose scene graph to address the limitations of conventional end effector control, enabling stable robotic grasping.

JP7798613B2Active Publication Date: 2026-01-14HONDA MOTOR CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2022035673
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2026-01-14
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

Conventional end effector control methods fail to determine the wrist posture for grasping unknown objects due to the lack of orientation information in offline databases.

Method used

An attitude determination method that includes measuring the target object using a sensor, calculating multiple points indicating the center of the object, determining a grasp representative point, and using a pose scene graph to determine the attitude of the moving mechanism based on the object's posture and its relationship with surrounding objects.

Benefits of technology

Enables online determination of the wrist posture for grasping unknown objects, allowing stable grasping by considering the geometric structure, taxonomy, and environmental context, thereby improving the efficiency of robotic grasping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798613000001
    Figure 0007798613000001
  • Figure 0007798613000002
    Figure 0007798613000002
  • Figure 0007798613000003
    Figure 0007798613000003
Patent Text Reader

Abstract

To provide a method for determining a posture that can determine a posture online.SOLUTION: A method for determining a posture includes: a measurement step of measuring an object matter with a sensor; a calculation step of calculating a plurality of points indicating a center of a measured object matter; a gripping representative point determination step of defining one of the points as a gripping representative point in accordance with one of the selected operation shapes; and a posture determination step of determining a posture of a movement mechanism from the gripping representative point.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an attitude determination method. [Background technology]

[0002] In conventional end effector control, the wrist posture determined from the fingertip contact position that allows stable grasping is stored in a database using polygon information (known object) obtained offline. Additionally, online, the wrist positions stored offline in the database are used through conditional search. A device for manipulating the posture of a robot has also been proposed (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6476358 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in conventional technology, when an end effector is remotely operated to grasp an unknown object, the wrist posture data for known objects stored offline in a database does not contain information such as the orientation in which the unknown object can be grasped, making it impossible to determine the posture of the end effector, etc.

[0005] The present invention has been made in consideration of the above problems, and has as its object to provide an attitude determination method capable of determining attitude online. [Means for solving the problem]

[0006] (1) In order to achieve the above-mentioned object, one aspect of the present invention provides an attitude determination method for a moving mechanism that is capable of grasping and operating a target object, is connected to an end effector having a plurality of operation attitudes, and is capable of moving the position of the end effector, and includes a measurement step of measuring the target object using a sensor, a calculation step of calculating a plurality of points that indicate the center of the measured target object, a grasp representative point determination step of determining one of the plurality of points as a grasp representative point corresponding to one of a plurality of selected operation shapes, and an attitude determination step of determining the attitude of the moving mechanism from the grasp representative point.

[0007] (2) Furthermore, in the posture determination method according to one aspect of the present invention, the posture determination process may determine the posture of the moving mechanism based on a taxonomy regarding the grasping of the grasped object, based on the postures of the target object and the moving mechanism.

[0008] (3) In the posture determination method according to one aspect of the present invention, the posture determination step may determine the grasping representative point using a model trained in advance using a pose scene graph in which a relationship between the end effector and the grasped object, a relationship between the target object and objects around the target object, and a posture of the target object are represented by nodes and edges.

[0009] (4) In the posture determination method according to one aspect of the present invention, the nodes of the pose scene graph may be based on coordinates of the target object as seen from the end effector, and the edges of the pose scene graph may be costs based on a relationship between the end effector and the grasped object, a relationship between the target object and objects surrounding the target object, and the posture of the target object.

[0010] (5) In addition, in the posture determination method according to one aspect of the present invention, the posture determination step may determine the grip representative point using a model that is pre-trained using a cost scene graph that represents the relative pose of the end effector and the target object, the target object and a plurality of points indicating the center of the target object, and the postures that the moving mechanism can take, using nodes and edges.

[0011] (6) In addition, in the posture determination method according to one aspect of the present invention, the cost scene graph may have, as costs on edges, information based on a relative pose of a grip center with respect to a plurality of points indicating the center of the target object, information based on a difference between a curved surface of the target object around the plurality of points indicating the center of the target object and a finger curved surface of a finger portion of the end effector, and a value of a Quality Measure. [Effects of the Invention]

[0012] According to (1) to (6), the posture can be determined online. [Brief explanation of the drawings]

[0013] [Figure 1] 10A and 10B are diagrams for explaining an example of a method for determining a point indicating a grip center according to an embodiment. [Figure 2] 10A and 10B are diagrams illustrating examples of center lines when the size of an object to be grasped changes. [Figure 3] 10A and 10B are diagrams illustrating examples of centerlines for three-dimensional writing styles. [Figure 4] FIG. 10 is a diagram for explaining selection of a grip representative point. [Figure 5] FIG. 10 is a diagram for explaining determination of a grasp representative point in grasping an object. [Figure 6] FIG. 10 is a diagram showing examples of names in a taxonomy. [Figure 7] This is an image diagram of the context of the relationship between objects and the relationship between objects and actions. [Figure 8] FIG. 1 is a diagram for explaining a scene. [Figure 9] FIG. 10 is a diagram for explaining a graph. [Figure 10] 10A and 10B are diagrams for explaining learning of a hand direction and an object posture according to an embodiment. [Figure 11] FIG. 10 is a diagram for explaining an example of use of a learning result according to the embodiment. [Figure 12] 1 is a diagram showing a scene of the relationship between a hand and an object and a simple graph representation thereof; [Figure 13] FIG. 10 is a diagram illustrating an overview of a wrist posture estimation process. [Figure 14] FIG. 10 is a diagram for explaining a cost scene graph. [Figure 15] 1 is a diagram illustrating an example of the configuration of an end effector that performs work according to an embodiment. FIG. [Figure 16] FIG. 2 is a diagram illustrating an example of the configuration of a control device according to the embodiment. [Figure 17] 10 is a flowchart of a processing procedure performed by a wrist posture determination device during learning according to an embodiment. [Figure 18] 10 is a flowchart of a processing procedure performed by a wrist posture determination device during work according to an embodiment. [Figure 19] FIG. 10 is a diagram for explaining a method for determining a wrist posture when no model is used. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).

[0015] <Summary> In the posture determination method of the embodiment, the wrist posture, which can be processed online, is determined based on the geometric structure of the object (mainly skeleton information of unknown polygons), the taxonomy grasp center, the fingertip curvature during grasp, the quality measure around the fingertip contact point (minimum sum of squares of lateral forces (friction forces)), and a scene graph of the hand, object, and surrounding environment. In addition, in the posture determination method of the embodiment, the relative pose between the hand and the grasped object, the relative pose between the grasped object and the surrounding environment, and the relative pose differences between the grasped object and the floor are defined as a pose graph. The posture determination method of the embodiment defines the relative pose of the grasp center relative to the skeleton, the convolution of the difference between the surface curves of the objects around the skeleton and the finger curves, and the value of the Quality Measure as a cost. Furthermore, in the posture determination method of the embodiment, the wrist posture is determined by convolving the edges of the two scene graphs around a node to find the node with the lowest cost.

[0016] <Example of how to determine the point indicating the grip center> FIG. 1 is a diagram for explaining an example of a method for determining a point indicating a grip center according to this embodiment. First, a black-and-white image g101 of the object to be grasped viewed from above is created to represent it as a polygon, as in image g100. Voronoi processing is then used to determine the center of the polygon using line g102, a set of nearby points. An axis g103 is then defined, either perpendicular to line g102, a set of points indicating the center of the object, or perpendicular to the outer surface of the object to be grasped. These are used to calculate frictional forces, etc., when pinching the long and short axes of the two fingers.

[0017] Image g110 is an example of a 2D or 3D copy of the centerline generated in this way. Image g120 is an image (polygon) of a rock, which is an object to be grasped g121, viewed obliquely from the side. A set of multiple adjacent points g122 is the target grasp center, and point g123 is the grasp center of the end effector. In this manner, the multiple adjacent points g122 that represent the center of the object are also referred to as a center line in this embodiment.

[0018] The rule for creation is that if the center line is short (e.g., 5 samples or less), the object-centered frame is selected as the target grasp pose. Also, above the target grasp pose, only the quaternion (0,0,0,1) is used, and the finger direction is the same as the x-axis of the body frame.

[0019] FIG. 2 is a diagram showing an example of the center line when the size of the object to be grasped changes. As shown in Figure 2, the smaller the target object, the shorter the length of the center line. Also, if the target object is a sphere or circle, the center line will have one point because multiple points converge at the center. Note that the center line, which is made up of multiple adjacent points, also contains information about the size, longitudinal direction, and orientation of the object. In FIG. 2, reference symbol g11 denotes the target grip center of the target object, and reference symbol g12 denotes the grip center of the end effector.

[0020] FIG. 3 shows an example of a center line for a three-dimensional writing style. To detect a center line, as described above, the image is converted to black and white and a center line is created. Image g201 is an example of a detected center line for a torus-shaped three-dimensional object. Reference symbol g202 represents a three-dimensional object, line g203 represents a center line candidate, and line g204 represents example information on a center line obtained by averaging line g202. Image g205 is an example of a detected center line for a screw-shaped three-dimensional object. Reference symbol g206 represents a three-dimensional object, line g207 represents a center line candidate, and line g208 represents example information on a center line obtained by averaging line g207. Center lines can be generated using well-known techniques for objects wrapped around thin strings, doll-like objects, animals, and even non-continuous objects.

[0021] <Determination of gripping representative point> As mentioned above, the center line is made up of multiple points. When grasping, it is necessary to select one point (the grasp representative point) from these multiple points. The selection of the grasp representative point is explained below. Figure 4 is a diagram for explaining the selection of the grasp representative point. When a trained model is not used, a point is selected from multiple points that is near the center of gravity of the object model, in the direction in which the end effector is approaching, and that is likely to be able to grasp stably without tilting when grasped.

[0022] However, in reality, it may be difficult to determine the grip representative point using only such a center line. Fig. 5 is a diagram for explaining how to determine a grasp representative point when grasping an object, and Fig. 6 is a diagram showing examples of names in the taxonomy.

[0023] Image g221 shows an example of a medium-wrap grip in the taxonomy (see Reference 1) that grips the side of a bottle from the side. Image g222 is an example of a bottle being held in a holder or similar, and gripped from the side of the upper part of the bottle's side using a medium-wrap in the taxonomy. Image g223 is an example of pinching the side of a bottle from above using Precision-3finger in the taxonomy when the bottle is lying down. Image g224 is an example of a Tripod in the taxonomy grasping a bottle from above. Image g225 is an example of a medium-wrap in the taxonomy supporting a bottle from below.

[0024] In images g221 to g225, the areas indicated by symbols g231 to g235 are areas that are easy to grasp. This area cannot be determined solely by the area surrounding the center; it is necessary to consider the state of the target object and the direction from which the end effector is approaching. For this reason, grasping requires a function that can understand the context and change the way of holding.

[0025] Reference 1; Thomas Feix, Javier Romero, et al., “The GRASP Taxonomy of Human Grasp Types” IEEE Transactions on Human-Machine Systems (Volume: 46, Issue: 1, Feb. 2016), IEEE, p66-77

[0026] As explained using Figure 5, the grasping scenes are holding a bottle, holding a bottle in a holder, holding a bottle on the ground, pinching the bottle cap, supporting the bottom of the bottle, and so on. These common elements show that it is necessary to understand two contexts: the relationship between objects and the relationship between objects and actions.

[0027] Figure 7 shows an image of the context of the relationship between objects and the relationship between objects and actions. Image g251 is an image of the context of the relationship between objects. Image g252 is an image of the context of the relationship between objects and actions. As in image g251, there are also relationships between objects, such as when an object is sandwiched between something else or placed on a desk or floor. Such relationships (states) are meanings, and meanings are added as costs to the relationships between objects. For example, in image g252, when picking up an object, there is an object and the action of picking up has meaning. The meaning is added as a cost to the relationship between the object and the action.

[0028] <Scene Graph> Next, the scene graph will be described. FIG. 8 is a diagram for explaining a scene. In this embodiment, a scene graph is used to represent the environment surrounding an object. For example, as shown in FIG. 8, a scene represents a state in which a bottle, a glass, and a bowl are placed on a desk. This scene can be represented graphically as shown in image g271 in FIG. 9. The bottle has a lid. FIG. 9 is a diagram for explaining a graph. The graph represents the relationship between objects. For example, when viewed from the bottle, the lid is on top, and when viewed from the lid, the bottle is on the bottom. Furthermore, when viewed from the bottle, the glass is on the right, and when viewed from the glass, the bottle is on the left. The bottle, glass, and bowl are on the desk.

[0029] The relationship between objects in image g271 in Figure 9 is represented by nodes, edges, and nodes, as in image g272 in Figure 9, and these edges correspond to the feature value, which is the cost. Note that the words that express the relationship are expressed numerically as labels. In this way, in a scene graph, the relationship between objects can be represented graphically using nodes and labels.

[0030] <Learning the end effector and object pose> FIG. 10 is a diagram illustrating the learning of hand direction and object posture according to this embodiment. As shown in image g301, this embodiment learns the orientation of the hand relative to the object in order to perform online processing on an unknown object. Furthermore, for example, if a glass is overturned, the meaning of "up" may not be clear from the object's orientation, such as whether the top of the glass is up or whether the side of the glass is up. For this reason, this embodiment learns the direction and object posture separately. Therefore, in this embodiment, the relative poses of the hand relative to the object and the object posture are learned in advance. Note that nodes are determined by captured images and sensor values. Edges are generated using a trained model. The graph representation is for the object as seen from the hand.

[0031] Image g302 is an image of learning the direction of the hand from the perspective of the target object (mass system). Note that the feature quantities of the edge are, for example, the distance between nodes, the angle between nodes, and the position vector. Image g303 is an image diagram of object posture learning. The object posture may be, for example, an object standing, an object not touching the floor, an object lying down (also referred to as "lying down" in the embodiment), or an object lying down and not touching the ground. In this case, the edge feature amount may be, for example, the object quaternion, the contact point and normal in the object coordinate system, gravitational acceleration, etc.

[0032] In this embodiment, the results of this learning are used as shown in FIG. FIG. 11 is a diagram for explaining an example of using the learning results according to this embodiment. As a result of learning using a scene graph, the mathematical model can be used as follows. The relationship between the object and the hand is, for example, "the direction is that the hand is on top of or beside the object" and "the posture is that the object is standing." In this case, the grasping information is to place a capsule, which is a grasping candidate area, on the lid or side. Also, the relationship between the object and the hand is, for example, "the direction is that the hand is on top of or beside the object" and "the posture is that the object is lying down." In this case, the grasping information is to place a capsule on the side.

[0033] <Scene graph of hand-object relationships> Next, the relationship between the hand and the object can be expressed in a scene graph as shown in Figure 12. Figure 12 is a diagram showing a scene of the relationship between the hand and the object and a simple graph representation. Note that the simple graph representation in Figure 12 shows a simplified version of the link arrows between the nodes in Figure 9. In this embodiment, this type of scene graph is called a pose scene graph. Note that the configuration of the pose graph is such that the nodes are, for example, the coordinates of the object as seen from the hand, and the edges are the coordinates obtained by adding the relative pose between the hand and the object to be grasped, the relative pose between the object to be grasped and the surrounding environment, and the relative pose between the object to be grasped and the floor.

[0034] <Wrist posture estimation processing> Next, an overview of the wrist posture estimation process will be described. Fig. 13 is a diagram showing an overview of the wrist posture estimation process. As shown in Fig. 13, in this embodiment, the above-mentioned pose scene graph and cost scene graph are used to estimate the wrist posture.

[0035] Now, let us explain the cost scene graph. FIG. 14 is a diagram for explaining the cost scene graph. Image g351 shows a state in which a hand is on an object (bowl). Image g352 shows a state in which a hand is on the side of an object (bottle). Image g353 is a graph representation. In such cases, the relationship between the hand and the object requires other information (hand position, finger angle, hand pose, friction, etc.). As a result, multiple states may be possible for a single node, as shown in image g354 of image g253. In this embodiment, a cost is assigned for each of these multiple states. The cost is calculated using the value of a quality measure (minimum sum of squares of lateral forces) around the fingertip contact point. The cost graph is constructed by multiplying the points that make up the center line by the wrist pose, and the edges are coordinates obtained by adding the relative pose of the grasp center with respect to the skeleton, the convolution of the difference between the object surface curve around the skeleton and the finger curve, and the value of the quality measure. The nodes are, for example, positions that can be grasped by the hand, positions that can be gripped, etc. The number of nodes is the number of points that make up the center line multiplied by the number of possible wrist poses. The reason for this is that even if the hand and target are positioned on an object and the gripping point is fixed, the way the object is gripped changes depending on the wrist posture.

[0036] In this embodiment, the pose graph is used to create the relative pose between the hand and the object to be grasped, or the relative pose between the object to be grasped and the surrounding environment, or the relative pose between the object to be grasped and the floor surface. Next, in this embodiment, a cost graph is used to create a scene graph in which the relative pose of the grasp center with respect to the skeleton, the convolution of the difference between the object (or surface, curved surface) around the skeleton and the finger curved surface (whether or not the finger is near the object, and if so, which point among the multiple points that make up the center line should be used as the grasp representative point), and the value of the Quality Measure are assigned as costs to the edges. Furthermore, in this embodiment, the wrist posture is determined by folding the edge costs and finding the node that can have the minimum cost. In addition, in this embodiment, training data, which is the result of actually operating the end effector and gripping a target object, is stored in a database.

[0037] In this embodiment, a network layer (model) is provided that uses a pose scene graph and a cost scene graph to determine the grip style online, and includes attention nodes that are associated with the correct answer data, using accumulated training data as correct answer data.

[0038] During operation, the actual current wrist posture can be input into this network (a trained model) to estimate the wrist posture during grasping.

[0039] <Example of end effector configuration> Fig. 15 is a diagram showing an example of the configuration of an end effector that performs work according to this embodiment. As shown in Fig. 15, end effector 1 (hand) includes finger portion 101, finger portion 102, finger portion 103, finger portion 104, and base 111. End effector 1 is connected via a joint to arm 121, which is a movement mechanism that can move the position of end effector 1.

[0040] Furthermore, finger portion 101 is provided with a force sensor 141, for example, at its fingertip. Finger portion 102 is provided with a force sensor 142, for example, at its fingertip. Finger portion 103 is provided with a force sensor 143, for example, at its fingertip. Finger portion 104 is provided with a force sensor 144, for example, at its fingertip. Finger portion 101 corresponds to, for example, a human thumb, finger portion 102 corresponds to, for example, a human index finger, finger portion 103 corresponds to, for example, a human middle finger, and finger portion 104 corresponds to, for example, a human ring finger. The end effector 1 has at least two fingers, but the number of fingers may be three or more.

[0041] The arm 121 is also a mechanism that can change the wrist position of the end effector 1 . 16, the end effector 1 and the arm 121 are controlled by a control device 200. Fig. 16 is a diagram showing an example of the configuration of the control device according to this embodiment.

[0042] The control device 2 is connected to an imaging device 5 (sensor), an instruction device 7, an end effector 1, and an arm 121 by wire or wirelessly. The end effector 1 includes an actuator 161 and a sensor 151 in addition to the fingers 101 to 104 (FIG. 15) and the base body 111 (FIG. 15). The arm 121 includes an actuator 171 and a sensor 181 .

[0043] The instruction device 7 includes, for example, a sensor 71 and a communication unit 72.

[0044] The control device 2 includes an acquisition unit 21, a control unit 22, an end effector driving unit 23, an arm driving unit 24, and a wrist posture determination device 3. The wrist posture determination device 3 includes a center line creation unit 31, a grasp representative determination unit 32, a learning unit 33, a memory unit 34, a determination unit 37, and a taxonomy determination unit 38. The memory unit 33 stores a model 35 and training data 36.

[0045] The instruction device 7 is a data glove that is worn by the worker on the hand. The instruction device 7 may include, for example, an HMD (head mounted display) that detects the line of sight. The sensor 71 is, for example, a six-axis sensor, a gyro sensor, a pressure sensor, etc. The sensor 71 detects at least the positions of the fingers and the wrist and their trajectories. The communication unit 72 transmits the sensor value detected by the sensor 71 to the control device 2.

[0046] The image capturing device 5 is an RGB-D camera that is also capable of depth measurement. The image capturing device 5 determines the position of the target object using the captured image. The image capturing device 5 also determines the position of the environment, such as the floor or ground, where the target object is located, using the captured image.

[0047] The actuators 161 are provided, for example, on each of the fingers 101 to 103 and each of the joints of the end effector. The sensor 151 is, for example, a six-axis sensor, a position sensor, a pressure sensor, or the like.

[0048] The actuator 171 is provided at a connection portion with the end effector 1, for example. The sensor 181 is, for example, a six-axis sensor, a position sensor, a pressure sensor, or the like.

[0049] The acquisition unit 21 acquires an image captured by the imaging device 5. The acquisition unit 21 acquires a sensor value detected by the sensor 151. The acquisition unit 21 acquires a sensor value detected by the sensor 181. The acquisition unit 21 acquires the sensor value from the instruction device 7.

[0050] The control unit 22 generates an arm drive command using the information acquired by the acquisition unit 21 and the information determined by the wrist posture determination device 3. The control unit 22 generates an end effector drive command using the information acquired by the acquisition unit 21.

[0051] The end effector driving unit 23 controls the end effector 1 based on the end effector driving command generated by the control unit 22.

[0052] The arm driving unit 24 controls the arm 121 based on the arm driving command generated by the control unit 22 .

[0053] Next, the wrist posture determination device 3 will be described. The center line creating unit 31 performs image processing on the captured image to create a center line that indicates the center and includes at least one point.

[0054] The grip representative determination unit 32 determines the grip representative point using a trained model 35.

[0055] During learning, the learning unit 33 generates a pose scene graph and a cost scene graph as described above. During learning, the learning unit 33 learns a model 35 using training data 36. The learning unit 33 stores the learned model 35 in the storage unit 34.

[0056] During work, the determination unit 37 acquires the actual wrist posture from the sensor 181. The determination unit 37 inputs the acquired actual wrist posture into the trained model 35 to determine the wrist posture online.

[0057] The taxonomy determination unit 38 estimates the position, posture, etc. of the target object based on the captured image. The taxonomy determination unit 38 determines the taxonomy based on the estimation result.

[0058] <Example of processing procedure> Next, an example of the processing procedure performed by the wrist posture determination device 3 during learning will be described. Fig. 17 is a flowchart of the processing procedure performed by the wrist posture determination device 3 during learning according to this embodiment.

[0059] (Step S101) The image capturing device 5 captures an image of a target object.

[0060] (Step S102) The image capturing device 5 determines the position of the target object and the position of the environment.

[0061] (Step S103) The learning unit 33 creates a pose scene graph using, for example, the coordinates of the object as seen from the hand as nodes, and the coordinates of "the relative pose between the hand and the object to be grasped, the relative pose between the object to be grasped and the surrounding environment, and the relative pose between the object to be grasped and the floor" as edges.

[0062] (Step S104) The learning unit 33 creates a cost scene graph using the points that make up the center line multiplied by the wrist pose as nodes, and the coordinates of the edges as "the relative pose of the grip center with respect to the skeleton, the convolution of the difference between the surface curve of the object around the skeleton and the finger curve, and the value of the Quality Measure."

[0063] (Step S105) The learning unit 33 uses the training data 36, ​​the pose scene graph, and the cost scene graph to learn the model 35, and stores the learned model 35 in the storage unit 34.

[0064] Next, an example of the processing procedure performed by the online wrist posture determination device 3 during work will be described. Fig. 18 is a flowchart of the processing procedure performed by the wrist posture determination device during work according to this embodiment.

[0065] (Step S201) The image capturing device 5 captures an image of a target object.

[0066] (Step S202) The taxonomy determination unit 38 estimates the position, posture, etc. of the target object based on the captured image. The taxonomy determination unit 38 determines the taxonomy based on the estimation result.

[0067] (Step S203) The center line creation unit 31 performs image processing on the captured image to create a center line that indicates the center and includes at least one point. Note that the wrist posture determination device 3 performs the processes of steps S201 to S203 at the start of grasping.

[0068] (Step S204) During work, the actual wrist posture is acquired from the sensor 181.

[0069] (Step S205) The grip representative point determination unit 32 inputs the actual current wrist posture into the trained model 35 to determine the grip representative point for the unknown object online.

[0070] (Step S205) The determination unit 37 determines the wrist posture of the hand with respect to the unknown object so that the target object is grasped with the determined grasp representative point as the grasp center. The determination unit 37 determines the wrist posture online, for example, by inputting the scene graph, cost scene graph, and actual current wrist posture into the trained model 35. The determination unit 37 determines the wrist posture by folding the edges of the two scene graphs around a node and finding the node that can have the minimum cost.

[0071] As described above, in this embodiment, a wrist posture that can be processed online is determined online by inputting the scene graph, cost scene graph, and actual current wrist posture into the trained model 35. When the model 35 is used, the grasp representative point is determined by the model 35, and the optimal wrist posture is determined using the model 35 from among the postures that the wrist of the end effector can take as it approaches the target object.

[0072] In this embodiment, the pose is determined based on the geometric structure of the object (mainly skeleton information of unknown polygons), the grasp center of the taxonomy, the fingertip curvature during grasp, a quality measure (minimum sum of squares of lateral forces) around the fingertip contact point, and a scene graph of the hand, object, and surrounding environment. Also, in this embodiment, a pose graph is created based on the relative pose between the hand and the grasped object, the relative pose between the grasped object and the surrounding environment, and the relative pose between the grasped object and the floor.

[0073] In this embodiment, a cost scene graph is created in which the relative pose of the grasp center (grasp representative point) with respect to the skeleton, the convolution of the difference between the surface curve of the object around the skeleton and the finger curve, and the value of the Quality Measure are used as the cost. The relative pose of the grasped object and the surrounding environment is calculated by determining a provisional wrist posture and calculating the difference in distance between the finger and the provisional grasp center (grasp representative point) based on the provisional grasp center between the fingers. The wrist posture determination device 3 calculates the cost for each of multiple points on the center line. The wrist posture determination device 3 also convolves multiple differences between the curved surface of the object surface around the center line of the object and the curved surface of the fingers of the end effector 1 to generate a single numerical value so that friction occurs. When the grasp center is near the center line, it is necessary to consider the balance of the total lateral friction force at the contact point. Therefore, the wrist posture determination device 3 calculates this as the Quality Measure value and uses it as the cost. In this embodiment, the wrist posture is determined by folding the edges of the two scene graphs around a node to find a node that can provide the minimum cost.

[0074] As a result, according to this embodiment, the wrist posture can be determined online. Also, according to this embodiment, the most suitable condition can be selected from a plurality of constraint conditions, and the method of moving the wrist of the end effector 1 can be determined. As a result, according to this embodiment, the processing can be easily converged.

[0075] Although the above description has been given of an example of the model 35 used to determine the grasp representative point and wrist posture, which uses both a pose scene graph and a cost scene graph, the present invention is not limited to this. Depending on the shape of the target object, the posture of the target object, the current posture of the wrist, etc., the model 35 may be only a pose scene graph.

[0076] <Attitude determination method without using a model> In the above example, the wrist posture is determined using the model 35, but this is not limiting. Fig. 19 is a diagram for explaining a method of determining the wrist posture without using a model.

[0077] 19, image g300 is an example in which the end effector 1 and the wrist are positioned above the target object obj, and then the end effector 1 is moved closer to the target object obj. Image g410 is an example in which the end effector 1 and the wrist are positioned diagonally to the upper right of the target object obj, and then the end effector 1 is moved closer to the target object obj. When the model 35 is not used, the wrist posture determination device 3 selects one of the multiple points constituting the center line from the multiple selected operation shapes (possible postures of the wrist of the end effector approaching the target object) according to the taxonomy and determines it as the grasp representative point. Then, the wrist posture is tilted within an angle range in which the end effector 1 does not hit the floor surface on which the target object obj is placed. Alternatively, the range in which the grasp center can be tilted is determined. In FIG. 19, symbols g401 and g4111 indicate the grasp centers of the target object. Symbols g402 and g412 indicate the grasp center of the end effector 1. The grasp center of the end effector 1 is calculated based on the sensor value detected by the sensor 151 provided in the end effector 1.

[0078] Note that a program for realizing some or all of the functions of the control device 200 and wrist posture determination device 3 of the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the control device 200 and wrist posture determination device 3. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also refers to devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line.

[0079] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0080] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0081] 1...end effector, 101, 102, 103, 104...finger portion, 111...base body, 121...arm, 5...imaging device, 7...instruction device, 151...sensor, 161...actuator, 171...actuator, 181...sensor, 71...sensor, 72...communication unit, 21...acquisition unit, 22...control unit, 23...end effector drive unit, 24...arm drive unit, 3...wrist posture determination device, 31...center line creation unit, 32...grasp representative determination unit, 33...learning unit, 34...storage unit, 37...determination unit, 38...taxonomy determination unit, 35...model, 36...training data

Claims

1. A method for determining the attitude of a moving mechanism that is capable of grasping and manipulating a target object, is connected to an end effector having a plurality of operating attitudes, and is capable of moving the position of the end effector, comprising: a measuring step of measuring the target object by a sensor; a calculation step of calculating a plurality of points indicating the center of the measured object; a grip representative point determination step of determining one of the plurality of points as a grip representative point corresponding to one of the plurality of selected operation shapes; and a posture determination step of determining a posture of the moving mechanism from the grip representative point, The posture determining step determines the grip representative point using a model that has been trained in advance using a pose scene graph that represents the relationship between the end effector and the target object, the relationship between the target object and objects in the periphery of the target object, and the posture of the target object using nodes and edges. Posture determination method.

2. the attitude determination step determines a taxonomy related to grasping the target object based on the position of the target object and the attitude of the moving mechanism, and determines the attitude of the moving mechanism based on the taxonomy.

2. The method of claim 1.

3. The nodes of the pose scene graph are based on the coordinates of the target object as seen from the end effector, The edges of the pose scene graph are costs based on the relationship between the end effector and the target object, the relationship between the target object and objects in the vicinity of the target object, and the posture of the target object.

2. The method of claim 1.

4. The posture determination step determines the grip representative point using a model that has been trained in advance using a cost scene graph that represents the relative pose between the end effector and the target object, the target object and a plurality of points indicating the center of the target object, and possible postures of the moving mechanism using nodes and edges.

2. The method of claim 1.

5. The cost scene graph has, as costs on edges, information based on a relative pose of a grip center with respect to a plurality of points indicating the center of the target object, information based on a difference between a curved surface of the target object around the plurality of points indicating the center of the target object and a curved surface of a finger unit of the end effector, and a Quality Measure value.

5. The method of claim 4.

Citation Information

Patent Citations

  • Qualitative inference processing method

    JP1989076358A

  • Holding attitude generation device, holding attitude generation method and holding attitude generation program

    JP2013182554A

  • Autonomous moving body, object information acquisition device, and object information acquisition method

    JP2014106597A

  • Robot, robot system, control device, and control method

    JP2015145055A

  • Handling device, control device and program

    JP2020032523A