End effector control method
The end effector control method improves grasping accuracy by using sensor-acquired visual information and feedback processes to minimize misalignment, addressing the issue of object misalignment in conventional robot grasping systems.
Patent Information
- Application Number
- JP2022035670
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Conventional methods for robot end effector control fail to accurately grasp small objects due to misalignment between the hand and the target object, leading to unsuccessful grasping.
An end effector control method that utilizes sensors to acquire visual information, determine grasping shape and position, calculate virtual positions, and feedback deviation amounts to improve alignment, using known or unknown object models and minimizing Euclidean distances between image features.
Reduces misalignment between the hand and the target object, enhancing the accuracy of grasping small objects by using virtual visual information and feedback processes.
Smart Images

Figure 0007798612000001 
Figure 0007798612000002 
Figure 0007798612000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an end effector control method. [Background technology]
[0002] When a robot with an end effector is made to perform a task, object estimation is performed only on real images when position estimation is performed using multiple cameras external to the robot or attached to the robot. For example, a control method has been proposed in which the state of a target object is estimated from images taken by a camera attached to the robot's head, and the wrist position determined from the estimated state of the target object is moved as a target (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-66632 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with conventional technology, when grasping a small object, if the state of the object as seen by the hand deviates even slightly from the target state, the fingertips of the hand cannot grasp the target object.
[0005] The present invention has been made in consideration of the above-mentioned problems, and an object of the present invention is to provide an end effector control method that can reduce the misalignment between the hand and the target object. [Means for solving the problem]
[0006] (1) In order to achieve the above-mentioned object, an end effector control method according to one embodiment of the present invention is an end effector control method capable of grasping and operating a target object, and includes an acquisition process for acquiring visual information of the target object using a sensor; a determination process for determining a grasping shape and grasping position of the end effector from the acquired visual information; a virtual position determination process for determining a virtual position from the current position of the end effector to the grasping position; a virtual visual information calculation process for calculating virtual visual information of the sensor at the virtual position based on the acquired visual information and the virtual position; a deviation amount calculation process for comparing the visual information of the sensor at the virtual position with the virtual visual information to calculate a deviation amount; and a feedback process for feeding back the deviation amount to the control amount of the end effector.
[0007] (2) In addition, in an end effector control method according to one aspect of the present invention, the determination step may determine the gripping shape and gripping position using information on a set of points indicating the center of the target object of a known model if the target object is a known object, and may determine the gripping shape and gripping position using information on a set of points indicating the center of the target object or a skeleton line of an unknown object if the target object is an unknown object.
[0008] (3) In addition, in the end effector control method according to one aspect of the present invention, the deviation amount calculation step may include calculating corner features of the target object included in the virtual visual information to calculate virtual features, finding an object in the real image that is highly correlated with the corner features of the target object included in the virtual visual information, and estimating the deviation amount so that the Euclidean distance between the pair of corner features of the two images is minimized.
[0009] (4) In the end effector control method according to one aspect of the present invention, the virtual positions may be at least two for each interval based on the resolution of the visual information.
[0010] (5) In addition, in the end effector control method according to one aspect of the present invention, the virtual position determination step may determine the virtual position using the visual information acquired at an initial position of the end effector that is away from a grasping position.
[0011] (6) In addition, in the end effector control method according to one aspect of the present invention, the determination step may determine a taxonomy based on the visual information, and determine the gripping shape and gripping position of the end effector based on the determined taxonomy. [Effects of the Invention]
[0012] According to (1) to (6), the misalignment between the hand and the target object can be reduced. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 10 is a diagram for explaining displacement caused by a head camera. [Figure 2] 1A and 1B are diagrams illustrating an example of a robot according to an embodiment and an example of a location where a camera is placed on the ground. [Figure 3] 1A and 1B are diagrams illustrating an example of the configuration of an end effector according to an embodiment. [Figure 4] 10A and 10B are diagrams for explaining a predetermined distance when the hand according to the embodiment is brought closer to the target object. [Figure 5] FIG. 10 is a diagram for explaining an image estimation process. [Figure 6] FIG. 10 is a diagram illustrating an example of the configuration of a network used for recognition error estimation. [Figure 7] 1A and 1B are diagrams for explaining a virtual position and a virtual viewpoint image according to an embodiment. [Figure 8] 10A and 10B are conceptual diagrams illustrating an operation when gripping a bolt according to an embodiment. [Figure 9] FIG. 1 is a diagram illustrating an example of the configuration of an end effector control device according to an embodiment. [Figure 10] 10 is a flowchart of a processing procedure of an end effector control device according to an embodiment. [Figure 11] FIG. 10 is a diagram showing examples of names in a taxonomy. [Figure 12] FIG. 10 is a diagram for explaining a set of points indicating the center of a target object according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).
[0015] <Summary> In this embodiment, the pose of the real object is recognized while moving the arm during grasping, using images captured by cameras equipped on the robot's head and the end effector. Furthermore, in this embodiment, control is performed by detecting the amount of deviation between the virtual information generator that generates virtual information (virtual viewpoint image, virtual feature amount) and the virtual information.
[0016] In this embodiment, for example, when the posture of the arm's real end effector is brought closer to the posture of the end effector capable of grasping an object, the corner features of the virtual viewpoint image and the object information shown in the virtual viewpoint image are calculated at predetermined intervals and treated as virtual features. Furthermore, in this embodiment, an object in the real image that is highly correlated with the object corners in the virtual image is identified, and a deviation estimation is trained and used to minimize the Euclidean distance between the pair of corner features in the two images. This makes it possible to reduce the deviation in the relative pose between the hand and the target object.
[0017] <Misalignment due to head camera> Here, we will explain the deviation that occurs when the hand is controlled based on an image captured by a camera installed on the head. FIG. 1 is a diagram illustrating the misalignment caused by the head camera. The position of image g11 is the initial position of the hand. From this initial position, the hand is brought closer to the target object based on images captured by the camera installed on the head (g11 to g13). Then, as shown in image g13, when the hand approaches the target object, a misalignment occurs between the actual pose g21 of the target object and the pose g22 estimated at the initial position if the target object is small.
[0018] <Robot configuration example> First, an example of a robot and an example of a camera's grounding location will be described. FIG. 2 is a diagram showing an example of a robot according to this embodiment and an example of a location where a camera is placed on the ground. The robot 2 includes, for example, a body 21, a head 22, arms 121 (121L, 121R), and hands (end effectors) 1 (1L, 1R). The hand 1 has a camera 130 (sensor) installed on the palm, for example, and cameras 140 (141, 142, 143, 144, 145) (sensors) installed on the fingertips, for example, of each finger portion. Also, a camera 151 (sensor) is installed on the wrist. Furthermore, a camera 161 (sensor) is installed on the head 22.
[0019] In the example shown in FIG. 2, the hand 1 has five fingers, but the number of fingers may be two or more. Also, in FIG. 2, an example of a robot with two arms is shown, but it may have one arm. Also, the number of cameras is an example and is not limited to this. Furthermore, the number of cameras may be at least one of the cameras on the palm, fingertips, and wrist.
[0020] <Example of end effector configuration> Next, an example of the configuration of the end effector will be described. Fig. 3 is a diagram showing an example of the configuration of an end effector according to this embodiment. As shown in Fig. 3, end effector 1 (hand) includes finger portion 101, finger portion 102, finger portion 103, finger portion 104, base 111, camera 131, camera 132, camera 133, camera 134, camera 141, camera 142, camera 143, camera 144, and camera 151. End effector 1 is connected via a joint to arm 121, which is a movement mechanism that can move the position of end effector 1.
[0021] Furthermore, finger unit 101 is equipped with a force sensor 141, for example, at its fingertip. Finger unit 102 is equipped with a force sensor 142, for example, at its fingertip. Finger unit 103 is equipped with a force sensor 143, for example, at its fingertip. Finger unit 104 is equipped with a force sensor 144, for example, at its fingertip. Finger unit 101 corresponds, for example, to a human thumb, finger unit 102 corresponds, for example, to a human index finger, finger unit 103 corresponds, for example, to a human middle finger, and finger unit 104 corresponds, for example, to a human ring finger. Arm 121 is also a mechanical unit that can change the wrist posture of end effector 1. The end effector 1 has at least two fingers, but the number of fingers may be three or more.
[0022] Cameras 131 to 134, 141 to 144, and 151 are RDG-D cameras that can obtain RGB information and depth information. Alternatively, cameras 141 to 144 may be, for example, a combination of a CCD (Charge Coupled Device) imaging device or a CMOS (Complementary MOS) imaging device and a depth sensor.
[0023] The camera 131 is installed at a position corresponding to the ball of a human's foot, for example. Note that the camera 131 may also be installed at a position corresponding to the outer side of the proximal joint of a human's thumb, not on the index finger side.
[0024] The camera 132 is installed at a position corresponding to the ball of a human's foot, for example. Alternatively, the camera 132 may be installed at a position corresponding to the index finger side of the proximal joint of a human's thumb.
[0025] The camera 133 is installed, for example, at a position corresponding to the side of the thumb side of the base of the four human fingers, or the side of the thumb and index finger at the thenar eminence.
[0026] The camera 134 is installed at a position corresponding to the side of the ring finger at the base of the human finger and the side of the little finger thermae, for example. The end effector 1 may be provided with at least one of the cameras 131 to 134 on the palm.
[0027] The camera 141 is installed at a position corresponding to the tip of a human thumb, for example. The camera 141 may be installed in an area 161 including the tip, distal part, and distal part of the finger corresponding to the human thumb.
[0028] The camera 142 is installed at a position corresponding to the tip of a human index finger, for example. The camera 142 may be installed in an area 162 including the tip, distal part, and distal part of the finger corresponding to the human index finger.
[0029] The camera 143 is installed at a position corresponding to the tip of a human middle finger or ring finger, for example. The camera 143 may be installed in an area 162 including the tip, distal phalanx, and distal phalanx corresponding to the human middle finger or ring finger.
[0030] The camera 144 is installed at a position corresponding to the tip of a human little finger or ring finger, for example. The camera 144 may also be installed in an area 162 including the tip, distal phalanx, and distal phalanx of the finger corresponding to the human little finger or ring finger.
[0031] The cameras 141 to 144 are installed near the contact points of the fingers (places where the fingers come into contact with the target object for grasping).
[0032] The camera 155 is placed at the wrist.
[0033] In addition, the fingertips, joints, wrists, etc. are equipped with six-axis sensors, position sensors, etc. The fingertips are also equipped with pressure sensors.
[0034] <Spacing when bringing hands closer> In this embodiment, when the hand is brought closer to the target object, the processing is switched at predetermined intervals. FIG. 4 is a diagram for explaining the predetermined intervals when the hand according to this embodiment is brought closer to the target object. As shown in FIG. 4, the predetermined intervals have, for example, three levels, such as long distance, medium distance, and short distance. Note that these intervals are, for example, intervals that align level 0 of the image resolution at long distance with level 2 of the image resolution at short distance, and level 0 of the image resolution at medium distance with level 1 of the image resolution at short distance, just as changing the resolution in pyramid processing during image recognition.
[0035] Image g51 is an example of an image taken by each camera at a long distance, image g52 is an example of an image taken by each camera at a medium distance, and image g53 is an example of an image taken by each camera at a close distance.
[0036] Figure 5 is a diagram illustrating image estimation processing. If images in a pyramid are represented as a convolutional network, the information input from the sensors is connected to each other at the layers of the pyramid in a pyramidal convolution with different resolutions, as shown in Figure 5. The middle layer corresponds to an information exchange layer for features with different resolutions. The lower middle layer processes features with fine resolution. The upper middle layer processes features with sparse resolution. This network outputs object information by performing convolution on the processing results of these middle layer networks. Therefore, by inputting an image taken at a long distance into the middle layer of a model of a network trained at a short distance, it is possible to estimate how a close-up image looks from a long-distance image. Alternatively, by inputting an image taken at a medium distance into the middle layer of a model of a network trained at a short distance, it is possible to estimate how a close-up image looks from a medium-distance image.
[0037] FIG. 6 is a diagram showing an example of the configuration of a network used for recognition error estimation. Network g61 is used for recognition. Network g62 is used for control. Learning is performed using images captured by a camera and correct images. The network configuration in FIG. 6 is an example and is not limited to this.
[0038] FIG. 7 is a diagram for explaining a virtual position and a virtual viewpoint image according to this embodiment. In FIG. 7, position g71 is the initial position of the hand 1, position g72 is the estimated mid-distance position of the hand 1, and position g73 is the estimated gripping position (virtual position). Transition g74 is the planned route for moving the hand 1, passing through these positions. However, if a deviation, such as trajectory g75, is calculated as a result of calculation using a model based on an image captured at a long distance and the gripping position, then this deviation amount g76 is calculated. When correcting the trajectory in this way, a virtual branch image after the position correction is calculated, and the image created internally when corner detection is performed is the virtual viewpoint image. In other words, the virtual view (virtual viewpoint image) is the view (image that would be captured) from the imaging unit 100 for the posture of the hand 1 at each position, simulated by simulating the position of the hand 1.
[0039] By performing such processing, in this embodiment, an object in the real image that is highly correlated with the corner of the object in the virtual image is found, and a deviation estimation is trained so that the Euclidean distance between the corner feature pairs of the two images is minimized.
[0040] The estimation of the position and viewpoint image is not limited to the initial position at a long distance. For example, while moving the hand 1, photographing may be performed again at a position at a medium distance to obtain an image, and the position and viewpoint image may be estimated again using the obtained image to calculate the deviation and provide feedback. Furthermore, in the above example, three examples of long distance, medium distance, and short distance have been described as examples of predetermined intervals, but the number of intervals may be four or more.
[0041] <Operation image> Next, an operation image will be explained. FIG. 8 is a conceptual diagram of the operation when gripping a bolt according to this embodiment. Image g80 is an image of the object before grasping. Dashed triangles g81 to g84 are examples of the capture range of the cameras installed on the fingertips, palm, and wrist. Screw image g86 represents the actual position and posture of the target object, while screw image g87 represents the estimated position and posture of the misaligned target object. Alternatively, screw image g86 represents the actual visible position, and screw image g87 represents the desired position to approach.
[0042] Image g90 is an image of the target object after being grasped. In this embodiment, the deviation between the actual pose of the target object and the estimated pose of the target object is corrected using a trained model and an image taken from a distant position.
[0043] <End effector control device> Next, a configuration example of the end effector control device will be described. Fig. 9 is a diagram showing a configuration example of the end effector control device according to this embodiment. As shown in FIG. 9, the end effector control device 200 includes, for example, an acquisition unit 201, an image processing unit 202, a model storage unit 203, a learning unit 204, a recognition unit 205, a control unit 206, an arm driving unit 207, and a hand driving unit 208.
[0044] The end effector control device 200 is connected to the imaging unit 100, the arm 121, and the hand (end effector) 1 by wire or wirelessly.
[0045] The photographing unit 3 includes cameras 131 to 134 installed on the palms, cameras 141 to 144 installed on the fingertips, camera 151 installed on the wrist, and camera 161 installed on the head 22 of the robot 2, as described with reference to Figures 2 and 3.
[0046] The acquisition unit 201 acquires the image captured by the imaging unit 100. The acquisition unit 201 acquires the sensor value detected by the sensor provided in the end effector 1.
[0047] The image processing unit 202 performs image processing on the image captured by the image capturing unit 100 .
[0048] The model storage unit 203 stores a model based on a network learned by the learning unit 204. The model storage unit 203 may also store correct answer data used during learning. The model stored in the model storage unit 203 may be stored on the cloud. The model storage unit 203 may be connected to the end effector control device 200 via a wired or wireless network.
[0049] During learning, the learning unit 204 learns the model using the image captured by the murderous intent unit 100, the correct image, the estimated position, the correct position, and the like.
[0050] The recognition unit 205 recognizes, for example, the position, size, shape, and posture of an object from the seismic intensity information contained in the captured image and image processing (binarization, edge detection, cluster processing, feature extraction, etc.).
[0051] The control unit 206 switches the image resolution depending on the distance between the hand 1 and the target object. The control unit 206 estimates a medium-distance or short-distance image from a long-distance image. The control unit 206 generates a control command to reduce the difference between the estimated image and the intended image.
[0052] The arm driving unit 207 drives the arm 121 in response to a control command.
[0053] The hand driving unit 208 drives the hand 1 in response to a control command.
[0054] <Processing Procedure> Next, a description will be given of an example of a processing procedure performed by the end effector control device 200. Fig. 10 is a flowchart of the processing procedure of the end effector control device according to this embodiment.
[0055] (Step S11) The acquisition unit 201 acquires a captured image (visual information) of the target object to be grasped by the imaging unit 100 (sensor).
[0056] (Step S12) The recognition unit 205 determines a taxonomy (see Reference 1) that indicates the gripping state from the acquired and image-processed visual information. The recognition unit 205 determines the gripping shape and gripping position of the end effector 1 from the determined taxonomy and the acquired and image-processed visual information.
[0057] (Step S13) The control unit 206 inputs the gripping shape and gripping position of the end effector 1 into the trained model, and obtains a virtual position (trajectory) passing from the current position of the end effector 1 to the gripping position.
[0058] (Step S14) The control unit 206 inputs the acquired image and the determined virtual position into the trained model to determine virtual visual information at the virtual position. For example, the control unit 206 inputs an image captured at a long distance into the model and estimates captured images (virtual viewpoint information (image visual information)) at predetermined intervals up to a close distance position (virtual position). For example, the control unit 206 estimates virtual viewpoint images at each of the medium distance and close distance positions (virtual positions). Then, the control unit 206 calculates corner features of the target object captured in the virtual viewpoint image to determine virtual features.
[0059] (Step S15) The control unit 206 compares the visual information of the image information at the virtual position with the virtual viewpoint information to calculate the amount of misalignment. For example, the control unit 206 finds an object in the real image that is highly correlated with the corner feature of the target object captured in the virtual viewpoint image, and estimates the amount of misalignment so that the Euclidean distance between the pair of corner feature values of the two images is minimized. Note that the control unit 206 may, for example, perform an optimization process using the same method as in optical flow calculations to determine the similarity of brightness or contour within the kernel. Alternatively, the control unit 206 may, for example, use machine learning technology to build a learning model of the correlation level and use the learned model to perform the calculation.
[0060] (Step S16) The control unit 206 feeds back the calculated amount of deviation to the control amount of the end effector 1, and generates a control command.
[0061] Reference 1; Thomas Feix, Javier Romero, et al., “The GRASP Taxonomy of Human Grasp Types” IEEE Transactions on Human-Machine Systems (Volume: 46, Issue: 1, Feb. 2016), IEEE, p66-77
[0062] FIG. 11 is a diagram showing examples of names in the taxonomy. From these multiple options, the recognition unit 205 selects one based on the state of the target object (standing, lying down, tilted), the shape of the target object, and the task content (grasping the middle of the bottle, grasping the narrow top of the bottle, lifting the target object, grabbing a bolt and inserting it into a screw hole, etc.) based on the captured image and the current position and state of the hand, thereby determining the taxonomy.
[0063] <Method for determining the gripping shape and gripping position of the end effector> Next, an example of a method for determining the gripping shape and gripping position of the end effector will be described. FIG. 12 is a diagram for explaining a set of points indicating the center of a target object according to this embodiment. First, a black-and-white image g101 of the object to be grasped viewed from above is created to represent it as a polygon, as in image g100. Voronoi processing is then used to determine the center of the polygon using line g102, a set of nearby points. An axis g103 is then defined, either perpendicular to line g102, a set of points indicating the center of the object, or perpendicular to the outer surface of the object to be grasped. These are used to calculate frictional forces, etc., when pinching the long and short axes of the two fingers.
[0064] Image g110 is an example of a 2D or 3D copy of the centerline generated in this way. Image g120 is an image (polygon) of a rock, which is an object to be grasped g121, viewed obliquely from the side. A set of multiple adjacent points g122 is the target grasp center, and point g123 is the grasp center of the end effector. In this manner, the multiple points g122 that represent the center of the target object are also referred to as a center line in this embodiment.
[0065] The rule for creation is that if the center line is short (e.g., 5 samples or less), the object-centered frame is selected as the target grasp pose. Also, above the target grasp pose, only the quaternion (0,0,0,1) is used, and the finger direction is the same as the x-axis of the body frame.
[0066] In this embodiment, if the target object is a known object, the recognition unit 205 determines the grip shape and grip position using information about a taxonomy and a set of points indicating the center of the target object of the known model. Alternatively, if the target object is an unknown object, the recognition unit 205 determines the grip shape and grip position using information about a taxonomy and a set of points indicating the center of the target object or a skeleton line of the unknown object.
[0067] As described above, in this embodiment, multiple virtual hands are placed around the target object at a long distance, for example, at the initial position of the hand 1 away from the target object, between the target object and the real robot hand. In this embodiment, the object displacement is estimated from the difference between the appearance of the placed hand 1 and the virtual and actual appearances when the real hand arrives at that position. For example, in this embodiment, when the actual end effector posture of the arm is brought closer to the end effector posture capable of grasping an object, the virtual viewpoint image and the corner features of the object information shown in the virtual viewpoint image are calculated at predetermined intervals to calculate virtual feature amounts. Then, in this embodiment, an object in the real image that is highly correlated with the corners of the object in the virtual image is found, and deviation estimation is trained so that the Euclidean distance between the pair of corner features of the two images is minimized.
[0068] As a result, according to this embodiment, the end effector posture is constantly corrected based on the actual image, thereby reducing the deviation in the relative pose between the hand and the object.
[0069] A program for implementing all or part of the functions of the end effector control device 200 of the present invention may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed to perform all or part of the processing performed by the end effector control device 200. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a web page provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line.
[0070] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0071] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0072] 1, 1L, 1R... End effector 1, 101, 102, 103, 104... Finger portion, 111... Base body, 100, 131, 132, 134, 140, 141, 142, 143, 144, 142, 151, 161... Camera, 121, 121L, 121R... Arm, 2... Robot, 21... Body, 22... Head, 200... End effector control device, 201... Acquisition unit, 202... Image processing unit, 203... Model memory unit, 204... Learning unit, 205... Recognition unit, 206... Control unit, 207... Arm drive unit, 208... Hand drive unit
Claims
1. An end effector control method capable of grasping and manipulating a target object, comprising: an acquisition step of acquiring visual information of the target object by a sensor; a determination step of determining a gripping shape and a gripping position of the end effector from the acquired visual information; a virtual position determination step of determining a virtual position passing from the current position of the end effector to a gripping position; a virtual visual information calculation step of calculating virtual visual information of the sensor at the virtual position based on the acquired visual information and the virtual position; a deviation amount calculation step of calculating a deviation amount by comparing the visual information of the sensor at the virtual position with the virtual visual information; a feedback process of feeding back the deviation amount to a control amount of the end effector; and The determining step If the target object is a known object, a gripping shape and a gripping position are determined using information on a set of points indicating the center of the target object of a known model; If the target object is an unknown object, a gripping shape and a gripping position are determined using a set of points or a skeleton line of the unknown object that indicates the center of the target object. End effector control method.
2. A method for controlling an end effector capable of grasping and manipulating a target object, comprising: an acquisition step of acquiring visual information of the target object by a sensor; a determination step of determining a gripping shape and a gripping position of the end effector from the acquired visual information; a virtual position determination step of determining a virtual position passing from the current position of the end effector to a gripping position; a virtual visual information calculation step of calculating virtual visual information of the sensor at the virtual position based on the acquired visual information and the virtual position; a deviation amount calculation step of calculating a deviation amount by comparing the visual information of the sensor at the virtual position with the virtual visual information; a feedback process of feeding back the deviation amount to a control amount of the end effector; and The deviation amount calculation step includes: calculating a corner feature of the target object included in the virtual visual information to calculate a virtual feature; an object in a real image that is highly correlated with a corner feature of the target object included in the virtual visual information is found, and a deviation amount is estimated so that a Euclidean distance between the pair of corner features of the two images is minimized; End effector control method.
3. The virtual positions are at least two for each interval based on the resolution of the visual information. The end effector control method according to claim 1 or 2.
4. The virtual position determining step includes: determining the virtual position using the visual information acquired at an initial position of the end effector away from a grasping position; The end effector control method according to any one of claims 1 to 3.
5. The determination step determines a taxonomy based on the visual information, and determines a gripping shape and a gripping position of the end effector based on the determined taxonomy. The end effector control method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Robot, robot control method, and robot control program
JP2015066632A
Machine learning control of object handovers
JP2022024952A
Method for determining a grasping hand model
US20220009091A1