Image processing apparatus and image processing method
By acquiring and matching image features through an image processing device, the position and posture of the workpiece can be estimated, which solves the problem of high cost of 3D measuring instruments, realizes low-cost workpiece position and posture recognition, and improves the accuracy and efficiency of robotic arm gripping.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MINEBEAMITSUMI INC
- Filing Date
- 2020-10-29
- Publication Date
- 2026-04-28
AI Technical Summary
In the prior art, the cost of using 3D measuring instruments to identify the position and posture of workpieces is high, resulting in high costs when they are introduced in large quantities in factories. Therefore, it is preferable to use 2D images captured by ordinary cameras to identify the position and posture of workpieces.
The image processing device acquires first and second images of the bulk workpiece, generates a matching map of feature quantities, uses an estimation unit to estimate the position, posture and category classification score of the workpiece, estimates the position of the workpiece based on the matching result and the position estimation result, and generates control signals for the robotic arm to grasp the workpiece.
It enables the estimation of workpiece position and orientation through image processing, reducing recognition costs and improving the accuracy and efficiency of workpiece gripping.
Smart Images

Figure CN114631114B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an image processing apparatus and an image processing method. Background Technology
[0002] A technique is known for identifying the three-dimensional position and orientation of multiple loose objects (workpieces) in order to grasp them using a robotic arm or similar device. In this technique, the three-dimensional position and orientation of the workpieces can be identified using a three-dimensional measuring instrument.
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 2019-058960 Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] However, 3D measuring instruments are expensive, so large-scale adoption in factories and other facilities would be costly. Therefore, it is preferable to identify the position and pose of objects based on 2D images captured by ordinary cameras or other imaging devices.
[0008] Taking the above-mentioned problem as an example, the present invention aims to provide an image processing apparatus and an image processing method for estimating the position of an object.
[0009] Solution for solving the problem
[0010] An image processing apparatus according to one aspect of the present invention includes an acquisition unit and an estimation unit. The acquisition unit acquires a first image and a second image obtained by photographing bulk workpieces. The estimation unit generates a matching map of feature values of the first image and feature values of the second image, and estimates the position, posture, and category classification score of each target workpiece based on the first image and the second image, respectively. The position of the workpiece is estimated based on the matching result obtained using the matching map and the position estimation result.
[0011] Invention Effects
[0012] According to one aspect of the present invention, the position of an object can be estimated through image processing. Attached Figure Description
[0013] Figure 1 This is a diagram illustrating an example of an object gripping system equipped with the image processing apparatus of the first embodiment.
[0014] Figure 2 This is a block diagram illustrating an example of the structure of the object gripping system according to the first embodiment.
[0015] Figure 3This is a flowchart representing an example of learning processing.
[0016] Figure 4 This is a diagram representing an example of three-dimensional data of an object.
[0017] Figure 5 This is a diagram representing an example of a captured image of a virtual space configured with multiple objects.
[0018] Figure 6 This is a diagram illustrating an example of the processing related to the control of a robotic arm.
[0019] Figure 7 This is another example of the processing related to the control of a robotic arm.
[0020] Figure 8 This is a diagram illustrating an example of the detection model of the first embodiment.
[0021] Figure 9 This is a diagram showing an example of a feature map output by the feature detection layer (u1) of the first embodiment.
[0022] Figure 10 This is a diagram showing an example of the estimated position and orientation of the object in the first embodiment.
[0023] Figure 11 This is another example of the estimated result of the gripping position of the object in the first embodiment.
[0024] Figure 12 This is a diagram showing an example of a bulk image captured by a stereoscopic camera according to the first embodiment.
[0025] Figure 13 This is a diagram illustrating an example of the relationship between a bulk image and a matching diagram in the first embodiment.
[0026] Figure 14 This is a flowchart illustrating an example of the presumed processing of the first embodiment.
[0027] Figure 15 This is a diagram illustrating an example of the presumed processing of the first embodiment.
[0028] Figure 16 This is an example of a modified image of a bulk shipment containing a pallet.
[0029] Figure 17 This is a diagram illustrating an example of a positional offset estimation model for a variant.
[0030] Figure 18 This is a diagram illustrating another example of the position offset estimation model for the variant. Detailed Implementation
[0031] Hereinafter, the image processing apparatus and image processing method according to embodiments will be described with reference to the accompanying drawings. It should be noted that the present invention is not limited to these embodiments. Furthermore, the dimensional relationships and proportions of the elements in the drawings may sometimes differ from reality. Sometimes the drawings may also include portions with different dimensional relationships or proportions. Moreover, the content described in one embodiment or modification is generally applicable to other embodiments and modifications as well.
[0032] (First Implementation)
[0033] The image processing device in the first embodiment is used, for example, in the object gripping system 1. Figure 1 This is a diagram illustrating an example of an object gripping system equipped with the image processing apparatus of the first embodiment. Figure 1 The object gripping system 1 shown includes an image processing unit 10 (not shown), a camera 20, and a robotic arm 30. The camera 20 is positioned, for example, to capture images of both the robotic arm 30 and the bulk workpieces 41, 42, etc., gripped by the robotic arm 30. The camera 20 captures images of the workpieces 41, 42 and the robotic arm 30 and outputs these images to the image processing unit 10. It should be noted that the robotic arm 30 and the bulk workpieces 41, 42, etc., can also be captured by different cameras. Figure 1 As shown, the camera 20 in the first embodiment uses a camera capable of capturing multiple images, such as a known stereoscopic camera. The image processing device 10 uses the images output from the camera 20 to estimate the position and posture of workpieces 41, 42, etc. Based on the estimated position and posture of workpieces 41, 42, etc., the image processing device 10 outputs signals to control the movement of the robotic arm 30. The robotic arm 30 performs the action of grasping workpieces 41, 42, etc., based on the signals output from the image processing device 10. It should be noted that... Figure 1 The invention discloses several different types of workpieces 41, 42, etc., but the type of workpiece can also be singular. In the first embodiment, the case where there is only one type of workpiece will be described. Furthermore, workpieces 41, 42, etc., are configured with irregular positions and orientations. Figure 1 As shown, for example, multiple workpieces can also be configured to overlap in a top view. Furthermore, workpieces 41 and 42 are examples of objects.
[0034] Figure 2 This is a block diagram illustrating an example of the structure of the object gripping system according to the first embodiment. For example... Figure 2 As shown, the image processing device 10 is communicatively connected to the camera 20 and the robotic arm 30 via a network NW. Furthermore, as... Figure 2As shown, the image processing device 10 includes a communication I / F (interface) 11, an input I / F 12, a display 13, a storage circuit 14, and a processing circuit 15.
[0035] The communication I / F11 controls the data input and output communication with external devices via the network NW. For example, the communication I / F11 is implemented by a network card, network adapter, network interface controller (NIC), etc., receiving image data output from the camera 20 and sending signals to the robotic arm 30.
[0036] Input I / F12 is connected to processing circuit 15, converting input operations received from the administrator (not shown) of image processing device 10 into electrical signals and outputting them to processing circuit 15. For example, input I / F12 can be a switch button, mouse, keyboard, touch panel, etc.
[0037] The display 13 is connected to the processing circuit 15 and displays various information and image data output from the processing circuit 15. For example, the display 13 can be implemented as a liquid crystal monitor, a CRT (cathode ray tube) monitor, a touch panel, etc.
[0038] The storage circuit 14 is implemented, for example, by a storage device such as a memory. Various programs executed by the processing circuit 15 are stored in the storage circuit 14. Furthermore, various data used by the processing circuit 15 when executing the various programs are temporarily stored in the storage circuit 14. The storage circuit 14 has a mechanical (deep) learning model 141. Moreover, the mechanical (deep) learning model 141 has a neural network structure 141a and learning parameters 141b. The neural network structure 141a uses, for example… Figure 8 The well-known convolutional neural network b1, as discussed later, provides a basis for further research. Figure 15 The network structure is shown. The learning parameters 141b are, for example, the weights of the convolutional filters in a convolutional neural network, which are learned to estimate the position and pose of an object and are parameters to be optimized. The neural network structure 141a can also be set in the estimation unit 152. It should be noted that the mechanical (deep) learning model 141 in this invention is described using a learned model as an example, but is not limited to this. It should be noted that the mechanical (deep) learning model 141 will sometimes be simply referred to as "learning model 141" below.
[0039] The learning model 141 is used for processing that infers the position and posture of a workpiece based on images output from the camera 20. The learning model 141 is generated, for example, by learning the positions and postures of multiple workpieces and using images obtained by photographing these multiple workpieces as teacher data. It should be noted that in the first embodiment, the learning model 141 is generated, for example, by the processing circuit 15, but is not limited thereto, and may also be generated by an external computer. Hereinafter, an embodiment in which the learning model 141 is generated and updated by a learning device (not shown) will be described.
[0040] In the first embodiment, a large number of images for generating the learning model 141 are generated, for example, by configuring multiple artifacts in a virtual space and capturing images of that virtual space. Figure 3 This is a flowchart illustrating an example of learning processing. For example... Figure 3 As shown, the learning device acquires the three-dimensional data of the object (step S101). The three-dimensional data can be acquired, for example, by known methods such as 3D scanning. Figure 4 This is a diagram representing an example of three-dimensional data of an object. By acquiring three-dimensional data, the posture of the workpiece can be arbitrarily changed and configured in virtual space.
[0041] Next, the learning device sets various conditions for configuring objects in the virtual space (step S102). The configuration of objects in the virtual space can be performed, for example, using known image generation software. The number, position, and pose of the configured objects can be set to be randomly generated by the image generation software, but this is not limited to this; they can also be arbitrarily set by the administrator of the image processing device 10. Next, the learning device configures the objects in the virtual space according to the set conditions (step S103). Next, the learning device acquires, for example, the image, position, and pose of the configured objects by capturing the virtual space containing multiple objects (step S104). In the first embodiment, the position and pose of the objects are represented, for example, by three-dimensional coordinates (x, y, z), and the pose of the objects is represented by four-dimensional numbers, i.e., quaternions (qx, qy, qz, qw), representing the pose or rotation state of the object. Figure 5 This is an example diagram representing a captured image of a virtual space configured with multiple objects. For example... Figure 5As shown, multiple objects W1a and W1b are randomly positioned and positioned in the virtual space. Furthermore, the images of the randomly positioned objects will sometimes be referred to as "bulk images". Next, the learning device saves the acquired images and the positions and positions of the positioned objects in the storage circuit 14 (step S105). Then, the learning device repeats steps S102 to S105 a predetermined number of times (step S106). It should be noted that, here, the combination of the images acquired through the above steps and the positions and positions of the positioned objects saved in the storage circuit 14 is sometimes referred to as "teacher data". By repeatedly performing the learning process by repeating steps S102 to S105 a predetermined number of times, a sufficient amount of teacher data is generated.
[0042] Furthermore, the learning device generates or updates the learning parameters 141b used as weights in the neural network structure 141a by performing a predetermined number of learning processes using the generated teacher data (step S107). In this way, by configuring the object with acquired 3D data on the virtual space, teacher data including a combination of the object's image, position, and pose can be easily generated for learning processing.
[0043] Back Figure 2 The processing circuit 15 is implemented by a processor such as a central processing unit (CPU). The processing circuit 15 controls the entire image processing device 10. The processing circuit 15 performs various processes by reading and executing various programs stored in the storage circuit 14. For example, the processing circuit 15 includes an image acquisition unit 151, an estimation unit 152, and a robot control unit 153.
[0044] Image acquisition unit 151 acquires bulk images, for example, via communication I / F11 and outputs them to estimation unit 152. Image acquisition unit 151 is an example of an acquisition unit.
[0045] The estimation unit 152 uses the output bulk image to estimate the position and orientation of the object. For example, the estimation unit 152 uses the learning model 141 to perform estimation processing on the image of the object and outputs the estimation result to the robot control unit 153. It should be noted that the estimation unit 152 can also further estimate the position and orientation of a tray or similar object placed on the object. The structure for estimating the position and orientation of the tray will be explained later.
[0046] The robot control unit 153 generates a signal to control the robotic arm 30 based on the estimated position and posture of the object, and outputs this signal to the robotic arm 30 via the communication I / F11. For example, the robot control unit 153 acquires information related to the current position and posture of the robotic arm 30. Furthermore, the robot control unit 153 generates a trajectory for the robotic arm 30 to move when grasping the object based on the current position and posture of the robotic arm 30 and the estimated position and posture of the object. It should be noted that the robot control unit 153 can also correct the trajectory of the robotic arm 30 based on the position and posture of the tray, etc.
[0047] Figure 6 This is a diagram illustrating an example of the processing related to the control of a robotic arm. (Example:) Figure 6 As shown, the estimation unit 152 estimates the position and orientation of the target object based on the bulk image. Similarly, the estimation unit 152 can also estimate the position and orientation of the tray or the like on which the target object is placed based on the bulk image. The robot control unit 153 calculates the position coordinates and orientation of the fingers of the robotic arm 30 based on the estimated model of the target object and the tray, and generates the trajectory of the robotic arm 30.
[0048] It should be noted that the robot control unit 153 can also further output signals to control the movement of the robotic arm 30 used to arrange the grasped objects after the robotic arm 30 grasps the objects. Figure 7 This is another example of the processing related to the control of a robotic arm. (See diagram below.) Figure 7 As shown, the image acquisition unit 151 acquires an image captured by the camera 20 of an object grasped by the robotic arm 30. The estimation unit 152 estimates the position and posture of the object grasped by the robotic arm 30 as the target and outputs it to the robot control unit 153. Furthermore, the image acquisition unit 151 may also acquire an image captured by the camera 20 of a tray or similar object as the arrangement destination, where the tray or similar object becomes the moving destination of the grasped object. At this time, the image acquisition unit 151 further acquires an image of the object already arranged on the tray or similar object as the arrangement destination (arrangement completion image). The estimation unit 152 estimates the position and posture of the tray or similar object as the arrangement destination, and the position and posture of the arranged object, based on the arrangement destination image or the arrangement completion image. Furthermore, the robot control unit 153 calculates the position coordinates and posture of the fingers of the robotic arm 30 based on the estimated position and posture of the object grasped by the robotic arm 30, the position and posture of the tray or similar object as the arrangement destination, and the position and posture of the arranged object, and generates the trajectory of the robotic arm 30 during object arrangement.
[0049] Next, the estimation process in estimation unit 152 will be explained. Estimation unit 152 uses, for example, a model employing a known object detection model with downsampling, upsampling, and skip connections to extract feature quantities of the object. Figure 8 This is a diagram illustrating an example of the detection model of the first embodiment. Figure 8 In the object detection model shown, layer d1, for example, uses convolutional neural network b1 to downsample the bulk image P1 (320×320 pixels) into a 40×40 grid, and calculates multiple features (e.g., 256) for each grid. Furthermore, layer d2, which is lower than layer d1, divides the grid from layer d1 into a coarser grid (e.g., 20×20 grid) and calculates the features for each grid. Similarly, layers d3 and d4, which are lower than layers d1 and d2 respectively, further coarseen the grid from layer d2. Layer d4 calculates the features through upsampling with a finer division, and then integrates these features with those from layer d3 via skip connection s3 to generate layer u3. Skip connection can be a simple addition, a combination of features, or a transformation of the features from layer d3 using a convolutional neural network. Similarly, by skipping connection s2, the feature values of layer d2 are integrated with the feature values calculated by upsampling layer u3 to generate layer u2. Then, layer u1 is generated in the same way. As a result, in layer u1, the feature values of each grid are calculated, which are divided into 40×40 grids just like in layer d1.
[0050] Figure 9 This is a diagram showing an example of a feature map output by the feature extraction layer (u1) of the first embodiment. Figure 9 The feature map shown represents the horizontal grid of the bulk image P1, which is divided into a 40×40 grid, and the vertical grid represents the vertical grid. Furthermore, Figure 9 The depth direction of the feature map shown represents the elements of the feature quantities in each grid.
[0051] Figure 10 This is a diagram illustrating an example of the estimated position and orientation of the object in the first embodiment. (Example) Figure 10As shown, the estimation unit outputs two-dimensional coordinates (Δx, Δy) representing the position of the object, quaternions (qx, qy, qz, qw) representing the pose of the object, and category classification scores (C0, C1, ..., Cn). It should be noted that in the first embodiment, the depth value representing the distance from the camera 20 to the object in the coordinates representing the position of the object is not calculated as an estimation result. The structure for calculating the depth value will be explained later. It should be noted that the depth referred to here is the distance from the z-coordinate of the camera to the z-coordinate of the object in the z-axis direction parallel to the optical axis of the camera. It should be noted that the category classification score is the value output for each grid cell, representing the probability that the center point of the object is contained within that grid cell. For example, if there are n types of objects, the probability of not containing the center point of the object is added to output n+1 category classification scores. For example, if there is only one type of workpiece as an object, two category classification scores are output. Furthermore, if multiple objects exist within the same grid cell, the probability of objects stacked higher is output.
[0052] exist Figure 10 In this context, point C represents the center of the grid Gx, and point ΔC, with coordinates (Δx, Δy), represents, for example, the center point of the detected object. That is, in... Figure 10 In the example shown, the center of the object is offset by Δx from the center point C of the grid Gx in the x-axis direction and by Δy in the y-axis direction.
[0053] It should be noted that it can also be used as a substitute. Figure 10 And such Figure 11 The program sets arbitrary points a, b, and c outside the center of the object and outputs the coordinates (Δx1, Δy1, Δz1, Δx2, Δy2, Δz2, x3, Δy3, Δz3) of these arbitrary points a, b, and c relative to the center point C of the grid Gx. It should be noted that these arbitrary points can be set at any location on the object; they can be a single point or multiple points.
[0054] It should be noted that when the grid division is coarser than the size of the object, multiple objects may enter into one grid, which may cause the features of each object to be mixed together and cause false detection. Therefore, in the first embodiment, only the fine feature amount (40×40 grid) generated at the end is used as the feature map output of the calculated feature extraction layer (u1).
[0055] Furthermore, in the first embodiment, the distance from the camera 20 to the object is determined, for example, by using a stereo camera to capture two images, left and right. Figure 12 This is a diagram illustrating an example of a bulk image captured by a stereoscopic camera according to the first embodiment. For example... Figure 12As shown, the image acquisition unit 151 acquires two separate images: a left image P1L and a right image P1R. Furthermore, the estimation unit 152 performs estimation processing on both the left image P1L and the right image P1R using the learning model 141. It should be noted that during the estimation processing, some or all of the learning parameters 141b used for the left image P1L can be shared as weights for the right image P1R. It should also be noted that instead of using a stereo camera, a single camera can be used, and the camera's position can be shifted to capture images equivalent to the left and right images at two different locations.
[0056] Therefore, in the first embodiment, the estimation unit 152 suppresses misidentification of objects by using a matching map obtained by combining the feature values of the left image P1L with the feature values of the right image P1R. In the first embodiment, the matching map shows the strength of the correlation between the feature values in the right image P1R and the left image P1L for each feature value. That is, by using the matching map, the matching between the left image P1L and the right image P1R can be sought by focusing on the feature values in each image.
[0057] Figure 13 This is a diagram illustrating an example of the relationship between the bulk image and the matching diagram in the first embodiment. For example... Figure 13 As shown, in the matching map ML obtained with the left image P1L as the reference and the right image P1R as the corresponding grid, grid MLA is highlighted. This grid MLA is the grid in the left image P1L containing the center point of object W1L with the grid containing the highest correlation between its feature value and the feature value contained in the right image P1R. Similarly, in the matching map MR obtained with the right image P1R as the reference and the left image P1L as the corresponding grid, grid MRa is also highlighted. This grid MRa is the grid in the right image P1R containing the center point of object W1R with the grid containing the feature value contained in the left image P1L. Furthermore, the grid MLA with the highest correlation in the matching map ML corresponds to the grid containing object W1L in the left image P1L, and the grid MRa with the highest correlation in the matching map MR corresponds to the grid containing object W1R in the right image P1R. Therefore, it can be determined that the grid containing object W1L in the left image P1L is the same as the grid containing object W1R in the right image P1R. That is, in Figure 12 In the image, the consistent grid consists of grid G1L of the left image P1L and grid G1R of the right image P1R. Therefore, based on the X coordinates of object W1L in the left image P1L and object W1R in the right image P1R, the parallax of object W1 can be determined, and thus the depth z from camera 20 to object W1 can be determined.
[0058] Figure 14 This is a flowchart illustrating an example of the presumed processing of the first embodiment. Furthermore, Figure 15This is a diagram illustrating an example of the presumed processing of the first embodiment. Next, using... Figures 12 to 15 Let's explain. First, the image acquisition unit 151, as shown in the description... Figure 12 The left and right images of the object are obtained as shown in the left image P1L and right image P1R (step S201). Next, the estimation unit 152 calculates feature quantities for each grid in the horizontal direction of each left and right image. Here, in the case that each image is divided into a 40×40 grid as described above and 256 feature quantities are calculated for each grid, a 40-row, 40-column matrix is obtained in the horizontal direction of each image as shown in the first and second terms on the left side of equation (1).
[0059]
[0060] Next, Presumption Department 152 executes. Figure 15 The process m is shown. First, the estimation unit 152 calculates the matrix product, for example, using equation (1), which is the matrix product obtained by transposing the feature quantities of the same column extracted from the right image P1R for the feature quantities of the defined column extracted from the left image P1L. In equation (1), in the first term on the left, each feature quantity l11 to l1n in the first grid in the horizontal direction of the defined column of the left image P1L is arranged along the row direction. On the other hand, in the second term on the left of equation (1), each feature quantity r11 to r1n in the first grid in the horizontal direction of the defined column of the right image P1R is arranged along the column direction. That is, the matrix of the second term on the left is the matrix after transposing the matrix obtained by arranging each feature quantity r11 to r1m in the horizontal direction of the defined column of the right image P1R along the row direction. Furthermore, the right side of equation (1) is the matrix obtained by calculating the matrix product of the matrix of the first term on the left and the matrix of the second term on the left. The first column on the right side of equation (1) represents the relationship between the feature quantity of the first grid extracted from the right image P1R and the feature quantities of each grid in the horizontal direction of a defined column extracted from the left image P1L. The first row represents the relationship between the feature quantity of the first grid extracted from the left image P1L and the feature quantities of each grid in the horizontal direction of a defined column extracted from the right image P1R. That is, the right side of equation (1) represents the relationship between the feature quantities of each grid in the left image P1L and the feature quantities of each grid in the right image P1R. It should be noted that in equation (1), the subscript "m" represents the position of the grid in the horizontal direction of each image, and the subscript "n" represents the number of the feature quantity in each grid. That is, m is 1 to 40, and n is 1 to 256.
[0061] Next, the estimation unit 152 uses the calculated association graph to calculate the matching graph ML of the left image P1L relative to the right image P1R shown in matrix (1). The matching graph ML of the right image P1R relative to the left image P1L is calculated, for example, by using the Softmax function for the row direction of the association graph. As a result, the values of the association in the horizontal direction are normalized. That is, the conversion is performed so that the total value in the row direction is 1.
[0062]
[0063] Next, the estimation unit 152, for example, uses equation (2) to convolve the feature values extracted from the right image P1R with the calculated matching map ML. The first term on the left side of equation (2) is the matrix obtained by transposing matrix (1), and the second term on the left side is the matrix of the first term on the left side of equation (1). It should be noted that in this invention, the same feature values are used for obtaining the association and for convolving with the matching map. However, it is also possible to generate feature values for re-obtaining the association and feature values for convolution based on the extracted feature values using a convolutional neural network or the like.
[0064] Next, the estimation unit 152 combines the feature values obtained from equation (2) with the feature values extracted from the left image P1L, for example, by generating new feature values through a convolutional neural network. Thus, by integrating the feature values from both images, the estimation accuracy of position and pose is improved. It should be noted that... Figure 15 The process m in the middle can also be repeated multiple times.
[0065]
[0066] Next, the estimation unit 152 estimates the position, pose, and category classification based on the feature quantities obtained therein, for example, through a convolutional neural network. In general, the estimation unit 152 uses the calculated association graph to calculate the matching graph MR of the left image P1L relative to the right image P1R, as shown in matrix (2) (step S202). The matching graph MR of the left image P1L relative to the right image P1R is also calculated in the same way as the matching graph ML of the right image P1R relative to the left image P1L, for example, by using the Softmax function for the row direction of the association graph.
[0067]
[0068] Next, the estimation unit 152 convolves the feature quantity of the left image P1L with the calculated matching map, for example, by using equation (3). The first term on the left side of equation (3) is matrix (2), and the second term on the left side is the matrix before transpose of the matrix of the second term on the left side of equation (1).
[0069]
[0070] Next, the estimation unit 152 selects the grid with the largest estimation result for the category classification of the target (object) estimated based on the left image P1L, and compares it with a preset threshold (step S203). If the threshold is not exceeded, it is set to no target and the process ends. If the threshold is exceeded, the grid with the largest value is selected based on the matching map ML between the grid and the right image P1R (step S204).
[0071] Next, in the selected grid, the estimated category classification result of the target in the right image P1R is compared with a preset threshold (step S208). If the threshold is exceeded, the grid with the largest value is selected based on the matching map ML of the left image P1L for that grid (step S209). If the threshold is not exceeded, the category classification score of the grid selected based on the estimated result of the left image P1L is set to 0 and the process returns to step S203 (step S207).
[0072] Next, the grids in the matching image ML selected in step S209 are compared with those selected in step S204 based on the inference result of the left image P1L (step S210). If the grids are different, the classification score of the grids selected in step S204 based on the inference result of the left image P1L is set to 0, and the process returns to the grid selection in step S203 (step S207). Finally, based on the position information of the grids selected in the left image P1L and the right image P1R (e.g., ...), Figure 1 The disparity is calculated from the detection results of the horizontal x-value in the middle (step S211).
[0073] Next, based on the disparity calculated in step S211, the depth of the target is calculated (step S212). It should be noted that when calculating the depth for multiple targets, after step S211, the classification score of the grid selected based on the inference results of the left image P1L and the right image P1R is set to 0, and then the process returns to step S203. After that, the process is repeated until step S212.
[0074] As described above, the image processing apparatus 10 in the first embodiment includes an acquisition unit and an estimation unit. The acquisition unit acquires a first image and a second image obtained by photographing bulk workpieces. The estimation unit generates a matching map of feature values of the first image and feature values of the second image, and estimates the position, pose, and category classification score of each workpiece as a target based on the first image and the second image, respectively. Based on the matching result obtained using the attention map and the position estimation result, the workpiece position is estimated, and the depth from the stereo camera to the workpiece is calculated. This suppresses false detections in object recognition.
[0075] (Modified Example)
[0076] The embodiments of the present invention have been described above, but the present invention is not limited to the above embodiments, and various modifications can be made without departing from its spirit. For example, in the first embodiment, the case where the object (workpiece) is of one type has been described, but it is not limited thereto, and the image processing apparatus 10 may also be configured to detect multiple types of workpieces. Furthermore, the image processing apparatus 10 can detect not only the object, but also the position and posture of the tray or the like on which the object is placed. Figure 16 This is an example diagram showing a modified example of a bulk image containing a pallet. Figure 16 In the example shown, the image processing device 10 can determine the position and orientation of the tray containing the object to set a trajectory so that the robotic arm 30 will not collide with the tray. It should be noted that the tray, as the object being detected, is an example of an obstacle. The image processing device 10 could also be a structure that detects other objects besides the tray that act as obstacles.
[0077] Furthermore, the image processing apparatus 10 has been described using, for example, dividing a bulk image into a 40×40 grid. However, it is not limited to this; it can also be divided into finer or coarser grids to detect objects. In addition, estimation processing can be performed on a pixel-by-pixel basis. As a result, the image processing apparatus 10 can calculate the distance between the camera and the object with higher accuracy. Figure 17 This is a diagram illustrating an example of a model for estimating the positional offset of a variant. For example... Figure 17 As shown, the image processing apparatus 10 can also combine the left image P1L and the right image P1R by removing portions of the periphery of the estimated position that are smaller than the grid. Furthermore, the estimation process can be performed in the same manner as in the first embodiment, and the position offset can be estimated based on the processing result.
[0078] Furthermore, when estimation is performed using fine or coarse grid units or pixel units, estimation can also be performed separately in the left image P1L and the right image P1R, just as in the first embodiment. Figure 18 This is a diagram illustrating another example of a presumed model for positional offset in a modified example. In Figure 18 In the example shown, the image processing apparatus 10 performs estimation processing on the left image P1L and the right image P1R separately. In this case, the image processing apparatus 10 also shares the weighting for the left image P1L and the weighting for the right image P1R when performing their respective estimation processing, just as in the first embodiment.
[0079] Alternatively, the estimation processing described above can be performed on the robotic arm 30, the workpieces 41 and 42 held on the robotic arm 30, or the workpieces 41 and 42 arranged at the arrangement destination, without performing the estimation processing on the images of the bulk workpieces 41 and 42 as described above.
[0080] Furthermore, the present invention is not limited to the embodiments described above. Inventions constructed by appropriately combining the aforementioned constituent elements are also included in the present invention. Moreover, those skilled in the art can readily derive further effects and modifications. Therefore, the broader scope of the present invention is not limited to the embodiments described above, and various modifications can be made.
[0081] Explanation of reference numerals in the attached figures
[0082] 1: Object gripping system;
[0083] 10: Image processing device;
[0084] 20: Camera;
[0085] 30: Robotic arm;
[0086] 41, 42: Workpiece.
Claims
1. An image processing apparatus, characterized in that, have: The acquisition unit acquires a first image and a second image obtained by photographing the bulk workpiece; and The estimation unit generates a matching map by combining the feature values of the first image and the feature values of the second image. It estimates the position, pose, and category classification score of each target workpiece based on both the first and second images, and estimates the workpiece position based on the matching results obtained using the matching map and the position estimation results. The matching map shows the strength of the correlation between feature quantities for each feature quantity in the first image and the second image. The first grid cell with the highest category classification score of the first image is selected. For the selected first grid cell, the second grid cell with the highest correlation of feature values between the second image and the first image is selected. For the selected second grid cell, the third grid cell with the highest correlation of feature values between the first image and the second image is selected. If the first grid cell and the third grid cell are equal, the workpiece position is estimated based on the positions of the first grid cell and the second grid cell. The workpieces include multiple different types of bulk workpieces. The posture is represented by a four-dimensional number that indicates the posture or rotational state of an object.
2. The image processing apparatus according to claim 1, characterized in that, The acquisition unit is a stereo camera. The estimation unit calculates the depth from the stereo camera to the workpiece.
3. The image processing apparatus according to claim 1, characterized in that, The estimation unit also detects obstacles other than the workpiece in at least one of the first image and the second image.
4. An image processing method, characterized in that, The computer acquires first and second images obtained by photographing the bulk workpieces. The computer generates a matching map by combining the feature values of the first image and the feature values of the second image. Based on the first image and the second image, it estimates the position, pose, and category classification score of each target workpiece. The workpiece position is then estimated based on the matching results and position estimation results obtained using the matching map. The matching map shows the strength of the correlation between feature quantities for each feature quantity in the first image and the second image. The first grid cell with the highest category classification score of the first image is selected. For the selected first grid cell, the second grid cell with the highest correlation of feature values between the second image and the first image is selected. For the selected second grid cell, the third grid cell with the highest correlation of feature values between the first image and the second image is selected. If the first grid cell and the third grid cell are equal, the workpiece position is estimated based on the positions of the first grid cell and the second grid cell. The workpieces include multiple different types of bulk workpieces. The posture is represented by a four-dimensional number that indicates the posture or rotational state of an object.
Citation Information
Patent Citations
Robot system and workpiece take-out method
JP2019058960A
Image detection method, neural network training method, image detection device, neural network training device and electronic apparatus
CN108229523A
Information processing system, information processing device, information processing method, and program
WO2019138835A1