How to generate data for machine learning

By coating objects with wavelength-specific paints and using stereo cameras for point cloud data matching, the method addresses the challenge of accurately measuring overlapping objects, enhancing object recognition and grasping precision in systems with randomly stacked items.

JP7724056B2Active Publication Date: 2025-08-15MINEBEAMITSUMI INC

Patent Information

Application Number
JP2020188150
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-11-11
Publication Date
2025-08-15
Estimated Expiration
2040-11-11

AI Technical Summary

Technical Problem

Creating training data or evaluation data for object grasping systems with randomly stacked objects is challenging due to overlapping 3D shapes of multiple objects, leading to inaccurate measurement of position and orientation.

Method used

A method involving coating objects with specific paints that emit light at different wavelengths, using a stereo camera to capture images, and performing point cloud data matching to separate and measure the 3D shapes accurately.

Benefits of technology

Enables accurate acquisition of image data for machine learning, improving the precision of object recognition and grasping in systems with randomly stacked objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007724056000006
    Figure 0007724056000006
  • Figure 0007724056000007
    Figure 0007724056000007
  • Figure 0007724056000008
    Figure 0007724056000008
Patent Text Reader

Abstract

To provide a data generation method for machine learning which correctly acquires image data including a plurality of objects.SOLUTION: A method includes a first step S324 of imaging a first object and a second object (S301, S311) by performing irradiation of light in a shorter wavelength than that of visible light with a prescribed pattern on the first object to which a first visible light is emitted when light in a shorter wavelength than that of the visible light is received and to which a colorless transparent first paint for the visible light is applied, and the second object to which a second visible light is emitted when light in a shorter wavelength than that of the visible light is received and to which a colorless transparent second paint for the visible light is applied, and measuring the position and three-dimensional shape of the first object and the position and three-dimensional shape of the second object from the captured image.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for generating data for machine learning. law Regarding. [Background technology]

[0002] In object grasping systems, which use a robot arm to grasp multiple randomly stacked objects (workpieces), artificial intelligence such as convolutional neural networks is sometimes used to recognize the type, position, and orientation of an object from a photograph of the object. For the proper use of systems using artificial intelligence, machine learning using a huge amount of training data is essential, and evaluation data is also required to assess whether the results of the machine learning are appropriate. This results in high costs for preparing the data. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-103284 [Patent Document 2] Japanese Patent Application Laid-Open No. 2008-281399 Summary of the Invention [Problem to be solved by the invention]

[0004] Creating training data or evaluation data requires captured images of an object and the relative position and orientation of the camera and object. One method for identifying the object's position and orientation is to measure its 3D shape using a 3D scanner camera or similar device and match it with 3D data of the object created in advance using CAD or other software. However, in the case of a group of objects containing multiple objects piled up in a random order, the overlapping of multiple objects can make it difficult to match the 3D shape measurement results of each object with the 3D data. For example, if multiple objects are placed close together, the 3D shape of the combined objects may be mistakenly matched with the 3D data of the object, making it difficult to accurately measure the object's position and orientation.

[0005] The present invention has been made in view of the above, and provides a method for generating machine learning data that can accurately acquire image data including a plurality of objects. law The purpose is to provide. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems and achieve the object, a method for generating data for machine learning according to one aspect of the present invention includes a first step of irradiating a first object coated with a first paint that emits light with a first visible light when irradiated with light having a wavelength shorter than that of visible light and that is colorless and transparent to visible light, and a second object coated with a second paint that emits light with a second visible light when irradiated with light having a wavelength shorter than that of visible light and that is colorless and transparent to visible light, with light having a wavelength shorter than that of visible light in a predetermined pattern, photographing the first object and the second object with a shape acquisition camera, and measuring the three-dimensional shapes of the first object and the second object from the photographed images; , a stereo camera Taken with a location acquisition camera Reflect a second step of a third step of separating the three-dimensional shape measured in the first step into a three-dimensional shape of the first object and a three-dimensional shape of the second object based on the difference between the first visible light and the second visible light, and matching the three-dimensional data of the first object and the three-dimensional data of the second object that have been stored in advance; Equipped with. The direction in which the first object and the second object are photographed is changed, and the measurement of the three-dimensional shapes of the first object and the second object in the first step and the photographing of the first object and the second object in the second step are repeated, and the data is linked and updated by point cloud data matching. The relative position or orientation of the object with the position acquisition camera is calculated by the third step of measuring the three-dimensional shapes of the first object and the three-dimensional shapes of the second object in the first step, photographing the first object and the second object in the second step, separating the point cloud data of the first object and the second object, and performing point cloud data matching with the three-dimensional data of the objects.

[0007] According to one aspect, image data including a plurality of objects can be accurately acquired. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is an external view showing an example of an object grasping system. [Figure 2] FIG. 2 is a block diagram showing an example of the configuration of an object grasping system. [Figure 3] FIG. 3 is a diagram illustrating an example of a process related to control of a robot arm. [Figure 4] FIG. 4 is a diagram illustrating another example of the process related to the control of the robot arm. [Figure 5] FIG. 5 is a diagram illustrating an example of a detection model. [Figure 6] FIG. 6 is a diagram showing an example of a feature map output by the feature detection layer (u1). [Figure 7] FIG. 7 is a diagram showing an example of the estimation result of the position and orientation of the object. [Figure 8] FIG. 8 is a diagram showing another example of the estimation result of the grip position of the target object. [Figure 9] FIG. 9 is a diagram showing an example of a bulk pile image captured by a stereo camera. [Figure 10] FIG. 10 is a diagram showing an example of the relationship between the bulk pile image and the matching map. [Figure 11] FIG. 11 is a flowchart illustrating an example of the estimation process. [Figure 12] FIG. 12 is a diagram illustrating an example of the estimation process. [Figure 13] FIG. 13 is a diagram showing an example of a bulk pile image including a tray according to a modified example. [Figure 14] FIG. 14 is a diagram showing an example of a positional deviation estimation model according to a modified example. [Figure 15] FIG. 15 is a diagram showing another example of a positional deviation estimation model according to a modified example. [Figure 16]FIG. 16 is a block diagram showing an example of the configuration of a system for acquiring (measuring) three-dimensional data of a group of objects and generating a three-dimensional model of the group of objects. [Figure 17] FIG. 17 is a diagram illustrating an example of the measurement process. [Figure 18] FIG. 18 is a diagram showing an example of three-dimensional data of an object. [Figure 19] FIG. 19 is a flowchart illustrating an example of the pre-processing. [Figure 20] FIG. 20 is a diagram illustrating an example of the calibration process. [Figure 21] FIG. 21 is a diagram showing an example of a plurality of objects to which paint is applied. [Figure 22] FIG. 22 is a flowchart showing an example of a process for acquiring three-dimensional data of an object group. [Figure 23] FIG. 23 is a diagram showing an example of a state in which a UV pattern is irradiated onto a group of subjects from a projector. [Figure 24] FIG. 24 is a diagram showing an example of a state in which visible light is irradiated onto a subject group. [Figure 25] FIG. 25 is a diagram showing an example of a captured image of a virtual space in which a group of objects is arranged. [Figure 26] FIG. 26 is a block diagram showing an example of an information processing device that executes a measurement program. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, a machine learning data generation method and a machine learning data generation device according to an embodiment will be described with reference to the drawings. Note that the present invention is not limited to these embodiments. Furthermore, the dimensional relationships and ratios of elements in the drawings may differ from reality. The drawings may also include portions with different dimensional relationships and ratios. Furthermore, the content described in one embodiment or variant applies, in principle, to other embodiments or variants as well.

[0010] (Object Grasping System) FIG. 1 is an external view showing an example of an object gripping system 1. The object gripping system 1 shown in FIG. 1 includes an image processing device 10, a camera 20, and a robot arm 30 (not shown). The camera 20 is installed in a position where it can capture images of both the robot arm 30 and the objects to be gripped by the robot arm 30, such as bulk-stacked workpieces 41 and 42. The camera 20 captures images of the robot arm 30 and the workpieces 41 and 42, and outputs the images to the image processing device 10. Note that the images of the robot arm 30 and the bulk-stacked workpieces 41 and 42 may be captured by separate cameras. Hereinafter, an image of a randomly arranged object may be referred to as a "bulk-stacked image." As shown in FIG. 1, the camera 20 may be a camera capable of capturing multiple images, such as a known stereo camera. The image processing device 10 estimates the positions and orientations of the workpieces 41 and 42 using the images output from the camera 20. The image processing device 10 outputs a signal to control the operation of the robot arm 30 based on the estimated positions and postures of the workpieces 41, 42, etc. The robot arm 30 performs an operation to grasp the workpieces 41, 42, etc. based on the signal output from the image processing device 10. Note that although FIG. 1 discloses a plurality of different types of workpieces 41, 42, etc., there may be only one type of workpiece. Here, a case where there is only one type of workpiece will be described. Furthermore, the workpieces 41, 42, etc. are arranged so that their positions and postures are irregular. For example, as shown in FIG. 1, multiple workpieces may be arranged so that they overlap in a top view. Furthermore, the workpieces 41, 42 are an example of a target object.

[0011] Fig. 2 is a block diagram showing an example of the configuration of the object grasping system 1. As shown in Fig. 2, an image processing device 10 is communicably connected to a camera 20 and a robot arm 30 via a network NW. Also, as shown in Fig. 2, the image processing device 10 includes a communication I / F (interface) 11, an input I / F 12, a display 13, a memory circuit 14, and a processing circuit 15.

[0012] The communication I / F 11 controls data input / output communication with external devices via the network NW. For example, the communication I / F 11 is realized by a network card, a network adapter, a NIC (Network Interface Controller), etc., and receives image data output from the camera 20 and transmits signals to be output to the robot arm 30.

[0013] The input I / F 12 is connected to the processing circuitry 15, converts input operations received from an administrator (not shown) of the image processing device 10 into electrical signals, and outputs the electrical signals to the processing circuitry 15. For example, the input I / F 12 is a switch button, a mouse, a keyboard, a touch panel, or the like.

[0014] The display 13 is connected to the processing circuit 15 and displays various information and image data output from the processing circuit 15. For example, the display 13 is realized by a liquid crystal monitor, a CRT (Cathode Ray Tube) monitor, a touch panel, or the like.

[0015] The memory circuitry 14 is realized by, for example, a storage device such as a memory. The memory circuitry 14 stores various programs executed by the processing circuitry 15. The memory circuitry 14 also temporarily stores various data used when the various programs are executed by the processing circuitry 15. The memory circuitry 14 has a machine (deep) learning model 141. The machine (deep) learning model 141 further includes a neural network structure 141a and learning parameters 141b. The neural network structure 141a is, for example, an application of a known network such as the convolutional neural network b1 in FIG. 5, and is the network structure shown in FIG. 12, which will be described later. The learning parameters 141b are, for example, weights of a convolutional filter in the convolutional neural network, and are parameters that are learned and optimized to estimate the position and orientation of an object. The neural network structure 141a may be provided in the estimation unit 152. Note that, although the machine (deep) learning model 141 in the present invention will be described as a trained model, the present invention is not limited thereto. In the following, the machine (deep) learning model 141 may be simply referred to as the "learning model 141."

[0016] The learning model 141 is used in a process of estimating the position and orientation of a workpiece from an image output from the camera 20. The learning model 141 is generated, for example, by learning using the positions and orientations of a plurality of workpieces and images of the plurality of workpieces as training data. Note that here, the learning model 141 is generated, for example, by the processing circuitry 15, but is not limited to this and may be generated by an external computer.

[0017] The processing circuitry 15 is realized by a processor such as a CPU (Central Processing Unit). The processing circuitry 15 controls the entire image processing device 10. The processing circuitry 15 reads various programs stored in the storage circuitry 14 and executes the read programs to perform various processes. For example, the processing circuitry 15 includes an image acquisition unit 151, an estimation unit 152, and a robot control unit 153.

[0018] The image acquisition unit 151 acquires a bulk pile image via the communication I / F 11, for example, and outputs the image to the estimation unit 152. The image acquisition unit 151 is an example of an acquisition unit.

[0019] The estimation unit 152 estimates the position and orientation of the object using the output image of the bulk pile. The estimation unit 152 performs estimation processing on the image of the object using, for example, the learning model 141, and outputs the estimation result to the robot control unit 153. Note that the estimation unit 152 may further estimate, for example, the position and orientation of a tray or the like on which the object is placed. The configuration for estimating the position and orientation of the tray will be described later.

[0020] The robot control unit 153 generates a signal to control the robot arm 30 based on the estimated position and orientation of the object, and outputs the signal to the robot arm 30 via the communication I / F 11. The robot control unit 153 acquires, for example, information on the current position and orientation of the robot arm 30. Then, the robot control unit 153 generates a trajectory along which the robot arm 30 will move when grasping the object, based on the current position and orientation of the robot arm 30 and the estimated position and orientation of the object. Note that the robot control unit 153 may correct the trajectory along which the robot arm 30 will move, based on the position and orientation of a tray or the like.

[0021] 3 is a diagram showing an example of processing related to control of a robot arm. As shown in FIG. 3, the estimation unit 152 estimates the position and orientation of a target object from the bulk pile image. Similarly, the estimation unit 152 may estimate the position and orientation of a tray or other object on which the object is placed from the bulk pile image. The robot control unit 153 calculates the coordinates of the position and orientation of the tip of the robot arm 30 based on the estimated models of the object and tray, and generates a trajectory for the robot arm 30.

[0022] After the robot arm 30 grasps the object, the robot control unit 153 may further output a signal to control the operation of the robot arm 30 to align the grasped object. FIG. 4 is a diagram showing another example of processing related to control of the robot arm. As shown in FIG. 4, the image acquisition unit 151 acquires an image of the object grasped by the robot arm 30, captured by the camera 20. The estimation unit 152 estimates the position and posture of the target object grasped by the robot arm 30, and outputs the image to the robot control unit 153. The image acquisition unit 151 may further acquire an image of a tray or the like to which the grasped object will be moved, captured by the camera 20. In this case, the image acquisition unit 151 further acquires an image (aligned image) of the object already aligned on the tray or the like to which the object will be aligned. The estimation unit 152 estimates the position and posture of the tray or the like to which the object will be aligned, as well as the position and posture of the object already aligned, from the image of the alignment destination or the aligned image. Then, the robot control unit 153 calculates the coordinates and posture of the position of the hand of the robot arm 30 based on the estimated position and posture of the object grasped by the robot arm 30, the position and posture of the tray or the like to be aligned, and the position and posture of the object that has already been aligned, and generates a trajectory for the robot arm 30 when aligning the object.

[0023] Next, the estimation process in the estimation unit 152 will be described. The estimation unit 152 extracts feature quantities of an object using, for example, a model that applies a known object detection model with downsampling, upsampling, and skip connections. FIG. 5 is a diagram illustrating an example of a detection model. In the object detection model illustrated in FIG. 5, the d1 layer divides, for example, a bulk pile image P1 (320 × 320 pixels) into 40 × 40 grids by downsampling via a convolutional neural network b1, and calculates multiple feature quantities (e.g., 256 types) for each grid. Furthermore, the d2 layer, which is a layer below the d1 layer, divides the grids divided in the d1 layer more coarsely (e.g., into 20 × 20 grids) and calculates feature quantities for each grid. Similarly, the d3 layer and d4 layer, which are layers below the d1 and d2 layers, divide the grids divided in the d2 layer more coarsely, respectively. The d4 layer uses upsampling to calculate features at finer divisions, and at the same time, skip connections s3 are used to integrate these with the features of the d3 layer to generate the u3 layer. The skip connections can be simple additions or concatenations of features, or they can apply transformations such as those used in convolutional neural networks to the features of the d3 layer. Similarly, the u2 layer is generated by integrating the features calculated by upsampling the u3 layer with the features of the d2 layer using skip connections s2. The u1 layer is then generated in a similar manner. As a result, in the u1 layer, features for each grid divided into a 40x40 grid are calculated, just like in the d1 layer.

[0024] Fig. 6 is a diagram showing an example of a feature map output by the feature detection layer (u1). The horizontal direction of the feature map shown in Fig. 6 indicates each horizontal grid of the bulk pile image P1 divided into a 40x40 grid, and the vertical direction indicates each vertical grid. Furthermore, the depth direction of the feature map shown in Fig. 6 indicates the feature element in each grid.

[0025] FIG. 7 illustrates an example of the estimation results of the position and orientation of an object. As shown in FIG. 7, the estimation unit outputs two-dimensional coordinates (Δx, Δy) indicating the position of the object, quaternions (qx, qy, qz, qw) indicating the orientation of the object, and classification scores (C0, C1, ..., Cn). Note that, among the coordinates indicating the position of the object, the depth value indicating the distance from the camera 20 to the object is not calculated as the estimation result. The configuration for calculating the depth value will be described later. Note that the depth here refers to the distance from the z coordinate of the camera to the z coordinate of the object in the z-axis direction parallel to the optical axis of the camera. Note that the classification score is a value output for each grid and is the probability that the center point of the object is included in that grid. For example, if there are n types of objects, the probability that the center point of the object is not included is added to this, and n+1 classification scores are output. For example, if there is only one type of object, two classification scores are output. Also, if multiple objects exist in the same grid, the probability of the object stacked higher is output.

[0026] 7, point C indicates the center of grid Gx, and point ΔC with coordinates (Δx, Δy) indicates, for example, the center point of a detected object. That is, in the example shown in FIG. 7, the center of the object is offset from center point C of grid Gx by Δx in the x-axis direction and by Δy in the y-axis direction.

[0027] 7, arbitrary points a, b, and c other than the center of the object may be set as shown in FIG. 8, and the coordinates (Δx1, Δy1, Δz1, Δx2, Δy2, Δz2, x3, Δy3, Δz3) of the arbitrary points a, b, and c from point C at the center of grid Gx may be output. The arbitrary points may be set at any position on the object, and may be one point or multiple points.

[0028] If the grid division is too coarse compared to the size of the object, multiple objects may be included in one grid, and the features of each object may be mixed together, resulting in a false detection. Therefore, here we use only the feature map, which is the output of the feature extraction layer (u1), which calculates the final fine feature values (40 x 40 grid).

[0029] Here, for example, a stereo camera is used to capture two types of left and right images, thereby determining the distance from camera 20 to the object. FIG. 9 is a diagram showing an example of a bulk pile image captured by a stereo camera. As shown in FIG. 9, image acquisition unit 151 acquires two types of bulk pile images, a left image P1L and a right image P1R. Furthermore, estimation unit 152 performs estimation processing using learning model 141 on both left image P1L and right image P1R. Note that when performing the estimation processing, some or all of learning parameters 141b used for left image P1L may be shared as weighting for right image P1R. Note that instead of a stereo camera, a single camera may be used, and images corresponding to two types of left and right images may be captured at two different positions by shifting the camera position.

[0030] Therefore, here, the estimation unit 152 suppresses erroneous recognition of the object by using a matching map that combines the feature amounts of the left image P1L and the feature amounts of the right image P1R. The matching map indicates the strength of correlation between the feature amounts of the right image P1R and the left image P1L for each feature amount. In other words, by using the matching map, it is possible to match the left image P1L and the right image P1R by focusing on the feature amounts in each image.

[0031] FIG. 10 is a diagram illustrating an example of the relationship between a bulk pile image and a matching map. As shown in FIG. 10, in a matching map ML that is based on the left image P1L and corresponds to the right image P1R, a grid MLa that has the highest correlation between the feature amount of a grid containing the center point of an object W1L in the left image P1L and the feature amount included in the right image P1R is highlighted. Similarly, in a matching map MR that is based on the right image P1R and corresponds to the left image P1L, a grid MRa that has the highest correlation between the feature amount of a grid containing the center point of an object W1R in the right image P1R and the feature amount included in the left image P1L is highlighted. Furthermore, the grid MLa with the highest correlation in the matching map ML corresponds to the grid where the object W1L in the left image P1L is located, and the grid MRa with the highest correlation in the matching map MR corresponds to the grid where the object W1R in the right image P1R is located. This makes it possible to determine that the grid on which object W1L is located in the left image P1L matches the grid on which object W1R is located in the right image P1R. That is, in Fig. 9, the matching grids are grid G1L in the left image P1L and grid G1R in the right image P1R. This makes it possible to determine the parallax with respect to object W1 based on the X coordinate of object W1L in the left image P1L and the X coordinate of object W1R in the right image P1R, and therefore the depth z from camera 20 to object W1 can be determined.

[0032] FIG. 11 is a flowchart showing an example of the estimation process. FIG. 12 is a diagram showing an example of the estimation process. Hereinafter, the description will be made with reference to FIGS. 9 to 12. First, the image acquisition unit 151 acquires left and right images of the object, such as the left image P1L and right image P1R shown in FIG. 9 (step S201). Next, the estimation unit 152 calculates feature amounts for each grid in the horizontal direction of each of the left and right images. Here, as described above, when each image is divided into a 40×40 grid and 256 feature amounts are calculated for each grid, a 40-by-40 matrix is obtained in the horizontal direction of each image, as shown in the first and second terms on the left side of equation (1) (the product of the matrix of the first term and the matrix of the second term).

[0033]

number

[0034] Next, the estimation unit 152 executes process m shown in FIG. 12. First, the estimation unit 152 calculates the matrix product of the feature quantities of a specific column extracted from the left image P1L transposed with the feature quantities of the same column extracted from the right image P1R, for example, using equation (1). In equation (1), the first term on the left side represents the feature quantities l11 to l1n of the first grid in the horizontal direction of a specific column in the left image P1L, each arranged in the row direction. Meanwhile, in equation (1), the feature quantities r11 to r1n of the first grid in the horizontal direction of a specific column in the right image P1R, each arranged in the column direction. That is, the matrix in the second term on the left side is the transpose of the matrix in which the feature quantities r11 to r1m of the grid in the horizontal direction of a specific column in the right image P1R are each arranged in the row direction. Furthermore, the right side of equation (1) is calculated by calculating the matrix product of the matrix in the first term on the left side and the matrix in the second term on the left side. The first column on the right side of equation (1) represents the correlation between the feature of the first grid extracted from the right image P1R and the feature of each horizontal grid in a specific column extracted from the left image P1L, and the first row represents the correlation between the feature of the first grid extracted from the left image P1L and the feature of each horizontal grid in a specific column extracted from the right image P1R. That is, the right side of equation (1) represents a correlation map between the feature of each grid in the left image P1L and the feature of each grid in the right image P1R. Note that in equation (1), the subscript "m" indicates the horizontal grid position of each image, and the subscript "n" indicates the feature number in each grid. That is, m is 1 to 40, and n is 1 to 256.

[0035] Next, the estimation unit 152 uses the calculated correlation map to calculate a matching map ML of the right image P1R with respect to the left image P1L, as shown in matrix (1). The matching map ML of the right image P1R with respect to the left image P1L is calculated, for example, by applying a Softmax function to the row direction of the correlation map. This normalizes the correlation values in the horizontal direction. In other words, the values in the row direction are converted so that the sum of all values becomes 1.

[0036]

number

[0037] Next, the estimation unit 152 convolves the calculated matching map ML with features extracted from the right image P1R, for example, using equation (2). The first term on the left side of equation (2) is the transposed matrix (1), and the second term on the left side is the matrix of the second term on the left side of equation (1). Note that in the present invention, the same features are used for determining correlation and for convolution into the matching map; however, new features for determining correlation and new features for convolution may be generated separately from the extracted features using a convolutional neural network or the like.

[0038] Next, the estimation unit 152 combines the feature obtained by equation (2) with the feature extracted from the left image P1L to generate a new feature, for example, using a convolutional neural network. Integrating the feature values of the left and right images in this way improves the accuracy of estimating the position and orientation. Note that the process m in FIG. 12 may be repeated multiple times.

[0039]

number

[0040] Next, the estimation unit 152 estimates the position, orientation, and class classification from the obtained feature amounts, for example, by using a convolutional neural network. Additionally, the estimation unit 152 calculates a matching map MR of the left image P1L with respect to the right image P1R, as shown in matrix (2), using the calculated correlation map (step S202). The matching map MR of the left image P1L with respect to the right image P1R is calculated, for example, by applying a Softmax function to the row direction of the correlation map, similar to the matching map ML of the right image P1R with respect to the left image P1L.

[0041]

number

[0042] Next, the estimation unit 152 convolves the feature amount of the left image P1L into the calculated matching map, for example, using equation (3). The first term on the left side of equation (3) is matrix (2), and the second term on the left side is the matrix of the second term on the left side of equation (1) before transposition.

[0043]

number

[0044] Next, the estimation unit 152 selects and compares a preset threshold with the grid with the largest estimated result of the class classification of the target (object) estimated from the left image P1L (step S203). If the threshold is not exceeded, it is determined that there is no target and the process ends. If the threshold is exceeded, the grid with the largest value is selected from the matching map ML with the right image P1R for that grid (step S204).

[0045] Next, for the selected grid, the estimation result of the target classification in the right image P1R is compared with a preset threshold (step S208). If the threshold is exceeded, the grid with the largest value is selected from the matching map ML for that grid and the left image P1L (step S209). If the threshold is not exceeded, the classification score of the grid selected from the estimation result of the left image P1L is set to 0, and the process returns to step S203 (step S207).

[0046] Next, the grid of the matching map ML selected in step S209 is compared with the grid selected from the estimation result of the left image P1L in step S204 to determine whether they are equal (step S210). If the grids are different, the classification score of the grid selected from the estimation result of the left image P1L in step S204 is set to 0, and the process returns to the grid selection in step S203 (step S207). Finally, the disparity is calculated from the detection results of position information (e.g., the value of the horizontal direction x in FIG. 1) of the grids selected in the left image P1L and right image P1R (step S211).

[0047] Next, the depth of the target is calculated based on the parallax calculated in step S211 (step S212). Note that if depths are to be calculated for multiple targets, after step S211, the class classification score of the grid selected from the estimation results of the left image P1L and the right image P1R is set to 0, and then the process returns to step S203, and thereafter the process up to step S212 is repeated.

[0048] As described above, the image processing device 10 includes an acquisition unit and an estimation unit. The acquisition unit acquires a first image and a second image of randomly stacked workpieces. The estimation unit generates a matching map between the feature amounts of the first image and the feature amounts of the second image, estimates the position, posture, and classification score of each target workpiece for each of the first and second images, and calculates the depth from the stereo camera to the workpiece by estimating the workpiece position based on the matching result and the position estimation result using the attention map. This makes it possible to reduce false detections in object recognition.

[0049] (Image processing variation) The object grasping system 1 has been described above, but is not limited to the above description and various modifications are possible without departing from the spirit thereof. For example, while the description has been given of a case in which there is one type of target object (workpiece), the present invention is not limited to this configuration and the image processing device 10 may be configured to detect multiple types of workpieces. Furthermore, the image processing device 10 may not only detect the target object but also detect the position and orientation of a tray or the like on which the target object is placed. FIG. 13 is a diagram showing an example of a bulk pile image including a tray according to a modified example. In the example shown in FIG. 13, the image processing device 10 identifies the position and orientation of the tray on which the target object is placed, thereby enabling the robot arm 30 to set a trajectory that will prevent the robot arm 30 from colliding with the tray. The tray, which is the target to be detected, is an example of an obstacle. The image processing device 10 may also be configured to detect obstacles other than trays.

[0050] Furthermore, while the image processing device 10 has been described as dividing a bulk image into, for example, a 40x40 grid, this is not limiting and the image may be divided into finer or coarser grids to detect the object, or estimation processing may be performed pixel by pixel. This allows the image processing device 10 to more accurately calculate the distance between the camera and the object. FIG. 14 is a diagram showing an example of a positional deviation estimation model according to a modified example. As shown in FIG. 14, the image processing device 10 may cut out and combine portions of the left image P1L and the right image P1R that are smaller in size than the grids around the estimated position. Then, estimation processing may be performed similarly to the above-described estimation processing, and the positional deviation may be estimated based on the processing results.

[0051] Furthermore, when performing estimation processing in fine or coarse grid units or pixel units, estimation processing may be performed separately for each of the left image P1L and the right image P1R, as described above. FIG. 15 is a diagram showing another example of a displacement estimation model according to a modified example. In the example shown in FIG. 15, the image processing device 10 performs estimation processing separately for each of the left image P1L and the right image P1R. In this case, too, the image processing device 10 may share the weighting for the left image P1L with the weighting for the right image P1R when performing each estimation processing, as described above.

[0052] Furthermore, the above-described fixed processing may be performed not on images of the randomly piled workpieces 41 and 42, but on the robot arm 30, the workpieces 41 and 42 held by the robot arm 30, or the workpieces 41 and 42 aligned at the alignment destination.

[0053] (Acquisition and learning of 3D data of the subject group) The following describes a configuration for acquiring three-dimensional data of an object group including multiple objects, and a configuration for generating three-dimensional data of the object group using the acquired data. FIG. 16 is a block diagram showing an example of the configuration of a system for acquiring (measuring) three-dimensional data of an object group and generating a three-dimensional model of the object group. In FIG. 16, a processing device 110 and a three-dimensional data measuring device 140 are communicably connected via a network NW. The processing device 110 also includes a communication I / F (interface) 111, an input I / F 112, a display 113, a memory circuit 114, and a processing circuit 115.

[0054] The three-dimensional data measuring device 140 is connected to the projector 120, the stereo cameras 131 and 132, and the 3D scanner camera 150. The three-dimensional data measuring device 140 includes a communication I / F 146, an input I / F 142, a display 143, a memory circuit 144, and a processing circuit 145. In the following description, when there is no need to distinguish between the stereo cameras 131 and 132, they may be simply referred to as the camera 130.

[0055] Projector 120 is capable of irradiating (projecting) a predetermined pattern using light with a shorter wavelength than visible light (for example, UV (ultraviolet) light), and is also capable of illumination using visible light. Illumination using visible light may be performed using white light, or may be illumination using red light, blue light, and green light, or illumination that matches the color of expected indoor lighting. A visible light illumination device may be provided separately from projector 120. Camera 130 is a color camera or monochrome camera that can capture visible light.

[0056] The stereo cameras 131 and 132 are color or monochrome cameras capable of capturing visible light. Distance information to a group of objects is acquired by the 3D scanner camera 150. The 3D scanner camera 150 is a color camera that captures images of a group of objects irradiated with light having a shorter wavelength than visible light by the projector 120. Note that the 3D scanner camera 150 may constitute a part of the stereo camera 130, or at least one of the cameras 131 and 132 may be configured to capture images of a group of objects irradiated with light having a shorter wavelength than visible light.

[0057] In this embodiment, the 3D scanner camera 150 captures an image of a group of objects including a plurality of objects (workpieces) illuminated with light having a wavelength shorter than that of visible light by the projector 120. Similarly, the stereo cameras 131 and 132 capture images of the group of objects illuminated with visible light by the projector 120. FIG. 17 is a diagram showing an example of measurement processing. As shown in FIG. 17, the stereo cameras 131 and 132 capture images of a group of objects including a plurality of objects W1 and W2. Similarly, the 3D scanner camera 150 captures images of the plurality of objects W1 and W2 to obtain point cloud data indicating the distance to each of the objects W1 and W2.

[0058] In the processing device 110, the communication I / F 111 controls communication of data input / output with an external device via the network NW. For example, the communication I / F 111 is realized by a network card, a network adapter, a NIC (Network Interface Controller), or the like.

[0059] The input I / F 112 is connected to the processing circuit 115, converts input operations received from an administrator (not shown) of the processing device 110 into electrical signals, and outputs the electrical signals to the processing circuit 115. For example, the input I / F 112 is a switch button, a mouse, a keyboard, a touch panel, or the like.

[0060] The display 113 is connected to the processing circuit 115 and displays various information and image data output from the processing circuit 115. For example, the display 113 is realized by a liquid crystal monitor, a CRT (Cathode Ray Tube) monitor, a touch panel, or the like.

[0061] The memory circuitry 114 is realized by, for example, a storage device such as a memory. The memory circuitry 114 stores various programs executed by the processing circuitry 115. The memory circuitry 114 also temporarily stores various data used when the processing circuitry 115 executes the various programs. The memory circuitry 114 has image, position, and orientation data 1141 of a group of objects and a machine (deep) learning model 1142.

[0062] Furthermore, the machine (deep) learning model 1142 includes a neural network structure 1142a and learning parameters 1142b. The neural network structure 1142a is, for example, an application of a known network such as the convolutional neural network b1 in FIG. 5, and is the network structure shown in FIG. 12. The learning parameters 1142b are, for example, weights of the convolutional filter of the convolutional neural network, and are parameters that are learned and optimized to estimate the position and orientation of each object included in the object group.

[0063] The machine (deep) learning model 1142 is used in the object grasping system 1 (FIGS. 1 and 2) for processing to estimate the position and orientation of a workpiece from images output from the camera 20 (FIGS. 1 and 2). The machine (deep) learning model 1142 is generated, for example, by learning using the positions and orientations of multiple workpieces and images of the multiple workpieces as training data. Note that here, the machine (deep) learning model 1142 is generated, for example, by the processing circuitry 115, but is not limited to this and may be generated by an external computer.

[0064] The processing circuitry 115 is realized by a processor such as a CPU (Central Processing Unit). The processing circuitry 115 controls the entire processing device 110. The processing circuitry 115 reads various programs stored in the memory circuitry 114 and executes the read programs to perform various processes. For example, the processing circuitry 115 includes a learning unit 1151 and a data output unit 1152.

[0065] The learning unit 1151 performs machine learning of a machine (deep) learning model 1142 based on image, position, and posture data 1141 of a group of objects (including both newly measured and accumulated data by the three-dimensional data measuring device 140 and data accumulated in the past), and updates learning parameters 1142b.

[0066] The data output unit 1152 outputs the image, position, and orientation data 1141 of the object group stored in the memory circuitry 114 and the data of the machine (deep) learning model 1142 in response to an instruction from an operator or an external request.

[0067] In the three-dimensional data measuring device 140, the communication I / F 146 controls communication of data input and output with an external device via the network NW. For example, the communication I / F 146 is realized by a network card, a network adapter, a NIC, etc. The communication I / F 146 also transmits control signals to be output to the projector 120 in accordance with standards such as HDMI (registered trademark) (High-Definition Multimedia Interface), and receives status signals from the projector 120. The communication I / F 146 also transmits control signals to the camera 130 and receives image data output from the camera 130.

[0068] The input I / F 142 is connected to the processing circuit 145, converts input operations received from an administrator (not shown) of the three-dimensional data measuring device 140 into electrical signals, and outputs the signals to the processing circuit 145. For example, the input I / F 142 is a switch button, a mouse, a keyboard, a touch panel, or the like.

[0069] The display 143 is connected to the processing circuit 145 and displays various information and image data output from the processing circuit 145. For example, the display 143 is realized by a liquid crystal monitor, a CRT monitor, a touch panel, or the like.

[0070] The memory circuitry 144 is realized by, for example, a storage device such as a memory. The memory circuitry 144 stores various programs executed by the processing circuitry 145. The memory circuitry 144 also temporarily stores various data used when the various programs are executed by the processing circuitry 145. The memory circuitry 144 has object group three-dimensional data 1441 and image, position, and orientation data 1442 of the object group. The object group image, position, and orientation data 1442 is original data for a portion of the object group image, position, and orientation data 1141 stored in the memory circuitry 114 of the processing device 110, and is accumulated in the memory circuitry 114 via the communication I / F 146, the network NW, and the communication I / F 111 of the processing device 110.

[0071] In this embodiment, the three-dimensional data of individual objects included in the object group may be, for example, already stored in the memory circuitry 144. Fig. 18 is a diagram showing an example of three-dimensional data of the objects. In this case, the object group three-dimensional data measurement unit 1451 may detect each of the objects W1 and W2 included in the object group by, for example, matching an image captured by the camera 130 or the 3D scanner camera 150 with the object three-dimensional model M shown in Fig. 18 that has already been stored in the memory circuitry 144.

[0072] The processing circuitry 145 of the three-dimensional data measuring device 140 is realized by a processor such as a CPU. The processing circuitry 145 controls the entire three-dimensional data measuring device 140. The processing circuitry 145 reads various programs stored in the memory circuitry 144 and executes the read programs to perform various processes. For example, the processing circuitry 145 has an object group three-dimensional data measuring unit 1451.

[0073] The object group 3D data measurement unit 1451 controls the projector 120 and the camera 130 via the communication I / F 146, captures images of the object group captured by the camera 130, and acquires point cloud data indicating the 3D shape of each object group from the images of the object group captured by the 3D scanner camera 150. The 3D shape of each object included in the object group is measured based on images obtained by irradiating each object with a predetermined pattern and capturing the image, for example, using a known grid method. The 3D shape of each object is measured while changing the capturing direction of each object. In this way, by changing the capturing direction of the camera 130 and the 3D scanner camera 150, even if some objects are piled up and therefore cannot be seen from a certain angle, capturing the object data from a different angle makes it easier to acquire a wide range of object data, thereby improving the accuracy of matching the point cloud data with the 3D data described below. In this process, the 3D scanner camera 150 distinguishes each object included in the object group based on differences in visible light, as described below.

[0074] As a preparation for the processing, the stereo cameras 131 and 132 and the 3D scanner camera 150 are calibrated, and a predetermined paint is applied to each object. The predetermined paint is a fluorescent paint that emits visible light when irradiated with light having a wavelength shorter than that of visible light (e.g., ultraviolet (UV) light) and is colorless and transparent to visible light. This can also be called invisible paint. In this embodiment, each object is coated with an invisible paint that emits light with a different visible light. For example, object W1 is coated with a first invisible paint that emits light with a first visible light, and object W2 is coated with a second invisible paint that emits light with a second visible light. In the following, light with a wavelength shorter than that of visible light is described as UV light, but this is not limited thereto. For example, fluorescent paint can be excited by two-photon excitation in the same way as light with a wavelength shorter than that of visible light, and this can be used to irradiate patterned light.

[0075] FIG. 19 is a flowchart showing an example of pre-processing. As shown in FIG. 19, the object group three-dimensional data measurement unit 1451 of the three-dimensional data measuring device 140 performs a calibration process for the stereo cameras 131 and 132 (step S101). Note that, hereinafter, the cameras 131 and 132 constituting the stereo camera may be referred to as a stereo set. FIG. 20 is a diagram showing an example of the calibration process. As shown in FIG. 20, the calibration process for the stereo set is performed, for example, by the stereo cameras 131 and 132 photographing a known calibration board CB under visible light. The calibration process is performed by a commonly known method.

[0076] Next, the object group three-dimensional data measurement unit 1451 performs a calibration process between the projector 120, which emits light with a wavelength shorter than that of visible light, and the 3D scanner camera 150, which acquires point cloud data (step S102). Note that, hereinafter, the combination of the projector 120 and the 3D scanner camera 150 may be referred to as a 3D scan set. Furthermore, the object group three-dimensional data measurement unit 1451 also performs a calibration process between the stereo set and the 3D scan set (step S103). Note that the calibration process of the 3D scan set and the calibration process between the 3D scan set and the stereo set in steps S102 and S103 may be performed with visible light, or may be performed with light shorter than visible light using a calibration board printed with fluorescent paint.

[0077] Next, an invisible paint is applied to each individual object included in the object group to be measured (step S111). In this embodiment, each object included in the object group is coated with an invisible paint that emits a different visible light. Each object coated with the invisible paint is placed at a position where 3D data measurement is performed (step S112). The processes of steps S111 and S112 are repeated until all objects included in the object group are placed (step S113). FIG. 21 is a diagram showing an example of multiple objects coated with paint. As shown in FIG. 21, multiple objects (workpieces) W1 and W2 included in the object group are coated with different invisible paints. In this embodiment, the objects W1 and W2 are placed at any position, such as a position where they appear to overlap each other when photographed by the stereo cameras 131 and 132, as shown in FIG. 18.

[0078] After the pre-processing, the object group three-dimensional data measurement unit 1451 of the three-dimensional data measuring device 140 performs processing to acquire three-dimensional data of the arranged object group. Fig. 22 is a flowchart showing an example of the processing to acquire object group three-dimensional data, and is a processing example when the image is captured by the color camera 150.

[0079] In FIG. 22, when the three-dimensional data measuring device 140 starts processing, it projects a UV pattern onto the object using the projector 120. Then, the 3D scanner camera 150, which is positioned in a predetermined positional relationship with the projector 120, captures an image of each object included in the object group (step S301). FIG. 23 is a diagram showing an example of a state in which the UV pattern is projected onto the object group from the projector. The UV light emitted from the projector 120 causes the objects W1 and W2 to emit light in a predetermined pattern using invisible paint, and this state is captured by the 3D scanner camera 150. Because the visible light emission is due to the invisible paint, it does not matter if the object W is glossy or black. Also, in step S301, the object W1, which is coated with a first invisible paint and emits first visible light, and the object W2, which is coated with a second invisible paint and emits second visible light, are simultaneously captured.

[0080] 22, the photographing of the object (step S301) is repeated by changing the UV pattern irradiated from projector 120 (step S303) until the predetermined UV pattern is completed (step S302: No). Then, three-dimensional data measuring device 140 calculates the three-dimensional shape of each object group based on the multiple images photographed for each UV pattern, and stores the calculated shape in the memory circuit (step S304).

[0081] When processing for the predetermined UV pattern is completed (step S302: Yes), the three-dimensional data measuring device 140 next performs processing by irradiating the object with white light from the projector 120 or an external light source and capturing an image with the camera 130 (step S311). The three-dimensional data measuring device 140 then stores the captured image in a memory circuit (step S312). FIG. 24 is a diagram showing an example of a state in which visible light is irradiated onto a group of objects. The irradiation of white light from the projector 120 causes the first invisible paint on the object W1 and the second invisible paint on the object W2 to become colorless and transparent. This allows for the generation of machine learning data similar to the original states of the objects W1 and W2, which is expected to improve the recognition accuracy of the position and orientation of the workpiece when the workpiece is actually grasped by a robot arm.

[0082] Next, returning to FIG. 22, the three-dimensional data measuring device 140 determines whether processing from a predetermined direction has been completed (step S321). If it determines that processing has not been completed (step S321: No), it changes the imaging direction for the object group (step S322) and repeats the process from irradiation of the UV pattern (step S301). The posture of the object may be changed manually by the operator or by a driving mechanism. If the three-dimensional data measuring device 140 determines that processing from a predetermined direction has been completed (step S321: Yes), it proceeds to step S323. Note that if the positions and postures of all objects included in the object group can be identified by imaging from only one direction, it is not necessary to repeat the processes from step S322 onwards.

[0083] Next, the three-dimensional data measuring device 140 performs point cloud data matching between point cloud data measured from different directions and updates the data (step S323). At this time, the three-dimensional data measuring device 140 may perform point cloud data matching individually for the point cloud data of the object W1 acquired based on the first visible light and the point cloud data of the object W2 acquired based on the second visible light, or may perform point cloud data matching for the entire object group. Next, point cloud matching is performed between the point cloud data acquired based on the first visible light and the pre-stored 3D data. From the position and orientation of the 3D model at this time, the relative position with the stereo camera at the time of image capture is calculated, and data on the position and orientation of the object is obtained (S324). Next, point cloud matching is similarly performed between the point cloud data acquired based on the second visible light and the 3D model. If there is other data separated by luminous color, matching with the 3D model is repeated for each luminous color (S331, S332).

[0084] Then, the three-dimensional data measuring device 140 determines whether or not a predetermined number of images for generating data have been captured (step S341). If the three-dimensional data measuring device 140 determines that the predetermined number of images have not been captured (step S341: No), the three-dimensional data measuring device 140 repeats the process from the irradiation of the UV pattern (step S301) after the bulk pile state of the objects has been changed (step S342). In step S342, for example, the processes of steps S112 to S113 shown in FIG. 19 are performed. On the other hand, if the three-dimensional data measuring device 140 determines that the predetermined number of images have been captured (step S341: Yes), it ends the process.

[0085] Although the embodiments of the present invention have been described above, the present invention is not limited to the above embodiments, and various modifications are possible without departing from the spirit of the present invention. For example, the number of objects included in the object group is not limited to those described, and the object group may include three or more objects.

[0086] As described above, the machine learning data generation method according to the embodiment includes a first step of irradiating a first object coated with a first paint that emits light using a first visible light when irradiated with light having a wavelength shorter than that of visible light and that is colorless and transparent to visible light, and a second object coated with a second paint that emits light using a second visible light when irradiated with light having a wavelength shorter than that of visible light and that is colorless and transparent to visible light, with light having a wavelength shorter than that of visible light in a predetermined pattern, and capturing images of the first object and the second object. This makes it possible to easily measure three-dimensional data of an object group including the three-dimensional shapes of each individual object from an actual object group including multiple objects.

[0087] The machine learning data generation method may further include a second step of irradiating the first object and the second object with visible light and capturing images of the first object and the second object. The first step may include measuring the position and three-dimensional shape of the first object by detecting the first visible light, and measuring the position and three-dimensional shape of the second object by detecting the second visible light. The machine learning data generation method may further include a third step of separating the three-dimensional shapes of the first object and the second object from the three-dimensional shapes measured in the first step based on the difference between the first visible light and the second visible light, and matching the three-dimensional data of the first object and the second object with previously stored three-dimensional data. This makes it possible to easily measure three-dimensional data of a group of objects, including the three-dimensional shapes of each individual object.

[0088] In the first step, the three-dimensional shapes of the first object and the second object are measured by switching between a plurality of predetermined patterns for one posture of the first object and the second object, thereby improving the accuracy of the measurement of the three-dimensional shapes.

[0089] The machine learning data generation method may also include changing the direction in which the first object and the second object are photographed, repeating the measurement of the three-dimensional shapes of the first object and the second object in the first step and the photographing of the first object and the second object in the second step, and linking and updating the data by point cloud data matching. The machine learning data generation method may also include repeating the measurement of the three-dimensional shapes of the first object and the second object in the first step and the photographing of the first object and the second object in the second step, and linking and updating the data by matching the point cloud data separated for each object with the three-dimensional data. This makes it possible to obtain the three-dimensional shapes of each object that cannot be captured by photographing from only one side.

[0090] In addition, an object machine learning data measurement device according to an embodiment includes a first processing unit that irradiates a first object coated with a first paint that emits first visible light when irradiated with light having a wavelength shorter than that of visible light and a second object coated with a second paint that emits second visible light when irradiated with light having a wavelength shorter than that of visible light, with light having a wavelength shorter than that of visible light, in a predetermined pattern, photographs the first object and the second object, and measures the position and three-dimensional shape of the first object and the position and three-dimensional shape of the second object from the photographed images. This allows the above-described object three-dimensional data measurement method to be realized as an apparatus.

[0091] The various processes described in each embodiment can be realized by executing a prepared program on an information processing device. An example of an information processing device 900 that executes a program having the same functions as the motor control device 100 in each of the above embodiments will be described below. Fig. 26 is a block diagram showing an example of an information processing device that executes a measurement program.

[0092] 26 includes a communication device 910, an input device 920, an output device 930, a ROM 940, a RAM 950, a CPU 960, and a bus 970. For example, the communication device 910 functions as the communication I / F 146 shown in FIG. 16, the input device 920 functions as the input I / F 142 shown in FIG. 16, and the output device 930 functions as the display 143 shown in FIG. 16.

[0093] A measurement program that performs the same functions as those of the above embodiment is pre-stored in the ROM 940. The measurement program may be stored in a storage device such as a hard disk built into the computer system, or in a recording medium readable by a drive (not shown), instead of in the ROM 940. The recording medium may be, for example, a CD-ROM, a DVD disk, a portable recording medium such as a USB memory or an SD card, or a semiconductor memory such as a flash memory. The measurement program is a measurement program 940A, as shown in FIG. 26. The measurement program 940A may be integrated or distributed as appropriate.

[0094] Then, the CPU 960 reads out the measurement program 940A from the ROM 940, and loads the read program onto the work area of the RAM 950. Then, as shown in Fig. 26, the CPU 960 causes the measurement program 940A loaded onto the RAM 950 to function as a measurement process 950A.

[0095] The CPU 960 determines whether the identification information included in the setting change information received from the maintenance device 200 is for the motor control device in question. The CPU 960 changes the setting using the setting change information corresponding to the identification information determined to be for the motor control device in question. The CPU 960 transfers the setting change information including identification information of a motor control device other than the motor control device in question to another motor control device.

[0096] 26 illustrates an information processing device 900 that executes a measurement program having the same functions as the three-dimensional data measuring device 140. However, the processing device 110 and the image processing device 10 can also be realized by executing a program prepared in advance in the information processing device. Furthermore, the information processing device 900 is, for example, a standalone computer, but is not limited to this. It may be realized by multiple computers that can communicate with each other, or may be realized in a virtual machine on the cloud. Furthermore, all or part of the functions of the information processing device 900 may be realized by an integrated circuit, such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0097] Furthermore, the present invention is not limited to the above-described embodiments. The present invention also includes configurations in which the above-described components are appropriately combined. Furthermore, further effects and modifications can be easily derived by those skilled in the art. Therefore, the broader aspects of the present invention are not limited to the above-described embodiments, and various modifications are possible. [Explanation of symbols]

[0098] 110 processing device, 114 memory circuit, 1141 image, position, and orientation data of object group, 1142 machine (deep) learning model, 1142a neural network structure, 1142b learning parameters, 115 processing circuit, 1151 learning unit, 1152 data output unit, 120 projector, 130 (131, 132) stereo camera, 140 3D data measuring device, 144 memory circuit, 1441 object group 3D data, 1442 image, position, and orientation data of object group, 145 processing circuit, 1451 object group 3D data measuring unit, 150 3D scanner camera, W1, W2 object, M object 3D model

Claims

1. a first step of irradiating a first object coated with a first paint that emits light using first visible light when irradiated with light having a wavelength shorter than that of visible light and that is colorless and transparent to visible light, and a second object coated with a second paint that emits light using second visible light when irradiated with light having a wavelength shorter than that of visible light and that is colorless and transparent to visible light, with light having a wavelength shorter than that of visible light in a predetermined pattern, photographing the first object and the second object with a shape acquisition camera, and measuring the three-dimensional shapes of the first object and the second object from the photographed images; a second step of irradiating the first object and the second object with visible light and capturing images of the first object and the second object with a position acquisition camera that is a stereo camera; a third step of separating the three-dimensional shape measured in the first step into a three-dimensional shape of the first object and a three-dimensional shape of the second object based on a difference between the first visible light and the second visible light, and matching the three-dimensional data of the first object and the three-dimensional data of the second object that have been stored in advance; Equipped with changing a direction in which the first object and the second object are photographed, repeating the measurement of the three-dimensional shape of the first object and the three-dimensional shape of the second object in the first step and the photographing of the first object and the second object in the second step, and linking and updating the data by point cloud data matching; measuring the three-dimensional shape of the first object and the three-dimensional shape of the second object in the first step, photographing the first object and the second object in the second step, and calculating the relative position or orientation of the object with the position acquisition camera in the third step, which separates point cloud data of the first object and the second object and performs point cloud data matching with the three-dimensional data of the object; Methods for generating data for machine learning.

2. The machine learning data generation method according to claim 1 , wherein the shape acquisition camera is a camera different from the position acquisition camera.

3. 3. The method for generating data for machine learning according to claim 1, wherein in the first step, the three-dimensional shape of the first object is measured by detecting the first visible light, and the three-dimensional shape of the second object is measured by detecting the second visible light.

4. 4. The method for generating data for machine learning according to claim 1, wherein in the first step, the three-dimensional shape of the first object and the three-dimensional shape of the second object are measured by switching between a plurality of the predetermined patterns for one posture of the first object and the second object.

Citation Information

Patent Citations

  • Shape examining device

    JP1985063404A

  • Three-dimensional image pickup device and method

    JP2002027501A

  • Three-dimensional measuring device and portable measuring device

    JP2008281399A

  • Determination device such as determination device for bone of meat, determination method such as determination method for bone of meat, and determination program such as determination program for bone of meat

    JP2020103284A

Cited By

  • Robot and method for controlling same

    US20240383148A1