Object recognition device, control device, and object recognition method

JPWO2025027797A5Active Publication Date: 2025-07-08MITSUBISHI ELECTRIC CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024550894
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-07-08
Estimated Expiration
2043-08-01

AI Technical Summary

Technical Problem

Conventional object recognition technologies face challenges in accurately distinguishing and separating objects from multiple objects due to insufficient contour information, leading to reduced recognition accuracy, especially when using images captured by different cameras.

Method used

The object recognition device enhances feature extraction by converting object images into converted images with feature-enhanced edges, emphasizing common features across edges, allowing for improved recognition accuracy through an image acquisition unit, image conversion unit, and recognition unit configuration.

Benefits of technology

This approach improves object recognition accuracy by emphasizing common features, enabling better separation and recognition of objects, even when captured by different cameras, thereby enhancing the success rate of object picking by industrial robots.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An object recognition device (10A) is provided with an image acquisition unit (11), an image conversion unit (12), and a recognition unit (13). The image acquisition unit (11) acquires an image obtained by imaging an object. The image conversion unit (12) converts the image into a post-conversion image in which the edge of the object is replaced with the edge having emphasized features. The recognition unit (13) recognizes the object, on the basis of the post-conversion image. The object recognition device (10A) can improve the accuracy of object recognition when recognizing an object from an image obtained by imaging the object.
Need to check novelty before this filing date? Find Prior Art

Description

Object recognition device, control device, and object recognition method

[0001] The present disclosure relates to an object recognition device, a control device, and an object recognition method for recognizing an object.

[0002] To automate work at production sites, etc., industrial robots that pick objects one by one from a plurality of objects are known. When an object is picked by an industrial robot, a recognition process is executed to recognize the object to be picked, and a gripping position when the industrial robot is to pick up the object is detected.

[0003] Patent Literature 1 discloses an information processing device that trains a recognizer that executes a recognition process for recognizing an object from an image of the object. The information processing device disclosed in Patent Literature 1 acquires training data having characteristics equivalent to those of the recognition target data by converting images captured by a camera used to collect training data so that the images have characteristics equivalent to those of the recognition target data input to the recognizer. Such characteristics include image quality characteristics such as noise, blur, color tone, and white balance. The technology disclosed in Patent Literature 1 can prevent a decrease in recognition accuracy even when the characteristics of the camera used to collect training data differ from those of the camera used to acquire the recognition target data. In other words, the technology disclosed in Patent Literature 1 can prevent a decrease in recognition accuracy due to individual differences between cameras.

[0004] Japanese Patent Application Laid-Open No. 2021-82068

[0005] However, even if an image is converted by the conventional technology disclosed in Patent Document 1, sufficient information regarding the contours of an object may not be obtained. If sufficient information regarding the contours of an object is not obtained, it may be difficult to individually separate and recognize an object from among multiple objects. Therefore, with the above conventional technology, there is a problem in that the difficulty in individually separating and recognizing objects may result in a decrease in object recognition accuracy.

[0006] The present disclosure has been made in view of the above, and aims to provide an object recognition device that can improve the accuracy of object recognition.

[0007] In order to solve the above-mentioned problems and achieve the objectives, the object recognition device of the present disclosure includes an image acquisition unit that acquires an image of an object, an image conversion unit that converts the image into a converted image in which the edges of the object are replaced with edges with emphasized features, and a recognition unit that recognizes the object based on the converted image.

[0008] The object recognition device according to the present disclosure has the effect of improving the object recognition accuracy.

[0009] FIG. 1 is a diagram showing a functional configuration of an object recognition device according to a first embodiment; FIG. 2 is a flowchart showing a processing procedure executed by the object recognition device according to the first embodiment; FIG. 3 is a diagram showing a functional configuration of an object recognition device according to a second embodiment; FIG. 4 is a flowchart showing a processing procedure executed by the object recognition device according to the second embodiment; FIG. 5 is a diagram for explaining an example of an image conversion method by the object recognition device according to the second embodiment;

[0010] An object recognition device, a control device, and an object recognition method according to embodiments will be described in detail below with reference to the accompanying drawings.

[0011] 1 is a diagram showing the functional configuration of an object recognition device 10A according to a first embodiment. The object recognition device 10A recognizes an object from a captured image of the object. The object recognition device 10A includes an image acquisition unit 11, an image conversion unit 12, and a recognition unit 13.

[0012] The object recognized by the object recognition device 10A may be an object to be grasped by an industrial robot. The industrial robot may grasp an object recognized by the object recognition device 10A from among a plurality of objects and pick the object. The industrial robot repeats this operation to pick the objects one by one. The plurality of objects may have the same shape or may have irregular shapes. In the following description, when a plurality of objects have the same shape, each object is referred to as a fixed-shaped object. When a plurality of objects have irregular shapes, each object is referred to as an irregular-shaped object. An example of a fixed-shaped object is a part of an industrial product. An example of an irregular-shaped object is food that can be sorted into pieces. Note that the object recognized by the object recognition device 10A is not limited to an object to be picked. The object recognition device 10A may also recognize objects other than objects to be picked.

[0013] When the object recognition device 10A recognizes objects, the arrangement of the multiple objects is arbitrary. The multiple objects may be, for example, loosely stacked or aligned. Bulk stacking means that the objects are stacked in random positions and orientations. Aligning the multiple objects means that each object is arranged in a predetermined position and in a predetermined orientation.

[0014] The image acquisition unit 11 acquires an image of an object. The image conversion unit 12 converts the image of the object into a converted image in which the edges of the object are replaced with edges whose features are emphasized. The recognition unit 13 recognizes the object based on the converted image.

[0015] In the first embodiment, the image conversion unit 12 converts an image of a photographed object into a converted image in which the edges of the object are replaced with edges in which a feature common to each of the edges of the plurality of objects is emphasized. That is, in the converted image, the common feature present in each of the edges of the plurality of objects is emphasized. Note that the edges in the converted image may be edges in which the edge features of an object having a predetermined shape are emphasized. That is, the image conversion unit 12 may convert an image of a photographed object into a converted image in which the edges of the object are replaced with edges in which a predetermined feature is emphasized.

[0016] Next, a procedure of processing executed by the object recognition device 10A will be described. Fig. 2 is a flowchart showing the procedure of processing executed by the object recognition device 10A according to the first embodiment.

[0017] In step S11, the image acquisition unit 11 acquires an image of an object and outputs the acquired image to the image conversion unit 12.

[0018] In step S12, the image conversion unit 12 performs an image conversion process to convert the image input to the image conversion unit 12. The image conversion unit 12 converts the image into a converted image in which the edges of the objects are replaced with edges that emphasize features common to the edges of the multiple objects. The image conversion unit 12 outputs the converted image to the recognition unit 13.

[0019] In step S13, the recognition unit 13 executes a recognition process to recognize an object based on the converted image input to the recognition unit 13. The recognition unit 13 outputs the object recognition result to the outside of the object recognition device 10A. With the above, the object recognition device 10A ends the process according to the procedure shown in FIG.

[0020] Next, details of processing in each part of the object recognition device 10A will be described. The device that captures an object is a two-dimensional (2D) camera or a three-dimensional (3D) sensor. The image acquisition unit 11 acquires, for example, an image captured by a general-purpose 2D camera or an image captured by a general-purpose 3D sensor. The image acquisition unit 11 acquires an image by, for example, inputting an image captured by a 2D camera or a 3D sensor external to the image acquisition unit 11. Alternatively, the 2D camera or the 3D sensor may be provided in the image acquisition unit 11. In this case, the image acquisition unit 11 acquires an image by the 2D camera or the 3D sensor capturing the image.

[0021] The image acquisition unit 11 acquires 2D data, which is data of a 2D image, using a 2D camera. The 2D image is, for example, a grayscale image or an RGB image, which is a color image. Alternatively, the image acquisition unit 11 acquires 3D data, which is data of a 3D image, using a 3D sensor. The 3D image is, for example, a three-dimensional point cloud or a range image. The following mainly describes a case where the image acquired by the image acquisition unit 11 is a range image. In the following description, the 2D camera may be simply referred to as a camera, and the 3D sensor may be simply referred to as a sensor. The image acquisition unit 11 may acquire an image captured by a single camera or a single sensor, or may acquire multiple images captured by multiple cameras or multiple sensors.

[0022] The image acquisition unit 11 may acquire images captured by multiple cameras with the same measurement method or specifications, or multiple sensors with the same measurement method or specifications. Alternatively, the image acquisition unit 11 may acquire images captured by multiple cameras with different measurement methods or specifications, or multiple sensors with different measurement methods or specifications. An example of sensors with the same measurement method or specifications is when all of the multiple sensors perform measurements using a structured illumination method. An example of sensors with different measurement methods or specifications is when some of the multiple sensors perform measurements using a structured illumination method, and the remaining sensors perform measurements using a time-of-flight (ToF) method.

[0023] The image input from the image acquisition unit 11 to the image conversion unit 12 is a single image or multiple images. When multiple images are input to the image conversion unit 12, the multiple images are, for example, images taken by multiple cameras with the same measurement method or specifications, or images taken by multiple sensors with the same measurement method or specifications. Alternatively, the multiple images may be images taken by multiple cameras with different measurement methods or specifications, or images taken by multiple sensors with different measurement methods or specifications.

[0024] The image conversion unit 12 applies deformation to the edges of objects depicted in the image to emphasize their features. As a result, the image conversion unit 12 replaces the edges of the objects with edges whose features are emphasized. The image conversion unit 12 emphasizes the edge features by simplifying the edges of the objects depicted in the image more than the actual edges. It can also be said that the image conversion unit 12 simplifies the edges of the objects based on their features.

[0025] When an object is an amorphous object, the image conversion unit 12 replaces the edge of each object with an edge that emphasizes its characteristic, thereby reducing the individual difference in the edge of each object. As a first example, the image conversion unit 12 converts an image into a converted image by replacing the edge of the object with an edge that emphasizes a predetermined characteristic. For example, when random irregularities exist in each object, the irregularities contained in the edge are omitted, and the edge of each object is converted into a smooth and simple edge. The image conversion unit 12 replaces the random shape of each object with a primitive shape, which is a simple shape such as a cylinder. In other words, the image conversion unit 12 converts the object captured in the image into a figure that roughly represents the shape of the object.

[0026] It should be noted that the image conversion unit 12 does not necessarily replace the edges of multiple objects with edges of the same shape. The replaced edges for each of the multiple objects may include edges of different shapes. For example, in a case where the edges of each object are replaced with cylinders, the edges may be of shapes with cylinders of different heights, or may be of shapes with cylinders of different radii. Furthermore, the image conversion unit 12 may replace the edges of each object with edges of a shape other than a cylinder. When the objects are fixed-shaped objects, the image conversion unit 12 may replace the edges of the objects with edges that are simplified from the actual edges.

[0027] In a second example, the image conversion unit 12 converts an image into a converted image by replacing the edges of the objects with edges that emphasize features common to the edges of the multiple objects. In this case, the image conversion unit 12 extracts information about the characteristic edges of the multiple objects to determine the features common to the edges of the multiple objects. The image conversion unit 12 converts the image into a converted image based on the features common to the edges of the multiple objects. In the second example, the image conversion unit 12 also replaces the random shapes of each object with a primitive shape, such as a simple cylinder. In the case where the edges of each object are replaced with a cylinder, the cylinder is a shape obtained by eliminating random elements in the edges of each object and can be said to be a feature common to the edges of the multiple objects. Note that, in the second example, the replaced edges for each of the multiple objects may include edges with different shapes. Furthermore, the image conversion unit 12 may replace the edges of each object with edges of a shape other than a cylinder.

[0028] The image conversion unit 12 detects edges by, for example, performing image conversion processing using a known method such as the Canny algorithm on a distance image acquired by a sensor. When performing image conversion processing on each image captured by multiple sensors, the image conversion unit 12 may determine in advance for each sensor a filter value that can facilitate edge detection and detect edges based on the determined filter value. In this case, the image conversion unit 12 extracts characteristic edge information for each edge of multiple objects from the edges acquired in each image and generates a converted image based on the extracted information. By extracting edge information from each image acquired by multiple sensors in this way, the image conversion unit 12 can mutually complement edge information that cannot be obtained from individual images using multiple images.

[0029] The image conversion unit 12 may extract edge information from each of two types of images selected from a 3D point cloud, a distance image, a grayscale image, and a color image. In this case, the image conversion unit 12 extracts characteristic edge information for each edge of a plurality of objects from the edges obtained in each of the two types of images, and generates a converted image based on the extracted information. In this case, the image conversion unit 12 can mutually complement edge information that cannot be obtained from the individual images using the two types of images.

[0030] A method utilizing machine learning may be applied to the image conversion by the image conversion unit 12. The image conversion unit 12 may convert a photographed image of an object into a converted image based on the results of learning between the photographed image of the object and the converted image. In a method utilizing machine learning, for example, a distance image acquired by a sensor and a distance image in which the edges of the object are emphasized may be input to a neural network, and learning may be performed to emphasize the edges of the distance image acquired by the sensor. Algorithms such as GAN (Generative Adversarial Networks) or U-Net may be used for such learning. In addition, learning algorithms such as support vector machines or random forests may be used to learn to estimate edges included in the distance image acquired by the sensor.

[0031] The recognition unit 13 performs recognition processing on the converted image converted by the image conversion unit 12. The recognition unit 13 may calculate the position and orientation of the hand when the robot hand grasps an object, or the position and orientation of the object, through the recognition processing. The recognition unit 13 may calculate the position and orientation of the hand when grasping an object, or the position and orientation of the object, using a neural network. The recognition processing by the recognition unit 13 is not limited to the processing described in embodiment 1. The recognition processing by the recognition unit 13 may be processing other than the processing described in embodiment 1.

[0032] According to the first embodiment, the image conversion unit 12 converts an image into a converted image in which the edges of an object are replaced with edges whose features are enhanced. The recognition unit 13 recognizes an object based on the converted image. The object recognition device 10A generates a converted image in which the edges of an object are replaced with edges whose features are enhanced, thereby enabling the object recognition device 10A to individually separate and recognize an object from among multiple objects. The object recognition device 10A can obtain a converted image in which the individual differences in the edges of each object are reduced and which facilitates individual object recognition, regardless of the type of sensor or camera. As described above, the object recognition device 10A has the effect of improving the object recognition accuracy. By using the object recognition device 10A to recognize objects when picking objects with an industrial robot, the object recognition accuracy can be improved, thereby increasing the success rate of picking by the industrial robot.

[0033] Second Embodiment In the second embodiment, an example will be described in which an image of an object and a converted image are learned. Fig. 3 is a diagram showing the functional configuration of an object recognition device 10B according to the second embodiment. In the second embodiment, the same components as those in the first embodiment are given the same reference numerals, and the configuration different from the first embodiment will be mainly described.

[0034] The object recognition device 10B recognizes an object from a captured image of the object. The object recognition device 10B has a configuration similar to that of the object recognition device 10A shown in Fig. 1. The object recognition device 10B further includes a learning unit 14, an image conversion evaluation unit 15, and a storage unit 16.

[0035] The image conversion unit 12 performs machine learning to generate a converted image in which the edges of an object are replaced with edges whose features are emphasized. The image conversion unit 12 performs machine learning to generate a converted image in which the edges of an object are replaced with edges whose features common to the edges of multiple objects are emphasized. Alternatively, the image conversion unit 12 performs machine learning to generate a converted image in which the edges of an object are replaced with edges whose predetermined features are emphasized.

[0036] The image conversion unit 12 learns image conversion parameters for converting a photographed image of an object into a converted image. The image conversion unit 12 stores the image conversion parameters, which are the learning results, in the storage unit 16. The image conversion unit 12 generates a converted image based on the image conversion parameters. The image conversion unit 12 stores the image conversion parameters in the storage unit 16.

[0037] The image transformation evaluation unit 15 evaluates the performance of the image transformation using the image transformation parameters by the image transformation unit 12. The image transformation evaluation unit 15 determines, through evaluation processing, whether the image transformation executed by the image transformation unit 12 is an image transformation that allows objects to be individually separated and recognized.

[0038] The learning unit 14 learns the image transformation parameters based on the results of the evaluation by the image transformation evaluation unit 15. The learning of the image transformation parameters by the learning unit 14 corresponds to relearning the image transformation parameters learned by the image transformation unit 12.

[0039] The memory unit 16 stores image transformation parameters. The memory unit 16 also stores an image of an object, which is an image acquired by the image acquisition unit 11, and a target transformation image, which is an image targeted for transformation by the image conversion unit 12. The target transformation image is an image in which the edges of the object are emphasized and noise is reduced. The target transformation image may be a CG (Computer Graphics) image in which the edges of the object are simplified based on features common to the edges of multiple objects. Alternatively, the target transformation image may be an image similar to a CG image. In the target transformation image, the shape of the object may be approximated to a primitive shape, which is a simple shape such as a cylinder. In the second embodiment, the post-transformation image may be a CG image or an image similar to a CG image. The target transformation image stored in the memory unit 16 is also utilized as learning data used for learning by the learning unit 14.

[0040] Next, a procedure of processing executed by the object recognition device 10B will be described with reference to Fig. 4. The procedure of processing executed by the object recognition device 10B according to the second embodiment will be described with reference to Fig. 4.

[0041] In step S21, the image acquisition unit 11 acquires an image of an object. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12. The image acquisition unit 11 also stores the acquired image in the storage unit 16.

[0042] In step S22, the image conversion unit 12 executes an image conversion process, which is a process of converting the image input to the image conversion unit 12. In the image conversion process, the image conversion unit 12 learns image conversion parameters for obtaining a converted image from an image of an object. The image conversion unit 12 generates a converted image based on the image input to the image conversion unit 12 and the image conversion parameters that are the learning results. The image conversion unit 12 outputs the converted image to the recognition unit 13.

[0043] In step S23, the recognition unit 13 executes a recognition process to recognize an object based on the converted image input to the recognition unit 13. The recognition unit 13 outputs the object recognition result to the outside of the object recognition device 10B. The recognition unit 13 also outputs the object recognition result to the image conversion evaluation unit 15.

[0044] In step S24, the image transformation evaluation unit 15 performs an evaluation process of the image transformation performed in step S22 based on the result of object recognition by the recognition unit 13. The image transformation evaluation unit 15 determines by the evaluation process whether the image transformation performed in step S22 is an image transformation that allows objects to be individually separated and recognized. Specifically, the image transformation evaluation unit 15 evaluates the appropriateness of the image transformation parameters used in the image transformation process.

[0045] If it is determined that the image transformation executed in step S22 is an image transformation that allows objects to be individually separated and recognized (step S25, Yes), the object recognition device 10B ends the processing according to the procedure shown in FIG.

[0046] On the other hand, if it is determined that the image transformation performed in step S22 is not an image transformation that allows objects to be individually separated and recognized (step S25, No), the image transformation evaluation unit 15 outputs an evaluation result to the learning unit 14 indicating that the image transformation is not an image transformation that allows objects to be individually separated and recognized.

[0047] Next, in step S26, the learning unit 14 executes a learning process to learn image transformation parameters. The learning unit 14 reads the image transformation parameters from the storage unit 16 and learns the image transformation parameters based on the results of evaluation by the image transformation evaluation unit 15. The learning unit 14 stores the image transformation parameters, which are the learning results, in the storage unit 16. After completing step S26, the object recognition device 10B returns the procedure to step S21. The object recognition device 10B repeats the processing according to the procedure from step S21 to step S25.

[0048] Next, details of the processing in each unit of the object recognition device 10B will be described. In the second embodiment, the image conversion unit 12 converts a captured image of an object into a converted image using a machine learning technique. Below, a method using a GAN will be described as an example of an image conversion method using machine learning.

[0049] 5 is a diagram for explaining an example of an image conversion method performed by the object recognition device 10B according to the second embodiment. Fig. 5 shows a conceptual diagram of image conversion processing using GAN. In the example shown here, it is assumed that the image acquired by the image acquisition unit 11 is a distance image.

[0050] The image conversion unit 12 acquires the distance image from the image acquisition unit 11. The image conversion unit 12 also receives a clear image, which is a distance image in which the edges of objects are emphasized and noise is reduced. The clear image is stored, for example, in the storage unit 16. The image conversion unit 12 may acquire the clear image by reading it from the storage unit 16.

[0051] The image conversion unit 12 inputs the distance image and the clear image to the generator G, causing the generator G to output a converted image. In this way, the image conversion unit 12 converts the acquired distance image into a converted image to which emphasized edges are added and noise is reduced.

[0052] The loss function of the generator G is expressed by the following equation (1): G is the loss function of the generator G, L is the labeling algorithm, G is the generator G, D is the discriminator D, x is the distance image acquired from the image acquisition unit 11, x' is the sharp image, α is a weighting coefficient, l is a logarithmic function, θ G are the parameters of the generator G, and θ D represents the parameters of the classifier D. The loss function shown in equation (1) uses a double error function of Least Square Generative Adversarial Networks (LSGAN).

[0053]

[0054] The labeling algorithm may be a known algorithm such as the Watershed algorithm. When distance images are captured by multiple sensors, the clear image input to the generator G may be a distance image captured by a sensor that is most likely to emphasize the edges of objects among the multiple sensors. Alternatively, the clear image input to the generator G may be a pseudo image in which the edges of objects are emphasized and noise is reduced. The pseudo image is an image that is artificially created.

[0055] The image conversion unit 12 calculates the loss function L of the generator G. G θ that minimizes G That is, the image conversion unit 12 obtains θ that satisfies the following equation (2): G Get.

[0056]

[0057] The image conversion unit 12 reads out the target converted image from the storage unit 16. The image conversion unit 12 performs edge detection processing on each of the converted image generated by the generator G and the target converted image. In the example shown in FIG. 5, the target converted image is a CG image. In a CG image, an object is represented by a primitive shape. A known method such as a Laplacian filter or the Canny method may be used for the edge detection processing.

[0058] The image conversion unit 12 inputs the converted image after edge detection and the CG image after edge detection to the classifier D. The image conversion unit 12 performs image conversion so that the smooth edges of the CG image are expressed in the converted image generated by the generator G.

[0059] The loss function of the classifier D for performing image transformation so that the transformed image expresses smooth edges is expressed by the following equation (3). D is the loss function of the classifier D, and y is the CG image.

[0060]

[0061] When edge images are used as both the converted image and the CG image, the amount of information contained in the patch image is extremely small. Here, each of the multiple patches cut out from the distance image, which is each of the converted image and the CG image, is referred to as a patch image. Also, an image in which an object is represented only by edges is referred to as an edge image. The image conversion unit 12 divides the classifier D into a local branch that performs classification for each patch image and a global branch that performs classification for the entire distance image, and then classifies the converted image and the CG image. In this way, the image conversion unit 12 can prevent loss of information contained in the converted image, capture edge characteristics locally and globally, and perform learning to obtain a converted image that can represent the smooth edges of the CG image.

[0062] The image conversion unit 12 calculates the loss function L D θ that minimizes D That is, the image conversion unit 12 obtains θ that satisfies the following equation (4): DGet.

[0063]

[0064] The network configuration, loss function, and parameters used in the machine learning are not limited to those described in the second embodiment, and may be changed as needed. The machine learning performed by the image conversion unit 12 is not limited to machine learning using a GAN, and may be machine learning using a neural network other than a GAN.

[0065] The image conversion unit 12 uses the above-mentioned image conversion process using the GAN to convert the image into a parameter θ G and θ, which is a parameter of the classifier D D In addition, the image transformation parameters learned by the image transformation unit 12 may be weighting coefficients used in the loss function of the neural network, or weighting coefficients set between units constituting the network.

[0066] The image transformation evaluation unit 15 evaluates whether the image transformation executed by the image transformation unit 12 is an image transformation that allows objects to be individually separated and recognized. In other words, the image transformation evaluation unit 15 evaluates the validity of the image transformation from the viewpoint of object recognition.

[0067] Here, an example of evaluation by the image conversion evaluation unit 15 will be described. An image acquired by the image acquisition unit 11 and a labeled image in which correct labels are attached to objects included in the acquired image are prepared. In addition, a labeling process is performed on a converted image obtained by image conversion of the image acquired by the image acquisition unit 11 by the image conversion unit 12. The image conversion evaluation unit 15 determines the accuracy of individual separation by comparing the result of performing the labeling process on the converted image with the true value of the labeled image. The image conversion evaluation unit 15 evaluates the performance of the image conversion based on the accuracy of individual separation.

[0068] An evaluation value when evaluating the performance of image conversion based on the accuracy of individual separation is expressed, for example, by the following formula (5). In formula (5), E represents the evaluation value, T represents the true value of the labeled image, R represents the result of performing the labeling process on the converted image, g(T) represents the center of gravity of the correct region, which is the region of an object in the labeled image, and g(R) represents the center of gravity of the region of an object in the converted image. In formula (5), S(T) represents the area of ​​the correct region, S(R) represents the area of ​​the region of an object in the converted image, α represents a weighting coefficient set for the center of gravity position, β represents a weighting coefficient set for the area of ​​the region of the object, and N represents the number of samples, which is the number of objects included in the acquired image. The first term of the calculation formula following the summation symbol on the right side of formula (5) represents an evaluation value for the center of gravity position. The second term of the calculation formula following the summation symbol on the right side of formula (5) represents an evaluation value for the area of ​​the region.

[0069]

[0070] The image conversion evaluation unit 15 evaluates the performance of the image conversion by comparing the calculated evaluation value with a preset threshold. In other words, the image conversion evaluation unit 15 evaluates the validity of the image conversion by comparing the evaluation value with the threshold. If the evaluation value calculated using equation (5) is less than the threshold, the image conversion evaluation unit 15 evaluates the performance of the image conversion as high. In other words, the image conversion evaluation unit 15 evaluates the image conversion performed by the image conversion unit 12 as an image conversion that allows objects to be individually separated and recognized.

[0071] On the other hand, if the evaluation value calculated by Equation (5) is equal to or greater than the threshold, the image transformation evaluation unit 15 evaluates the performance of the image transformation as low. In other words, the image transformation evaluation unit 15 evaluates that the image transformation performed by the image transformation unit 12 is not an image transformation that allows objects to be individually separated and recognized. Note that performance evaluation based on the evaluation value of the performance of the image transformation may be performed by a method other than the method of comparing the evaluation value with a threshold. The evaluation method used by the image transformation evaluation unit 15 is not limited to the above-described method exemplified in the second embodiment, and may be any method.

[0072] If the image transformation performed by the image transformation unit 12 is evaluated as not being an image transformation that allows objects to be individually separated and recognized, the learning unit 14 learns image transformation parameters. In this way, the learning unit 14 learns image transformation parameters based on the results of evaluation by the image transformation evaluation unit 15.

[0073] The learning unit 14 reads out the image of the object, which is the image acquired by the image acquisition unit 11, and the target transformation image from the storage unit 16. The learning unit 14 learns the image transformation parameters using the image of the object and the target transformation image.

[0074] When a sensor or a camera used to photograph an object is added, the image captured by the added sensor or the added camera is added to the storage unit 16. When an image is added to the storage unit 16, the learning unit 14 may learn image transformation parameters based on the image stored in the storage unit 16 and the target transformation image.

[0075] According to the second embodiment, the image conversion unit 12 converts the image acquired by the image acquisition unit 11 into a converted image based on the results of learning the image acquired by the image acquisition unit 11 and the converted image. The object recognition device 10B can obtain a converted image that reduces individual differences in the edges of each object and facilitates individual object recognition, regardless of the type of sensor or camera. The object recognition device 10B can also obtain a converted image that facilitates individual object recognition depending on the shape or size of the object to be recognized. Furthermore, the object recognition device 10B can obtain a converted image that enables easier individual object recognition by learning image transformation parameters based on the results of evaluation by the image transformation evaluation unit 15. As described above, the object recognition device 10B has the effect of improving the object recognition accuracy. By using the object recognition device 10B to recognize objects when picking objects with an industrial robot, the object recognition accuracy can be improved, thereby increasing the success rate of picking by the industrial robot.

[0076] Third Embodiment In a third embodiment, an example will be described in which learning data used in machine learning is automatically collected by simulation. Fig. 6 is a diagram showing the functional configuration of an object recognition device 10C according to the third embodiment. In the third embodiment, the same components as those in the first or second embodiment are denoted by the same reference numerals, and the configuration different from the first or second embodiment will be mainly described.

[0077] The object recognition device 10C recognizes an object from a captured image of the object. The object recognition device 10C has a configuration similar to that of the object recognition device 10B shown in Fig. 3. The object recognition device 10C further includes a simulation condition setting unit 17, a scene generation unit 18, and a data set generation unit 19. Note that in the third embodiment, simulation refers to simulating a scene in which an object is captured.

[0078] The simulation condition setting unit 17 sets simulation conditions for simulating the photographing of an object, including at least one of photographing device information, which is information about a device that photographs an image, and object information, which is information about the object. The photographing device information is information about a camera that photographs an image, or information about a sensor that photographs an image.

[0079] The scene generation unit 18 simulates a scene in which an object is photographed, based on the simulation conditions set by the simulation condition setting unit 17. The dataset generation unit 19 generates a dataset to be used for learning, based on the scene generated by the scene generation unit 18. The dataset generated by the dataset generation unit 19 is learning data to be used for machine learning. The learning unit 14 learns image transformation parameters using the dataset generated by the dataset generation unit 19.

[0080] A computer-aided design (CAD) model of an object is stored in the storage unit 16. The CAD model stored in the storage unit 16 is at least one of a 3D CAD model and a 2D CAD model. If the object is amorphous or if a 3D CAD model of the object does not exist, a model generated by a method such as mesh generation based on data obtained by measuring the object using a sensor or the like may be stored as the 3D CAD model. In the following description, the CAD model may be simply referred to as a model.

[0081] Next, a procedure of processing executed by the object recognition device 10C will be described. Fig. 7 is a flowchart showing the procedure of processing executed by the object recognition device 10C according to the third embodiment.

[0082] In step S31, the simulation condition setting unit 17 sets simulation conditions. The simulation condition setting unit 17 sets simulation conditions including at least one of imaging device information and object information. The simulation condition setting unit 17 reads a model from the storage unit 16 and sets simulation conditions based on the model. The simulation condition setting unit 17 outputs information on the simulation conditions to the scene generation unit 18.

[0083] In step S32, the scene generation unit 18 simulates a scene in which an object is photographed, based on the simulation conditions set by the simulation condition setting unit 17. The scene generation unit 18 reads a model from the storage unit 16. The scene generation unit 18 generates a scene based on the simulation conditions and the model. The scene generation unit 18 outputs information indicating the generated scene to the data set generation unit 19.

[0084] In step S33, the dataset generation unit 19 generates a dataset based on the scene generated by the scene generation unit 18. The dataset generation unit 19 outputs the generated dataset to the learning unit 14.

[0085] In step S34, the image acquisition unit 11 acquires an image of the object. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12. The image acquisition unit 11 also stores the acquired image in the storage unit 16.

[0086] In step S35, the image conversion unit 12 executes an image conversion process, which is a process of converting the image input to the image conversion unit 12. In the image conversion process, the image conversion unit 12 learns image conversion parameters for obtaining a converted image from an image of an object. The image conversion unit 12 generates a converted image based on the image input to the image conversion unit 12 and the image conversion parameters that are the learning results. The image conversion unit 12 outputs the converted image to the recognition unit 13.

[0087] In step S36, the recognition unit 13 executes a recognition process to recognize an object based on the converted image input to the recognition unit 13. The recognition unit 13 outputs the object recognition result to the outside of the object recognition device 10C. The recognition unit 13 also outputs the object recognition result to the image conversion evaluation unit 15.

[0088] In step S37, the image transformation evaluation unit 15 performs an evaluation process of the image transformation performed in step S35 based on the result of object recognition by the recognition unit 13. The image transformation evaluation unit 15 determines by the evaluation process whether the image transformation performed in step S35 is an image transformation that allows objects to be individually separated and recognized. Specifically, the image transformation evaluation unit 15 evaluates the appropriateness of the image transformation parameters used in the image transformation process.

[0089] If it is determined that the image transformation executed in step S35 is an image transformation that allows objects to be individually separated and recognized (step S38, Yes), the object recognition device 10C ends the processing according to the procedure shown in FIG.

[0090] On the other hand, if it is determined that the image transformation performed in step S35 is not an image transformation that allows objects to be individually separated and recognized (step S38, No), the image transformation evaluation unit 15 outputs an evaluation result to the learning unit 14 indicating that the image transformation is not an image transformation that allows objects to be individually separated and recognized.

[0091] Next, in step S39, the learning unit 14 executes a learning process to learn image transformation parameters. The learning unit 14 learns the image transformation parameters based on the data set input to the learning unit 14 and the results of the evaluation by the image transformation evaluation unit 15. The learning unit 14 stores the image transformation parameters, which are the learning results, in the storage unit 16. After completing step S39, the object recognition device 10C returns the procedure to step S34. The object recognition device 10C repeats the processing according to the procedure from step S34 to step S38.

[0092] Next, a detailed description will be given of the functions of the simulation condition setting unit 17, the scene generation unit 18, and the data set generation unit 19. Fig. 8 is a diagram illustrating an example of the configuration of the simulation condition setting unit 17 included in an object recognition device 10C according to the third embodiment.

[0093] The simulation condition setting unit 17 includes an image capture device information setting unit 21, an object information setting unit 22, and an environment information setting unit 23. The simulation condition setting unit 17 sets simulation conditions including image capture device information, object information, and environment information.

[0094] The imaging device information setting unit 21 sets imaging device information including at least one of information about the specifications of the device that captures the image and information about the installation of the device. The information about the device specifications includes parameters indicating the angle of view, working distance, installation angle, resolution, etc. The information about the installation of the device includes parameters indicating the position of the device in the simulation. The imaging device information setting unit 21 may read a model from the storage unit 16 and confirm that an object fits within the field of view of the device in the simulation.

[0095] The object information setting unit 22 sets object information. The object information includes information about the color or size of the object. The object information setting unit 22 reads a model from the storage unit 16 and sets object information based on the model. The object information set by the object information setting unit 22 may include information about the number of models used in the scene generated by the scene generation unit 18.

[0096] The environmental information setting unit 23 sets environmental information. The environmental information includes information about a light source, such as information about diffuse reflection or ambient light. The environmental information may include information indicating the size of the box in which the objects are stacked in bulk, or information indicating the method by which the objects are supplied to the location where they are picked up. Here, the supply method refers to the state in which the objects are arranged, such as a bulk stacked state or an aligned state. The environmental information setting unit 23 may read a model from the storage unit 16 and confirm the supply method in a simulation based on the model.

[0097] The simulation condition setting unit 17 outputs information on simulation conditions including photographing device information, object information, and environmental information to the scene generation unit 18. Note that the simulation conditions set by the simulation condition setting unit 17 are not limited to those including photographing device information, object information, and environmental information. The simulation conditions only need to include at least one of photographing device information and object information. The simulation condition setting unit 17 is not limited to those including all of the photographing device information setting unit 21, object information setting unit 22, and environmental information setting unit 23. The simulation condition setting unit 17 only needs to include at least one of the photographing device information setting unit 21 and the object information setting unit 22.

[0098] The scene generated by the scene generation unit 18 can be said to be a situation in which an object is photographed. For example, when objects are picked one by one from a pile of multiple objects, the scene is a situation in which the objects are arranged in random positions and orientations. The scene generation unit 18 reads a model from the storage unit 16 and simulates the scene using the simulation conditions and the model.

[0099] The scene generation unit 18 may generate a scene using multiple models. When an object is an amorphous object, the multiple models used to generate a scene may include models of the same shape or models of different shapes. For example, the scene generation unit 18 may read 10 models of different shapes from the storage unit 16 and generate five copies of each model, thereby using 50 models. In this way, by using models of different shapes to generate a scene, the object recognition device 10C can perform image conversion for each of multiple objects with irregular shapes.

[0100] 9 is a diagram illustrating a configuration example of a data set generation unit 19 included in an object recognition device 10C according to Embodiment 3. The data set generation unit 19 includes a 2D data generation unit 24, a 3D data generation unit 25, an annotation data generation unit 26, and a target transformed image generation unit 27.

[0101] The 2D data generation unit 24 and the 3D data generation unit 25 function as image data generation units that generate image data. The 2D data generation unit 24 generates 2D data, which is image data representing the scene generated by the scene generation unit 18. The 2D data generated by the 2D data generation unit 24 is image data such as a grayscale image or a color image. The 2D data generation unit 24 outputs the generated 2D data to the learning unit 14. The learning unit 14 may use the 2D data input to the learning unit 14 to learn image transformation parameters.

[0102] The 3D data generation unit 25 generates 3D data, which is image data representing the scene generated by the scene generation unit 18. The 3D data generated by the 3D data generation unit 25 is image data such as a three-dimensional point cloud or a range image. The 3D data generation unit 25 outputs the generated 3D data to the learning unit 14. The learning unit 14 may use the 3D data input to the learning unit 14 to learn image conversion parameters.

[0103] The 2D data generation unit 24 may store the generated 2D data in the storage unit 16. The 3D data generation unit 25 may store the generated 3D data in the storage unit 16. The image conversion unit 12 may use the 2D data or 3D data, which is image data stored in the storage unit 16, for machine learning in the image conversion process.

[0104] The image data generated by the image data generation unit may be at least one of 2D data representing the scene generated by the scene generation unit 18 and 3D data representing the scene generated by the scene generation unit 18. The data set generation unit 19 is not limited to being equipped with both the 2D data generation unit 24 and the 3D data generation unit 25. It is sufficient for the data set generation unit 19 to be equipped with at least one of the 2D data generation unit 24 and the 3D data generation unit 25.

[0105] The annotation data generation unit 26 generates annotation data to be added to the image data generated by the image data generation unit. The annotation data includes labeling data that indicates the color when an object is colored, information indicating the position and posture of the object, and type name information. When an object is colored, for example, when multiple types of objects are included in multiple objects, the objects are color-coded according to their type. The type name is a name that indicates the type of object. The annotation data generation unit 26 outputs the generated annotation data to the learning unit 14. The learning unit 14 may use the annotation data input to the learning unit 14 to learn image conversion parameters.

[0106] The target transformation image generation unit 27 generates a target transformation image, which is an image that is a target for image transformation from image data. The target transformation image is a grayscale image, a color image, or a distance image, and is an image in which the edges of an object are emphasized. The target transformation image may be a CG image in which the edges of an object are simplified based on features common to the edges of multiple objects. The target transformation image may also be a CG image in which the edges of an object are replaced with edges in which predetermined features are emphasized. The target transformation image may be an image similar to these CG images. In the target transformation image, the shape of the object may be approximated to a primitive shape. The target transformation image generation unit 27 stores the generated target transformation image in the memory unit 16.

[0107] The learning unit 14 may read out the target transformed image generated by the target transformed image generation unit 27 from the storage unit 16, and learn the image transformation parameters using the image data and the target transformed image input to the learning unit 14. The image transformation unit 12 may read out the image data generated by the image data generation unit and the target transformed image generated by the target transformed image generation unit 27 from the storage unit 16, and perform image transformation processing using the image data and the target transformed image.

[0108] According to the above description, the dataset generated by the dataset generation unit 19 includes image data, annotation data, and a target transformed image. The dataset generated by the dataset generation unit 19 is not limited to one including image data, annotation data, and a target transformed image. The dataset generated by the dataset generation unit 19 does not have to include at least one of the image data, annotation data, and target transformed image.

[0109] The data generated by the dataset generating unit 19 is data generated by simulation. Therefore, it is conceivable that differences in the level of noise and the like will occur between the image data generated by the dataset generating unit 19 and the image that is actually captured. The dataset generating unit 19 may be provided with a function for reproducing noise measured in advance in the image data generated by the dataset generating unit 19. The dataset generating unit 19 may reproduce the measured noise in accordance with the measurement method of the sensor, such as the structured illumination method or the ToF method.

[0110] The image conversion unit 12 may input to the neural network the distance image, which is the image data generated by the image data generation unit, and the CG image, which is the target transformed image generated by the target transformed image generation unit 27. In this way, the image conversion unit 12 may learn image transformation parameters such that the smooth edges of the CG image are expressed in the transformed image.

[0111] The image conversion evaluation unit 15 may calculate an evaluation value using annotation data, which is labeling data attached to a correct answer region, which is a region of an object in the image data generated by the image data generation unit. The evaluation value is an evaluation value when evaluating the performance of image conversion based on the accuracy of individual separation. The image conversion unit 12 may learn image conversion parameters by reflecting the evaluation value in a loss function used for machine learning in the image conversion unit 12.

[0112] It should be noted that the method of utilizing the data generated by the dataset generation unit 19 for learning is not limited to the method described in embodiment 3. The object recognition device 10C may appropriately increase the variety of methods of utilizing the data generated by the dataset generation unit 19 in addition to the method described in embodiment 3.

[0113] According to the third embodiment, the object recognition device 10C includes a simulation condition setting unit 17 that sets simulation conditions, a scene generation unit 18 that simulates a scene, and a dataset generation unit 19 that generates a dataset used for learning image transformation parameters. The object recognition device 10C can automatically collect learning data by simulating an actual environment in which an object is photographed. The object recognition device 10C can reduce the number of personnel and costs required to collect learning data, and can obtain learning data that enables highly accurate learning. As described above, the object recognition device 10C has the effects of enabling labor savings, cost reduction, and improved object recognition accuracy.

[0114] Using the object recognition device 10C to recognize objects during picking by an industrial robot can reduce the labor required at the site where the industrial robot is operated, shorten the time required for verification before the industrial robot is put into operation, and reduce the cost required for verification before operation. Simulating the actual environment in which an object is photographed using the object recognition device 10C can improve the reproducibility of the simulation process before the industrial robot is introduced. This allows for accurate determination of the placement of the industrial robot or confirmation of the operation of the industrial robot.

[0115] Fourth Embodiment In the fourth embodiment, an example will be described in which a control device equipped with an object recognition device automatically adjusts picking-related parameters. FIG. 10 is a diagram showing the functional configuration of a control device 60 according to the fourth embodiment. In the fourth embodiment, the same components as those in the first to third embodiments are given the same reference numerals, and the configuration different from the first to third embodiments will be mainly described. In the following description, an industrial robot will be simply referred to as a robot.

[0116] The robot includes a robot main body 62 and a control device 60 that controls the robot main body 62. Fig. 10 shows the control device 60 and the robot main body 62. The control device 60 includes an object recognition device 10D and a robot control unit 61. The object recognition device 10D recognizes an object from a photographed image of the object. The object recognition device 10D outputs the object recognition result to the robot control unit 61. The robot control unit 61 controls the robot main body 62 based on the object recognition result obtained by the object recognition device 10D.

[0117] The robot main body 62 has a hand that grasps an object. The hand is not shown. The robot control unit 61 causes the robot main body 62 to perform an operation to grasp an object based on the grasping position detected by the object recognition device 10D. The robot main body 62 grasps the object and picks it up. The robot repeats this operation to pick up objects one by one from a plurality of objects.

[0118] As an example, the robot control unit 61 picks each of the multiple objects based on the results of object recognition by the object recognition device 10D. Note that the method for determining the picking order for each of the multiple objects is not limited to that described in embodiment 4. The robot control unit 61 may pick each of the multiple objects stacked vertically in order from the highest object. Alternatively, the picking order may be determined based on the likelihood and the position. In other words, the robot control unit 61 may pick each of the multiple objects in order from the object with the highest likelihood and the object with the highest position. The criteria for determining the picking order may be changed as necessary.

[0119] A force sensor, a proximity sensor, a tactile sensor, or the like may be attached to the robot body 62. A plurality of objects may be picked up by a single robot or by a plurality of robots.

[0120] The object recognition device 10D has a configuration similar to that of the object recognition device 10C shown in FIG. 6. Furthermore, the object recognition device 10D includes a parameter adjustment unit 31 that adjusts parameters related to picking. The parameter adjustment unit 31 adjusts gripping parameters, which are parameters related to the hand used for picking. The parameter adjustment unit 31 also adjusts recognition parameters, which are parameters used in the recognition process during picking. The parameter adjustment unit 31 stores the adjusted gripping parameters and the adjusted recognition parameters in the storage unit 16. The object recognition device 10D automatically adjusts the picking-related parameters through simulation.

[0121] 11 is a diagram illustrating an example of the configuration of the parameter adjustment unit 31 included in the control device 60 according to the fourth embodiment. FIG. 11 illustrates the storage unit 16, the data set generation unit 19, and the parameter adjustment unit 31. The parameter adjustment unit 31 includes a grip parameter adjustment unit 32 and a recognition parameter adjustment unit 33. The grip parameter adjustment unit 32 adjusts grip parameters, which are parameters for the hand of the robot main body 62 that grips an object. The recognition parameter adjustment unit 33 adjusts recognition parameters, which are parameters used in object recognition by the recognition unit 13.

[0122] A CAD model of the object is stored in the storage unit 16. The grip parameter adjustment unit 32 reads the CAD model of the object from the storage unit 16 and adjusts grip parameters using the CAD model of the object. The grip parameter adjustment unit 32 stores the adjusted grip parameters in the storage unit 16.

[0123] The hands possessed by the robot main body 62 may be tweezers hands, parallel hands, suction pads, or the like. The gripping parameter adjustment unit 32 adjusts gripping parameters for these various hands. The gripping parameters include parameters such as the hand opening width or hand claw width. The gripping parameter adjustment unit 32 adjusts the parameter values ​​such as the hand opening width or claw width to values ​​that make it easier to grip an object. The type of hand used to grip an object may be determined by the user of the robot.

[0124] The grip parameters include multiple types of parameters for the hand, such as an opening width parameter or a claw width parameter. In the following description, the grip parameter refers to each of the multiple types of parameters. In the following description, the multiple types of grip parameters are also referred to as multiple grip parameters. In the fourth embodiment, the grip parameter adjustment unit 32 adjusts each of the multiple grip parameters by determining the value of each of the multiple grip parameters.

[0125] The grip parameter adjustment unit 32 may adjust all of the grip parameters simultaneously, or may adjust the grip parameters individually in a predetermined order. In the following description, it is assumed that the grip parameter adjustment unit 32 adjusts the grip parameters individually in a predetermined order.

[0126] The recognition parameter adjustment unit 33 reads the CAD model of the object and the grip parameters adjusted by the grip parameter adjustment unit 32 from the storage unit 16. The dataset generated by the dataset generation unit 19 is input to the recognition parameter adjustment unit 33. The recognition parameter adjustment unit 33 adjusts the recognition parameters using the CAD model, the grip parameters, and the dataset. The recognition parameter adjustment unit 33 saves the adjusted recognition parameters in the storage unit 16.

[0127] The recognition parameter adjustment unit 33 adjusts, for example, recognition parameters, which are parameters for detecting grip position candidates or the position and orientation of each object from 2D data or 3D data. Alternatively, the recognition parameter adjustment unit 33 may adjust parameters of a learning model for detecting grip positions through machine learning. The recognition parameter adjustment unit 33 may adjust parameters of a learning model in machine learning used in auxiliary processing for improving the recognition rate. An example of auxiliary processing for improving the recognition rate is segmentation. In the following description, it is assumed that the recognition parameter adjustment unit 33 mainly adjusts parameters for detecting grip position candidates from 2D data.

[0128] The recognition parameters include multiple types of parameters used in object recognition by the recognition unit 13. In the following description, the recognition parameter refers to each of the multiple types of parameters. In the following description, the multiple types of recognition parameters are also referred to as multiple recognition parameters. In the fourth embodiment, the recognition parameter adjustment unit 33 adjusts each of the multiple recognition parameters by determining the value of each of the multiple recognition parameters.

[0129] The recognition parameter adjustment unit 33 may adjust all of the recognition parameters simultaneously, or may adjust the recognition parameters individually in a predetermined order. In the following description, it is assumed that the recognition parameter adjustment unit 33 adjusts the recognition parameters individually in a predetermined order.

[0130] Next, the details of the function of the grip parameter adjustment unit 32 will be described. Fig. 12 is a diagram showing an example of the configuration of the grip parameter adjustment unit 32 included in the control device 60 according to the fourth embodiment. Fig. 12 shows the storage unit 16, the data set generation unit 19, and the grip parameter adjustment unit 32 and the recognition parameter adjustment unit 33 of the parameter adjustment unit 31.

[0131] The grip parameter adjustment unit 32 includes a grip parameter adjustment range determination unit 41, a grip parameter change unit 42, a model rotation unit 43, a grip evaluation unit 44, a grip parameter value determination unit 45, and a grip parameter adjustment completion determination unit 46.

[0132] The grip parameter adjustment range determination unit 41 determines an adjustment range for each of a plurality of grip parameters. The grip parameter adjustment range determination unit 41 determines the adjustment range based on a model stored in the storage unit 16, etc. The grip parameter adjustment unit 32 adjusts the value of each of the plurality of grip parameters within the value range determined by the grip parameter adjustment range determination unit 41. In other words, the adjustment range of a grip parameter refers to a range within which the value of a grip parameter is adjusted. The grip parameter adjustment range determination unit 41 outputs information indicating the determined adjustment range to the grip parameter change unit 42.

[0133] The grip parameter change unit 42 changes the grip parameter to be adjusted from among the plurality of grip parameters. The grip parameter adjustment unit 32 adjusts each grip parameter by sequentially switching among the plurality of grip parameters in the grip parameter change unit 42. In this way, the grip parameter adjustment unit 32 individually adjusts the plurality of grip parameters in a predetermined order. The grip parameter change unit 42 outputs information indicating the grip parameter to be adjusted and information indicating the adjustment range of the grip parameter to be adjusted to the grip evaluation unit 44 via the model rotation unit 43.

[0134] The model rotation unit 43 rotates the model of the object. The model rotation unit 43 reads out the model stored in the storage unit 16 and executes a process of rotating the read out model.

[0135] The grip evaluation unit 44 evaluates whether the object will be grasped properly when grasped by the hand, assuming that the object will assume the same posture as the model rotated by the model rotation unit 43. The grip evaluation unit 44 evaluates whether the object will be grasped properly when grasped while adjusting the value of the grip parameter within the adjustment range, based on information indicating the grip parameter to be adjusted and information indicating the adjustment range of the grip parameter to be adjusted. The grip evaluation unit 44 outputs the result of the evaluation of whether the object will be grasped properly to the grip parameter value determination unit 45.

[0136] The grip parameter value determination unit 45 determines the value of each of the plurality of grip parameters based on the evaluation result by the grip evaluation unit 44. When the grip parameter value determination unit 45 determines the value of the grip parameter, it outputs the determined value to the grip parameter adjustment completion determination unit 46.

[0137] The grip parameter adjustment completion determination unit 46 determines whether adjustment of all of the plurality of grip parameters has been completed. The grip parameter adjustment completion determination unit 46 determines that adjustment of all of the plurality of grip parameters has been completed when the grip parameter values ​​have been determined for all of the plurality of grip parameters.

[0138] FIG. 13 is a flowchart illustrating a procedure of processing executed by the grip parameter adjustment unit 32 of the control device 60 according to the fourth embodiment.

[0139] In step S41, the grip parameter adjustment range determination unit 41 determines the adjustment range of the grip parameter. The grip parameter adjustment range determination unit 41 reads out the model stored in the storage unit 16 and determines the adjustment range of the grip parameter based on the size of the model, etc. The grip parameter adjustment range determination unit 41 may also determine the initial value of the grip parameter based on the size of the model, etc. The initial value of the grip parameter is a value that serves as a reference when adjusting the value of the grip parameter.

[0140] In step S42, the grip parameter change unit 42 changes the type of grip parameter to be adjusted. That is, the grip parameter change unit 42 changes the grip parameter to be adjusted from among the multiple grip parameters.

[0141] In step S43, the model rotation unit 43 rotates the object model. For example, the model rotation unit 43 rotates the model randomly. The model rotation unit 43 may also rotate the model in a predetermined order.

[0142] In step S44, the grip evaluation unit 44 evaluates whether the grip is good or bad when the object is gripped by the hand. The grip evaluation unit 44 applies the hand model to the model rotated by the model rotation unit 43, and assumes a situation in which the hand grips an object in the same posture as the model, thereby evaluating whether the grip is good or bad. The grip evaluation unit 44 evaluates whether the grip is good or bad, for example, by calculating an evaluation function of the grip parameters. The evaluation function of the grip parameters is, for example, the frequency of misalignment F Ang and interference frequency F Col The deviation occurrence frequency F Ang represents the frequency of occurrence of misalignment between the object's main axis and the gripping direction. Col represents the frequency of interference between the object and the hand.

[0143] If the model rotation by the model rotation unit 43 has not finished (step S45, No), the grip parameter adjustment unit 32 returns the procedure to step S43. On the other hand, if the model rotation by the model rotation unit 43 has finished (step S45, Yes), in step S46, the grip parameter value determination unit 45 determines the value of the grip parameter. The grip parameter value determination unit 45 determines the value of the grip parameter based on the result of the evaluation in step S44. The grip parameter value determination unit 45, for example, searches for the value of the grip parameter that minimizes the evaluation function of the grip parameter. The grip parameter value determination unit 45 determines the value of the grip parameter that minimizes the evaluation function of the grip parameter as the adjusted value of the grip parameter.

[0144] Next, the grip parameter adjustment completion determination unit 46 determines whether adjustment of all grip parameters to be adjusted in the grip parameter adjustment unit 32 has been completed. If adjustment of all grip parameters to be adjusted has not been completed (step S47, No), the grip parameter adjustment unit 32 returns the procedure to step S42. The grip parameter adjustment unit 32 executes the processes from step S42 to step S47 for grip parameters for which adjustment has not been completed. On the other hand, if adjustment of all grip parameters to be adjusted has been completed (step S47, Yes), the grip parameter adjustment unit 32 ends the processing according to the procedure shown in FIG.

[0145] Next, a detailed description will be given of the function of the recognition parameter adjustment unit 33. Fig. 14 is a diagram showing an example of the configuration of the recognition parameter adjustment unit 33 included in the control device 60 according to the fourth embodiment. Fig. 14 shows the storage unit 16, the data set generation unit 19, and the grip parameter adjustment unit 32 and the recognition parameter adjustment unit 33 of the parameter adjustment unit 31.

[0146] The recognition parameter adjustment unit 33 includes a recognition parameter adjustment range determination unit 51 , a recognition parameter change unit 52 , a recognition trial unit 53 , a recognition evaluation unit 54 , a recognition parameter value determination unit 55 , and a recognition parameter adjustment end determination unit 56 .

[0147] The recognition parameter adjustment range determination unit 51 determines an adjustment range for each of a plurality of recognition parameters. The recognition parameter adjustment range determination unit 51 determines the adjustment range based on a model stored in the storage unit 16, etc. The recognition parameter adjustment unit 33 adjusts the value of each of the plurality of recognition parameters within the value range determined by the recognition parameter adjustment range determination unit 51. In other words, the recognition parameter adjustment range refers to a range within which the value of a recognition parameter is adjusted. The recognition parameter adjustment range determination unit 51 outputs information indicating the determined adjustment range to the recognition parameter change unit 52.

[0148] The recognition parameter change unit 52 changes the recognition parameter to be adjusted from among the plurality of recognition parameters. The recognition parameter adjustment unit 33 adjusts each recognition parameter by sequentially switching among the plurality of recognition parameters in the recognition parameter change unit 52. In this way, the recognition parameter adjustment unit 33 individually adjusts the plurality of recognition parameters in a predetermined order. The recognition parameter change unit 52 outputs information indicating the recognition parameter to be adjusted and information indicating the adjustment range of the recognition parameter to be adjusted to the recognition trial unit 53.

[0149] The recognition trial unit 53 attempts a recognition process using the grip parameters adjusted by the grip parameter adjustment unit 32. The recognition trial unit 53 reads the grip parameters adjusted by the grip parameter adjustment unit 32 from the storage unit 16. The recognition trial unit 53 acquires a dataset from the dataset generation unit 19. The recognition trial unit 53 detects candidate grip positions by attempting a recognition process using the grip parameters on image data included in the dataset. The recognition trial unit 53 detects candidate grip positions while adjusting the values ​​of the recognition parameters within the adjustment range based on information indicating the recognition parameters to be adjusted and information indicating the adjustment range of the recognition parameters to be adjusted. The recognition trial unit 53 outputs the detection result of the candidate grip positions, which is the result of the attempted recognition process, to the recognition evaluation unit 54.

[0150] The recognition evaluation unit 54 evaluates the quality of the result of the recognition process attempted by the recognition attempt unit 53. The recognition evaluation unit 54 evaluates the quality of the result of the recognition process attempted by, for example, evaluating the suitability of the candidate grip positions detected by the recognition attempt unit 53 as grip positions. The recognition evaluation unit 54 outputs the result of the evaluation of the quality of the result of the recognition process attempted to the recognition parameter value determination unit 55.

[0151] The recognition parameter value determination unit 55 determines the values ​​of each of the plurality of recognition parameters based on the evaluation results by the recognition evaluation unit 54. After determining the values ​​of the recognition parameters, the recognition parameter value determination unit 55 outputs the determined values ​​to the recognition parameter adjustment completion determination unit 56.

[0152] The recognition parameter adjustment completion determination unit 56 determines whether or not the adjustment of all of the plurality of recognition parameters has been completed. When the recognition parameter values ​​for all of the plurality of recognition parameters have been determined, the recognition parameter adjustment completion determination unit 56 determines that the adjustment of all of the plurality of recognition parameters has been completed.

[0153] FIG. 15 is a flowchart illustrating a procedure of processing executed by the recognition parameter adjustment unit 33 of the control device 60 according to the fourth embodiment.

[0154] In step S51, the recognition parameter adjustment range determination unit 51 determines the adjustment range of the recognition parameter. The recognition parameter adjustment range determination unit 51 reads out a model stored in the storage unit 16 and determines the adjustment range of the recognition parameter based on the size of the model, etc. The recognition parameter adjustment range determination unit 51 may also determine the initial value of the recognition parameter based on the size of the model, etc. The initial value of the recognition parameter is a value that serves as a reference when adjusting the value of the recognition parameter.

[0155] In step S52, the recognition parameter change unit 52 changes the type of recognition parameter to be adjusted. That is, the recognition parameter change unit 52 changes the recognition parameter to be adjusted from among the plurality of recognition parameters.

[0156] In step S53, the recognition trial unit 53 attempts a recognition process using the grip parameters adjusted by the grip parameter adjustment unit 32. The recognition trial unit 53 adjusts the values ​​of the recognition parameters within the adjustment range and detects candidate grip positions for each value of the recognition parameters. If the recognition trial process has not been completed for each value of the recognition parameters within the adjustment range (No in step S54), the recognition trial unit 53 returns the procedure to step S53 and continues adjusting the values ​​of the recognition parameters and attempting the recognition process.

[0157] On the other hand, when the trial of the recognition process has been completed for each value of the recognition parameter within the adjustment range (Yes in step S54), in step S55, the recognition evaluation unit 54 evaluates the quality of the result of the trial of the recognition process in step S53. The recognition evaluation unit 54 evaluates the quality of the result of the trial of the recognition process, for example, by calculating an evaluation function of the recognition parameter.

[0158] The evaluation function of the recognition parameters is, for example, the deviation E between the candidate gripping position detected in step S53 and the optimal gripping position of the object. Pos and the deviation E between the object's main axis and the gripping direction Ang and the evaluation value E of the number of recognized objects. Num The number of recognized objects is the number of objects recognized by the object recognition device 10D. The evaluation value E of the number of recognized objects Num is an evaluation value for the number of recognized objects. In picking by a robot, it is required that no objects to be picked are left behind. When the number of recognized objects is small, there are many objects that are not recognized among the objects to be picked, and it is easy for objects to be left behind. Therefore, the evaluation value E Num On the other hand, when the number of recognized objects is large, more objects to be picked are recognized and it is difficult for objects to be left behind. Therefore, the evaluation value E Num The optimum grip position of an object is, for example, a position determined based on the center of gravity of the object. The optimum grip position of an object may also be a position arbitrarily determined by the user.

[0159] Next, in step S56, the recognition parameter value determination unit 55 determines the value of the recognition parameter. The recognition parameter value determination unit 55 determines the value of the recognition parameter based on the result of the evaluation in step S54. For example, the recognition parameter value determination unit 55 searches for the value of the recognition parameter that minimizes the evaluation function of the recognition parameter. The recognition parameter value determination unit 55 determines the value of the recognition parameter that minimizes the evaluation function of the recognition parameter as the adjusted value of the recognition parameter.

[0160] Next, the recognition parameter adjustment completion determination unit 56 determines whether or not adjustment of all recognition parameters to be adjusted by the recognition parameter adjustment unit 33 has been completed. If adjustment of all recognition parameters to be adjusted has not been completed (step S57, No), the recognition parameter adjustment unit 33 returns the procedure to step S52. The recognition parameter adjustment unit 33 executes the processes from step S52 to step S57 for the recognition parameters for which adjustment has not been completed. On the other hand, if adjustment of all recognition parameters to be adjusted has been completed (step S57, Yes), the recognition parameter adjustment unit 33 ends the process shown in FIG. 15 .

[0161] 10 may use the recognition parameters adjusted by the recognition parameter adjustment unit 33 when detecting candidates for gripping positions from a transformed image obtained by transforming a photographed image of an object. The recognition unit 13 can read out the recognition parameters adjusted by the recognition parameter adjustment unit 33 from the storage unit 16 and detect candidates for gripping positions using the read-out recognition parameters.

[0162] 10 may evaluate the performance of image conversion using the image conversion parameters based on the operation results of the robot main body 62. The image conversion evaluation unit 15 calculates an evaluation value for evaluating the performance of image conversion based on the operation results when the robot main body 62 is actually operated.

[0163] For example, by having the robot body 62 pick up an object, it is possible to observe the probability that the robot body 62 succeeds in grasping the object, the operation time taken by the robot body 62 to perform the operation of grasping the object, and the cause of the robot body 62 failing to grasp the object. The image conversion evaluation unit 15 evaluates the performance of the image conversion based on the operation results including information on the probability that the robot body 62 succeeds in grasping the object, information on the operation time taken by the robot body 62 to perform the operation of grasping the object, and information indicating the cause of the robot body 62 failing to grasp the object.

[0164] The image conversion evaluation unit 15 calculates an evaluation value for evaluating the performance of the image conversion, for example, using the following formula (6): In formula (6), E is the evaluation value for evaluating the performance of the image conversion, P is the grasping success rate, which is the probability that an object is successfully grasped, T is the operation time during which the robot main body 62 performs the operation of grasping the object, and f1, f2, ... each represent the cause of the robot main body 62 failing to grasp the object. In formula (6), W1 is a weighting coefficient set for the grasping success rate, W2 is a weighting coefficient set for the operation time, W 31 , W 32 , . . . represent weighting coefficients set for each of f1, f2, .

[0165]

[0166] W is a weighting factor set for the cause of grasping failure 31 , W 32 Each value of , . . . may be set to a small value for hardware-related causes such as a drop during transport, and to a large value for causes related to the results of image conversion.

[0167] The operation results used for evaluation in the image conversion evaluation unit 15 may include at least one of information on the probability that the robot body 62 succeeds in grasping the object, information on the operation time during which the robot body 62 performs the operation to grasp the object, and information indicating the cause of the robot body 62 failing to grasp the object. The operation results used for evaluation in the image conversion evaluation unit 15 may also include information other than these. The method of calculating the evaluation value for evaluating the performance of image conversion is not limited to the method described in the fourth embodiment. The image conversion evaluation unit 15 may evaluate the performance of image conversion based on a combination of the operation results of the robot body 62 and the accuracy of individual separation. The image conversion evaluation unit 15 may evaluate the performance of image conversion based on the result of the recognition process performed by the recognition unit 13.

[0168] The image conversion evaluation unit 15 may evaluate the performance of the image conversion, i.e., the validity of the image conversion, by comparing the evaluation value calculated as described above with a preset threshold. If the evaluation value calculated using equation (6) is less than the threshold, the image conversion evaluation unit 15 evaluates the performance of the image conversion as high. In other words, the image conversion evaluation unit 15 evaluates the image conversion performed by the image conversion unit 12 as an image conversion that allows objects to be individually separated and recognized.

[0169] On the other hand, if the evaluation value calculated by Equation (6) is equal to or greater than the threshold, the image transformation evaluation unit 15 evaluates the performance of the image transformation as low. In other words, the image transformation evaluation unit 15 evaluates that the image transformation performed by the image transformation unit 12 is not an image transformation that allows objects to be individually separated and recognized. Note that performance evaluation based on the evaluation value of the performance of the image transformation may be performed by a method other than the method of comparing the evaluation value with a threshold. The evaluation method used by the image transformation evaluation unit 15 is not limited to the above-described method exemplified in embodiment 4, and may be any method.

[0170] If the image transformation performed by the image transformation unit 12 is evaluated as not being an image transformation that allows objects to be individually separated and recognized, the learning unit 14 may learn image transformation parameters. The learning unit 14 learns the image transformation parameters using the image data stored in the storage unit 16 and the dataset generated by the dataset generation unit 19. In this way, the learning unit 14 learns the image transformation parameters based on the results of the evaluation by the image transformation evaluation unit 15.

[0171] According to the fourth embodiment, the control device 60 includes a gripping parameter adjustment unit 32 that adjusts gripping parameters, which are parameters for a hand that grips an object, and a recognition parameter adjustment unit 33 that adjusts recognition parameters, which are parameters used in object recognition by the recognition unit 13. The control device 60 can automatically adjust the gripping parameters and recognition parameters by simulating the actual environment in which an object is photographed. The control device 60 can reduce the number of personnel and costs required to adjust the gripping parameters or recognition parameters, and can improve the success rate of picking by the robot. As described above, the control device 60 has the effects of enabling labor savings, cost reduction, and an improvement in the success rate of picking.

[0172] The control device 60 includes an image transformation evaluation unit 15 that evaluates the performance of image transformation using the image transformation parameters based on the operation results of the robot main body 62. The control device 60 can confirm the validity of the image transformation parameters based on operation results such as the probability of successfully grasping an object, the operation time required to grasp the object, or the cause of failure to grasp the object. This allows the control device 60 to obtain a transformed image in which objects are individually separated and easily recognized, further improving the success rate of picking.

[0173] By using the control device 60 to control the industrial robot, it is possible to reduce the number of people required at the site where the industrial robot is operated, shorten the time required to verify the industrial robot before it is put into operation, and reduce the costs required for verification before it is put into operation.

[0174] In the above description, the control device 60 is described as including the object recognition device 10D, but is not limited to this. The control device 60 may also be provided with the object recognition device 10A shown in Fig. 1, the object recognition device 10B shown in Fig. 3, or the object recognition device 10C shown in Fig. 6.

[0175] Next, a hardware configuration for realizing the object recognition devices 10A, 10B, 10C, and 10D according to the first to fourth embodiments will be described. The processing units included in each of the object recognition devices 10A, 10B, 10C, and 10D are realized by a processing circuit. The processing units included in the object recognition device 10A include an image acquisition unit 11, an image conversion unit 12, and a recognition unit 13 shown in FIG. 1. The processing units included in the object recognition device 10B include an image acquisition unit 11, an image conversion unit 12, a recognition unit 13, a learning unit 14, and an image conversion evaluation unit 15 shown in FIG. 3. The processing units included in the object recognition device 10C include an image acquisition unit 11, an image conversion unit 12, a recognition unit 13, a learning unit 14, an image conversion evaluation unit 15, a simulation condition setting unit, a scene generation unit 18, and a data set generation unit 19 shown in FIG. 10, the processing units of the object recognition device 10D include an image acquisition unit 11, an image conversion unit 12, a recognition unit 13, a learning unit 14, an image conversion evaluation unit 15, a simulation condition setting unit 17, a scene generation unit 18, a data set generation unit 19, and a parameter adjustment unit 31. The processing circuitry may be a circuit in which a processor executes software, or may be a dedicated circuit.

[0176] 16 is a diagram illustrating a first example of a hardware configuration for realizing the object recognition devices 10A, 10B, 10C, and 10D according to embodiments 1 to 4. In the first example, the object recognition devices 10A, 10B, 10C, and 10D include a processing circuit 70 having a processor 72 and a memory 73, an input unit 71, and an output unit 74.

[0177] The input unit 71 is an interface circuit that receives data transmitted to the object recognition devices 10A, 10B, 10C, and 10D from outside the object recognition devices 10A, 10B, 10C, and 10D and provides the data to the processor 72. The output unit 74 is an interface circuit that outputs data from the processor 72 or the memory 73 to outside the object recognition devices 10A, 10B, 10C, and 10D.

[0178] The processing units of each of the object recognition devices 10A, 10B, 10C, and 10D are realized by software, firmware, or a combination of software and firmware. The software or firmware is written as a program and stored in the memory 73. In the processing circuit 70, the processor 72 reads and executes the program stored in the memory 73, thereby realizing each function of the processing unit. That is, the processing circuit 70 includes the memory 73 for storing a program that results in the processing of each function of the processing unit. This program can be said to be a program that causes a computer to execute the procedures and methods of the object recognition devices 10A, 10B, 10C, and 10D. The memory 73 is also used as temporary memory when the processor 72 executes various processes. The storage unit 16 of each of the object recognition devices 10B, 10C, and 10D is realized by the memory 73.

[0179] The processor 72 includes, for example, one or more of a central processing unit (CPU), a digital signal processor (DSP), and a system large-scale integration (LSI). The memory 73 includes one or more of a random access memory (RAM), a read-only memory (ROM), a flash memory, an erasable programmable read-only memory (EPROM), and an electrically erasable programmable read-only memory (EEPROM). The memory 73 also includes a recording medium on which a computer-readable program is recorded. Such a recording medium includes one or more of a nonvolatile or volatile semiconductor memory, a magnetic disk, a flexible memory, an optical disk, a compact disk, and a digital versatile disc (DVD).

[0180] 16 shows an example of hardware in which a processing unit is realized by a general-purpose processor 72 and a memory 73, but the processing unit may also be realized by a dedicated hardware circuit. Fig. 17 is a diagram showing a second example of a hardware configuration in which the object recognition devices 10A, 10B, 10C, and 10D according to the first to fourth embodiments are realized. In the second example, the object recognition devices 10A, 10B, 10C, and 10D include a processing circuit 75, which is a dedicated hardware circuit, an input unit 71, and an output unit 74.

[0181] The processing circuitry 75 may be a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), or a combination of these. Each function of the processing unit may be realized by the processing circuitry 75 individually, or all functions may be realized collectively by the processing circuitry 75. Each function of the processing unit may be realized by a combination of the processing circuitry 70 and the processing circuitry 75.

[0182] The control device 60 according to the fourth embodiment is realized by a hardware configuration similar to the hardware configuration illustrated in Fig. 16 or the hardware configuration illustrated in Fig. 17. In the control device 60, the processing unit of the object recognition device 10D and the robot control unit 61 are realized by a processing circuit 70 or a processing circuit 75. The output unit 74 transmits data to the robot main body 62.

[0183] The configurations shown in the above embodiments are examples of the contents of the present disclosure. The configurations of each embodiment can be combined with other known technologies. The configurations of each embodiment can also be combined as appropriate. Part of the configuration of each embodiment can be omitted or modified without departing from the gist of the present disclosure.

[0184] 10A, 10B, 10C, 10D Object recognition device, 11 Image acquisition unit, 12 Image conversion unit, 13 Recognition unit, 14 Learning unit, 15 Image conversion evaluation unit, 16 Memory unit, 17 Simulation condition setting unit, 18 Scene generation unit, 19 Data set generation unit, 21 Photographing device information setting unit, 22 Object information setting unit, 23 Environment information setting unit, 24 2D data generation unit, 25 3D data generation unit, 26 Annotation data generation unit, 27 Target conversion image generation unit, 31 Parameter adjustment unit, 32 Grasping parameter adjustment unit, 33 Recognition parameter adjustment unit, 41 Grasping parameter adjustment range determination unit, 42 Grasping parameter change unit, 43 Model rotation unit, 44 Grasping evaluation unit, 45 Grasping parameter value determination unit, 46 Grasping parameter adjustment completion determination unit, 51 Recognition parameter adjustment range determination unit, 52 Recognition parameter change unit, 53 Recognition trial unit, 54 recognition evaluation unit, 55 recognition parameter value determination unit, 56 recognition parameter adjustment completion determination unit, 60 control device, 61 robot control unit, 62 robot body, 70, 75 processing circuit, 71 input unit, 72 processor, 73 memory, 74 output unit.

Claims

1. An image acquisition unit that acquires an image of an object; An image conversion unit that converts the image into a converted image in which the edges of the object are replaced with edges in which features common to the edges of each of the plurality of objects are emphasized; A recognition unit that recognizes the object based on the converted image, characterized in that the object recognition device comprises: An object recognition device characterized by the above.

2. The image conversion unit converts the image into the converted image based on the result of learning the image and the converted image. The object recognition device according to claim 1, characterized in that the above is the case.

3. The image conversion unit converts the image into the converted image based on image conversion parameters that are the result of learning the image and the converted image. The object recognition device according to claim 2, characterized in that the above is the case.

4. A learning unit that learns the image conversion parameters used for the conversion from the image to the converted image; The image conversion unit converts the image into the converted image based on the image conversion parameters that are the result of learning by the learning unit. The object recognition device according to claim 3, characterized in that the above is the case.

5. An image conversion evaluation unit that evaluates the performance of image conversion using the image conversion parameters. The object recognition device according to claim 4, characterized in that the above is the case.

6. The learning unit learns the image conversion parameters based on the result of evaluation by the image conversion evaluation unit. The object recognition device according to claim 5, characterized in that the above is the case.

7. A simulation condition setting unit that sets simulation conditions that are conditions for simulating the photographing of the object and include at least one of photographing device information that is information about the device for photographing the image and object information that is information about the object; A scene generation unit that simulates and generates a scene in which the object is photographed based on the simulation conditions; A dataset generation unit that generates a dataset used for learning the image conversion parameters based on the scene generated by the scene generation unit, characterized in that the object recognition device according to any one of claims 4 to 6 comprises: The object recognition device according to any one of claims 4 to 6, characterized in that the above is the case.

8. The simulation condition setting unit: A photographing device information setting unit that sets the photographing device information including at least one of information about the specifications of the device and information about the installation of the device; An object information setting unit that sets the object information. An environment information setting unit that sets environment information, which is information about the environment at the time of shooting in the simulation, and sets the simulation conditions including the shooting device information, the object information, and the environment information The object recognition device according to claim 7, characterized in that.

9. The dataset generation unit An image data generation unit that generates image data that is at least one of two-dimensional data indicating the scene generated by the scene generation unit and three-dimensional data indicating the scene generated by the scene generation unit, An annotation data generation unit that generates annotation data to be attached to the image data, A target conversion image generation unit that generates a target conversion image that is a target image for image conversion from the image data, and includes The object recognition device according to claim 7, characterized in that.

10. The learning unit learns the image conversion parameters based on the target conversion image and the annotation data The object recognition device according to claim 9, characterized in that.

11. An image acquisition unit that acquires an image of the object, An image conversion unit that converts the image into a converted image in which the edges of the object are replaced with edges in which features common to the edges of each of the plurality of objects are emphasized, A recognition unit that recognizes the object based on the converted image, A robot control unit that controls a robot body that grips the object recognized by the recognition unit, and includes The control device characterized by that.

12. A simulation condition setting unit that sets simulation conditions that are conditions when simulating the shooting of the object and include at least one of shooting device information that is information about the device that shoots the image and object information that is information about the object, A scene generation unit that simulates and generates a scene in which the object is shot based on the simulation conditions, A dataset generation unit that generates a dataset used for learning image conversion parameters used for conversion from the image to the converted image based on the scene generated by the scene generation unit, and includes The control device according to claim 11, characterized in that.

13. A gripping parameter adjustment unit that adjusts gripping parameters, which are parameters for the hand that grips the object in the robot body, is provided The control device according to claim 12, characterized in that.

14. The gripping parameter adjustment unit includes: a gripping parameter adjustment range determination unit that determines the adjustment range of each of the plurality of gripping parameters; a gripping parameter change unit that changes the gripping parameter to be adjusted among the plurality of gripping parameters; a model rotation unit that rotates the model of the object; a gripping evaluation unit that evaluates the quality of gripping when the object is gripped by the hand on the assumption that the object takes the same posture as the model rotated by the model rotation unit; a gripping parameter value determination unit that determines the value of each of the plurality of gripping parameters based on the evaluation result by the gripping evaluation unit; a gripping parameter adjustment end determination unit that determines whether the adjustment of all of the plurality of gripping parameters has been completed, and adjusts each of the plurality of gripping parameters by determining the value of each of the plurality of gripping parameters The control device according to claim 13, characterized in that.

15. The control device according to claim 12, further comprising a recognition parameter adjustment unit that adjusts recognition parameters that are parameters used in the recognition of the object by the recognition unit. The control device according to claim 12, characterized in that.

16. The control device according to claim 13 or 14, further comprising a recognition parameter adjustment unit that adjusts recognition parameters that are parameters used in the recognition of the object by the recognition unit, The recognition parameter adjustment unit includes: a recognition parameter adjustment range determination unit that determines the adjustment range of each of the plurality of recognition parameters; a recognition parameter change unit that changes the recognition parameter to be adjusted among the plurality of gripping parameters; a recognition trial unit that performs a recognition process using the gripping parameters adjusted by the gripping parameter adjustment unit; a recognition evaluation unit that evaluates the quality of the result of the recognition process; a recognition parameter value determination unit that determines the value of each of the plurality of recognition parameters based on the evaluation result by the recognition evaluation unit; a recognition parameter adjustment end determination unit that determines whether the adjustment of all of the plurality of recognition parameters has been completed, and adjusts each of the plurality of recognition parameters by determining the value of each of the plurality of recognition parameters The control device according to claim 13 or 14, characterized in that.

17. The control device further comprises an image conversion evaluation unit that evaluates the performance of image conversion using the image conversion parameters based on the operation result of the robot main body. The control device according to any one of claims 12 to 15, characterized in that...

18. The operation result includes at least one of information on the probability of successful grasping of the object by the robot body, information on the operation time of the operation of grasping the object by the robot body, and information indicating the cause of failure of the robot body to grasp the object. The control device according to claim 17, characterized in that...

19. A step of acquiring an image obtained by photographing an object; A step of converting the image into a converted image in which the edge of the object is replaced with an edge in which features common to the edges of a plurality of the objects are emphasized; A step of recognizing the object based on the converted image, including: An object recognition method, characterized in that...