Object recognition device, control device, and object recognition method

The object recognition device enhances edge features in images to improve recognition accuracy by simplifying random shapes and emphasizing common edges, addressing the challenge of individual object separation and recognition in conventional technologies.

JP7857418B2Active Publication Date: 2026-05-12MITSUBISHI ELECTRIC CORP +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2023-08-01
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Conventional object recognition technologies fail to adequately obtain sufficient information regarding the contours of objects, leading to difficulties in individually separating and recognizing objects, which decreases recognition accuracy.

Method used

An object recognition device that includes an image acquisition unit, an image conversion unit, and a recognition unit, where the image conversion unit enhances the edges of objects by emphasizing common features, transforming the image into a form that simplifies random shapes and emphasizes common edges, facilitating accurate recognition.

Benefits of technology

Improves the accuracy of object recognition by reducing individual differences in object edges, enabling easier separation and recognition, thereby enhancing the success rate of object picking by industrial robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007857418000007
    Figure 0007857418000007
  • Figure 0007857418000008
    Figure 0007857418000008
  • Figure 0007857418000009
    Figure 0007857418000009
Patent Text Reader

Abstract

An object recognition device (10A) is provided with an image acquisition unit (11), an image conversion unit (12), and a recognition unit (13). The image acquisition unit (11) acquires an image obtained by imaging an object. The image conversion unit (12) converts the image into a post-conversion image in which the edge of the object is replaced with the edge having emphasized features. The recognition unit (13) recognizes the object, on the basis of the post-conversion image. The object recognition device (10A) can improve the accuracy of object recognition when recognizing an object from an image obtained by imaging the object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to an object recognition device, a control device, and an object recognition method for recognizing objects. [Background technology]

[0002] Industrial robots that pick up objects one by one from a group of objects are known for automating tasks in production sites and other environments. When an industrial robot picks up an object, a recognition process is performed to recognize the object to be picked up, and the gripping position when the industrial robot picks up the object is detected.

[0003] Patent Document 1 discloses an information processing device for training a recognition device that performs recognition processing, relating to a recognition process that recognizes an object from an image of the object. The information processing device disclosed in Patent Document 1 acquires training data having characteristics equivalent to the recognition target data by converting an image captured by a camera used when collecting training data so that it has characteristics equivalent to the recognition target data input to the recognition device. Such characteristics are image quality characteristics such as noise, blur, color tone, or white balance. According to the technology of Patent Document 1, even if the characteristics of the camera used when collecting training data and the camera used when acquiring the recognition target data are different, a decrease in recognition accuracy can be prevented. In other words, according to the technology of Patent Document 1, a decrease in recognition accuracy due to individual differences in cameras can be prevented. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2021-82068 [Overview of the project] [Problems that the invention aims to solve]

[0005] However, even when images are converted using the conventional technology disclosed in Patent Document 1, sufficient information regarding the contours of objects may not be obtained. If sufficient information regarding the contours of objects is not obtained, it may be difficult to individually separate and recognize objects from among multiple objects. Therefore, the above-mentioned conventional technology has the problem that the accuracy of object recognition may decrease due to the difficulty in individually separating and recognizing objects.

[0006] This disclosure has been made in view of the above, and aims to provide an object recognition device that can improve the accuracy of object recognition. [Means for solving the problem]

[0007] To solve the above-mentioned problems and achieve the objective, the object recognition device according to this disclosure comprises: an image acquisition unit that acquires an image of an object; an image conversion unit that converts the image into a converted image in which the edges of multiple objects are replaced with edges in which features common to each edge of the multiple objects are emphasized; and a recognition unit that recognizes an object based on the converted image. The image conversion unit replaces the edges of multiple objects with edges in which features common to each edge of the multiple objects are emphasized. And it actually exists The image is transformed into a post-transformation image by first identifying a shape from which random elements have been eliminated, and then replacing the shape of each object with this identified shape. [Effects of the Invention]

[0008] The object recognition device described herein has the effect of improving the accuracy of object recognition. [Brief explanation of the drawing]

[0009] [Figure 1] A diagram showing the functional configuration of the object recognition device according to Embodiment 1. [Figure 2] A flowchart showing the procedure of processing performed by the object recognition device according to Embodiment 1. [Figure 3] This figure shows the functional configuration of the object recognition device according to Embodiment 2. [Figure 4] Flowchart showing the procedure of the process executed by the object recognition device according to Embodiment 2 [Figure 5] Diagram for explaining an example of an image conversion method by the object recognition device according to Embodiment 2 [Figure 6] Diagram showing the functional configuration of the object recognition device according to Embodiment 3 [Figure 7] Flowchart showing the procedure of the process executed by the object recognition device according to Embodiment 3 [Figure 8] Diagram showing a configuration example of the simulation condition setting unit of the object recognition device according to Embodiment 3 [Figure 9] Diagram showing a configuration example of the dataset generation unit of the object recognition device according to Embodiment 3 [Figure 10] Diagram showing the functional configuration of the control device according to Embodiment 4 [Figure 11] Diagram showing a configuration example of the parameter adjustment unit of the control device according to Embodiment 4 [Figure 12] Diagram showing a configuration example of the gripping parameter adjustment unit of the control device according to Embodiment 4 [Figure 13] Flowchart showing the procedure of the process executed by the gripping parameter adjustment unit of the control device according to Embodiment 4 [Figure 14] Diagram showing a configuration example of the recognition parameter adjustment unit of the control device according to Embodiment 4 [Figure 15] Flowchart showing the procedure of the process executed by the recognition parameter adjustment unit of the control device according to Embodiment 4 [Figure 16] Diagram showing a first example of the hardware configuration for realizing the object recognition device according to Embodiments 1 to 4 [Figure 17] Diagram showing a second example of the hardware configuration for realizing the object recognition device according to Embodiments 1 to 4

Modes for Carrying Out the Invention

[0010] The object recognition device, control device, and object recognition method according to the embodiment will be described in detail below with reference to the drawings.

[0011] Embodiment 1. Figure 1 shows the functional configuration of the object recognition device 10A according to Embodiment 1. The object recognition device 10A recognizes an object from an image of the object. The object recognition device 10A comprises an image acquisition unit 11, an image conversion unit 12, and a recognition unit 13.

[0012] The objects recognized by the object recognition device 10A may be objects that are to be grasped by an industrial robot. The industrial robot may grasp the object recognized by the object recognition device 10A from among multiple objects and pick up the object. The industrial robot picks up objects one by one by repeating this operation. Each of the multiple objects may be the same shape as the others, or they may be of different shapes. In the following description, each object when each of the multiple objects is the same shape as the others will be referred to as a standard-shaped object. Each object when each of the multiple objects is of different shapes will be referred to as an irregular-shaped object. An example of a standard-shaped object is a part of an industrial product. An example of an irregular-shaped object is food that can be sorted into pieces. Note that the objects recognized by the object recognition device 10A are not limited to objects that are to be picked. The object recognition device 10A may recognize objects other than those that are to be picked.

[0013] The arrangement of multiple objects when objects are recognized by the object recognition device 10A is arbitrary. The multiple objects may be, for example, piled up randomly or arranged in a line. Piling up randomly means that the objects are piled up in a state where their positions and orientations are random. Arrangement of multiple objects means that each object is placed in a predetermined position and in a predetermined orientation.

[0014] The image acquisition unit 11 acquires an image of the object. The image conversion unit 12 converts the image of the object into a converted image in which the edges of the object are replaced with edges that have been enhanced for their features. The recognition unit 13 recognizes the object based on the converted image.

[0015] In Embodiment 1, the image conversion unit 12 converts an image of an object into a converted image in which the edges of the objects are replaced with edges that emphasize the common features of each edge of the multiple objects. That is, in the edges of the converted image, the common features present in each edge of the multiple objects are emphasized. The edges of the converted image may also be those that emphasize the edge features of an object of a predetermined shape. That is, the image conversion unit 12 may convert an image of an object into a converted image in which the edges of the objects are replaced with edges that emphasize predetermined features.

[0016] Next, the procedure for processing performed by the object recognition device 10A will be described. Figure 2 is a flowchart showing the procedure for processing performed by the object recognition device 10A according to Embodiment 1.

[0017] In step S11, the image acquisition unit 11 acquires an image of the object. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12.

[0018] In step S12, the image conversion unit 12 performs an image conversion process, which is the process of converting the image input to the image conversion unit 12. The image conversion unit 12 converts the image into a converted image in which the edges of multiple objects are replaced with edges that have been enhanced with features common to each edge of the objects. The image conversion unit 12 outputs the converted image to the recognition unit 13.

[0019] In step S13, the recognition unit 13 performs recognition processing to recognize an object based on the converted image input to the recognition unit 13. The recognition unit 13 outputs the result of object recognition to the outside of the object recognition device 10A. With this, the object recognition device 10A completes the processing according to the procedure shown in Figure 2.

[0020] Next, the details of the processing in each part of the object recognition device 10A will be described. The device that photographs objects is a two-dimensional (2D) camera or a three-dimensional (3D) sensor. The image acquisition unit 11 acquires images, for example, images taken by a general-purpose 2D camera or images taken by a general-purpose 3D sensor. The image acquisition unit 11 acquires images, for example, when images taken by an external 2D camera or 3D sensor are input to the image acquisition unit 11. Alternatively, the 2D camera or 3D sensor may be provided in the image acquisition unit 11. In this case, the image acquisition unit 11 acquires images when the 2D camera or 3D sensor takes images.

[0021] The image acquisition unit 11 acquires 2D data, which is the data of a 2D image, using a 2D camera. A 2D image is, for example, a grayscale image or a color image such as an RGB image. Alternatively, the image acquisition unit 11 acquires 3D data, which is the data of a 3D image, using a 3D sensor. A 3D image is, for example, a 3D point cloud or a depth image. The following description will mainly focus on the case where the image acquired by the image acquisition unit 11 is a depth image. In the following description, a 2D camera may be simply referred to as a camera, and a 3D sensor may be simply referred to as a sensor. The image acquisition unit 11 may acquire an image captured by a single camera or a single sensor, or it may acquire multiple images captured by multiple cameras or multiple sensors.

[0022] The image acquisition unit 11 may acquire images captured by multiple cameras or sensors having the same measurement method or specifications. Alternatively, the image acquisition unit 11 may acquire images captured by multiple cameras or sensors having different measurement methods or specifications. An example of multiple sensors having the same measurement method or specifications is when all of them perform measurement using a structured illumination method. An example of multiple sensors having different measurement methods or specifications is when some of the multiple sensors perform measurement using a structured illumination method, and the rest of the multiple sensors perform measurement using a Time of Flight (ToF) method.

[0023] The image input from the image acquisition unit 11 to the image conversion unit 12 may be a single image or multiple images. When multiple images are input to the image conversion unit 12, the multiple images may be, for example, images taken by multiple cameras with the same measurement method or specifications, or images taken by multiple sensors with the same measurement method or specifications. Alternatively, the multiple images may be images taken by multiple cameras with different measurement methods or specifications, or images taken by multiple sensors with different measurement methods or specifications.

[0024] The image transformation unit 12 applies a deformation to the edges of objects in the image to emphasize their features. In this way, the image transformation unit 12 replaces the object's edges with edges that have emphasized their features. The image transformation unit 12 emphasizes the features of the edges by simplifying the edges of objects in the image compared to their actual edges. It can also be said that the image transformation unit 12 simplifies the object's edges based on their features.

[0025] If an object is irregularly shaped, the image conversion unit 12 reduces individual differences in the edges of each object by replacing the edges of each object with edges that have emphasized features. As a first example, the image conversion unit 12 converts the image into a converted image by replacing the edges of an object with edges that have emphasized predetermined features. For example, if each object has random bumps and dips, the bumps and dips included in the edges are removed, and the edges of each object are converted into smooth and simple edges. The image conversion unit 12 replaces the random shapes of each object with primitive shapes, such as simple shapes like a cylinder. In other words, the objects depicted in the image are converted by the image conversion unit 12 into figures that represent the approximate shape of the objects.

[0026] Furthermore, the image conversion unit 12 does not necessarily replace each edge of multiple objects with edges of the same shape. The replaced edges of each of the multiple objects may include edges of different shapes. For example, in the case where the edges of each object are replaced with cylinders, the replaced edges may include edges of different cylinder heights, or edges of different cylinder radii. The image conversion unit 12 may also replace the edges of each object with edges of shapes other than cylinders. If the objects are standard shapes, the image conversion unit 12 may replace the edges of the objects with simplified edges compared to the actual edges.

[0027] As a second example, the image conversion unit 12 converts an image to a converted image by replacing the edges of multiple objects with edges that emphasize features common to each edge of multiple objects. In this case, the image conversion unit 12 extracts characteristic edge information from each edge of multiple objects to find features common to each edge of multiple objects. Based on the features common to each edge of multiple objects, the image conversion unit 12 converts the image to a converted image. In the second example as well, the image conversion unit 12 replaces the random shapes of each object with primitive shapes, such as simple shapes like cylinders. In the case where the edges of each object are replaced with cylinders, the cylinder is a shape obtained by eliminating random elements in the edges of multiple objects, and can be said to be a feature common to each edge of multiple objects. In the second example as well, the replaced edges of each of the multiple objects may include edges of different shapes. Furthermore, the image conversion unit 12 may replace the edges of each object with edges of shapes other than cylinders.

[0028] The image conversion unit 12 detects edges by, for example, applying image conversion processing to the depth image obtained by the sensor using a known method such as the Canny method. When the image conversion unit 12 applies image conversion processing to each image captured by multiple sensors, it may pre-determine filter values ​​for each sensor that make edge detection easier, and detect edges based on the determined filter values. In this case, the image conversion unit 12 extracts characteristic edge information from the edges obtained in each image for each of the multiple objects, and generates a converted image based on the extracted information. In this way, the image conversion unit 12 can extract edge information from each image obtained by multiple sensors, thereby mutually complementing edge information that could not be obtained from individual images using multiple images.

[0029] The image conversion unit 12 may extract edge information from each of two types of images selected from a 3D point cloud, a depth image, a grayscale image, and a color image. In this case, the image conversion unit 12 extracts characteristic edge information from the edges obtained in each of the two types of images for each of the multiple objects, and generates a converted image based on the extracted information. In this case as well, the image conversion unit 12 can mutually complement the edge information that could not be obtained from the individual images using the two types of images.

[0030] The image conversion performed by the image conversion unit 12 may utilize machine learning methods. The image conversion unit 12 may convert an image of an object into a transformed image based on the results of learning between an image of an object and a transformed image. In a machine learning-based method, for example, a distance image obtained by a sensor and a distance image with enhanced object edges may be input to a neural network, and learning may be performed to enhance the edges of the distance image obtained by the sensor. For such learning, algorithms such as GAN (Generative Adversarial Networks) or U-Net can be used. In addition, learning algorithms such as support vector machines or random forests may be used to estimate the edges contained in the distance image obtained by the sensor.

[0031] The recognition unit 13 performs recognition processing on the converted image converted by the image conversion unit 12. The recognition unit 13 may calculate the position and orientation of the robot's hand when grasping an object, or the position and orientation of the object, through the recognition processing. The recognition unit 13 may use a neural network to calculate the position and orientation of the hand when grasping an object, or the position and orientation of the object. The recognition processing by the recognition unit 13 is not limited to the processing described in Embodiment 1. The recognition processing by the recognition unit 13 may be any processing other than the processing described in Embodiment 1.

[0032] According to Embodiment 1, the image conversion unit 12 converts an image into a converted image in which the edges of an object are replaced with edges that have been enhanced with features. The recognition unit 13 recognizes an object based on the converted image. By generating a converted image in which the edges of an object are replaced with edges that have been enhanced with features, the object recognition device 10A can individually separate and recognize objects from among multiple objects. Regardless of the type of sensor or camera, the object recognition device 10A can obtain a converted image in which individual differences in the edges of each object are reduced and individual object recognition is easier. As a result, the object recognition device 10A has the effect of improving the accuracy of object recognition. By using the object recognition device 10A for object recognition when picking objects with an industrial robot, the accuracy of object recognition can be improved, and the success rate of picking by the industrial robot can be improved.

[0033] Embodiment 2. Embodiment 2 describes an example of learning from an image of an object and a converted image. Figure 3 is a diagram showing the functional configuration of the object recognition device 10B according to Embodiment 2. In Embodiment 2, the same reference numerals are used for the same components as in Embodiment 1, and the configuration that differs from Embodiment 1 will be mainly described.

[0034] The object recognition device 10B recognizes objects from images of objects that have been captured. The object recognition device 10B has the same configuration as the object recognition device 10A shown in Figure 1. Furthermore, the object recognition device 10B includes a learning unit 14, an image conversion and evaluation unit 15, and a storage unit 16.

[0035] The image conversion unit 12 performs machine learning to generate a converted image in which the edges of an object are replaced with edges that have been enhanced to highlight certain features. The image conversion unit 12 performs machine learning to generate a converted image in which the edges of an object are replaced with edges that have been enhanced to highlight features common to each of the edges of multiple objects. Alternatively, the image conversion unit 12 performs machine learning to generate a converted image in which the edges of an object are replaced with edges that have been enhanced to highlight predetermined features.

[0036] The image conversion unit 12 learns image conversion parameters for converting an image of an object into a converted image. The image conversion unit 12 stores the learned image conversion parameters in the storage unit 16. The image conversion unit 12 generates a converted image based on the image conversion parameters. The image conversion unit 12 stores the image conversion parameters in the storage unit 16.

[0037] The image conversion evaluation unit 15 evaluates the performance of the image conversion performed by the image conversion unit 12 using the image conversion parameters. The image conversion evaluation unit 15 determines, through evaluation processing, whether the image conversion performed by the image conversion unit 12 is an image conversion that allows for the individual separation and recognition of objects.

[0038] The learning unit 14 learns image conversion parameters based on the evaluation results from the image conversion evaluation unit 15. The learning of image conversion parameters by the learning unit 14 is equivalent to relearning the image conversion parameters learned by the image conversion unit 12.

[0039] The memory unit 16 stores image conversion parameters. The memory unit 16 also stores the image of the object, which is the image acquired by the image acquisition unit 11, and the target conversion image, which is the image targeted for conversion in the image conversion unit 12. The target conversion image is an image in which the edges of the object are emphasized and noise is reduced. The target conversion image may be a CG (Computer Graphics) image in which the edges of the object are simplified based on features common to the edges of each of multiple objects. Alternatively, the target conversion image may be an image similar to a CG image. In the target conversion image, the shape of the object may be approximated to a primitive shape, such as a cylinder. In Embodiment 2, the converted image may be a CG image or an image similar to a CG image. The target conversion image stored in the memory unit 16 is also used as training data for training in the learning unit 14.

[0040] Next, the procedure for processing performed by the object recognition device 10B will be described. Figure 4 is a flowchart showing the procedure for processing performed by the object recognition device 10B according to Embodiment 2.

[0041] In step S21, the image acquisition unit 11 acquires an image of the object. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12. The image acquisition unit 11 also saves the acquired image to the storage unit 16.

[0042] In step S22, the image conversion unit 12 performs an image conversion process, which is the process of converting the image input to the image conversion unit 12. In the image conversion process, the image conversion unit 12 learns image conversion parameters for obtaining a converted image from an image of an object. The image conversion unit 12 generates a converted image based on the image input to the image conversion unit 12 and the learned image conversion parameters. The image conversion unit 12 outputs the converted image to the recognition unit 13.

[0043] In step S23, the recognition unit 13 performs recognition processing to recognize an object based on the converted image input to the recognition unit 13. The recognition unit 13 outputs the result of object recognition to the outside of the object recognition device 10B. The recognition unit 13 also outputs the result of object recognition to the image conversion evaluation unit 15.

[0044] In step S24, the image conversion evaluation unit 15 performs an evaluation process of the image conversion performed in step S22 based on the result of object recognition by the recognition unit 13. The image conversion evaluation unit 15 determines through evaluation whether the image conversion performed in step S22 is an image conversion that allows for the individual separation and recognition of objects. Specifically, the image conversion evaluation unit 15 evaluates the appropriateness of the image conversion parameters used in the image conversion process.

[0045] If the image transformation performed in step S22 is determined to be an image transformation that allows for the individual separation and recognition of objects (step S25, Yes), the object recognition device 10B terminates the processing according to the procedure shown in Figure 4.

[0046] On the other hand, if the image transformation performed in step S22 is determined not to be an image transformation that allows for the individual separation and recognition of objects (step S25, No), the image transformation evaluation unit 15 outputs an evaluation result to the learning unit 14 indicating that the image transformation is not an image transformation that allows for the individual separation and recognition of objects.

[0047] Next, in step S26, the learning unit 14 performs a learning process to learn image transformation parameters. The learning unit 14 reads the image transformation parameters from the storage unit 16 and learns the image transformation parameters based on the evaluation results by the image transformation evaluation unit 15. The learning unit 14 stores the learned image transformation parameters in the storage unit 16. After completing step S26, the object recognition device 10B returns to step S21. The object recognition device 10B repeats the process from step S21 to step S25.

[0048] Next, we will describe the details of the processing in each part of the object recognition device 10B. In Embodiment 2, the image conversion unit 12 converts an image of an object into a converted image using machine learning techniques. Below, we will describe an example of an image conversion method utilizing machine learning, specifically a method using GAN.

[0049] Figure 5 is a diagram illustrating an example of an image transformation method by the object recognition device 10B according to Embodiment 2. Figure 5 shows a conceptual diagram of an image transformation process using GAN. In this example, the image acquired by the image acquisition unit 11 is assumed to be a depth image.

[0050] The image conversion unit 12 acquires a depth image from the image acquisition unit 11. The image conversion unit 12 also receives a sharp image, which is a depth image in which the edges of objects are emphasized and noise is reduced. The sharp image is stored, for example, in the storage unit 16. The image conversion unit 12 may also acquire a sharp image by reading the sharp image from the storage unit 16.

[0051] The image conversion unit 12 inputs the distance image and the sharp image to the generator G, causing the generator G to output a converted image. In this way, the image conversion unit 12 converts the acquired distance image into a converted image in which enhanced edges are added and noise is reduced.

[0052] The loss function of generator G is given by equation (1) below. In equation (1), L G θ is the loss function of the generator G, L is the labeling algorithm, G is the generator G, D is the discriminator D, x is the distance image acquired from the image acquisition unit 11, x' is the sharp image, α is the weight coefficient, l is the logarithmic function, θ G These are the parameters of generator G, and θ D represents the parameters of the classifier D. The loss function shown in equation (1) uses the double error function of LSGAN (Least Square Generative Adversarial Networks).

[0053]

number

[0054] A known algorithm, such as the Watershed algorithm, may be used for the labeling algorithm. When depth images are captured by multiple sensors, the sharp image input to the generator G may be the depth image captured by the sensor that best enhances the edges of the object among the multiple sensors. Alternatively, the sharp image input to the generator G may be a pseudo-image in which the edges of the object are enhanced and noise is reduced. A pseudo-image is an image that is artificially created.

[0055] The image conversion unit 12 calculates the loss function L of the generator G. G θ that minimizes G The image conversion unit 12 obtains θ that satisfies the following equation (2). G Obtain it.

[0056]

number

[0057] The image conversion unit 12 reads the target conversion image from the storage unit 16. The image conversion unit 12 performs edge detection processing on both the converted image generated by the generator G and the target conversion image. In the example shown in Figure 5, the target conversion image is assumed to be a computer graphics (CG) image. In CG images, objects are represented by primitive shapes. Known methods such as the Laplacian filter or the Canny method may be used for edge detection processing.

[0058] The image conversion unit 12 inputs the converted image after edge detection and the CG image after edge detection to the classifier D. The image conversion unit 12 performs image conversion on the converted image generated by the generator G so that the smooth edges of the CG image are represented.

[0059] The loss function of the discriminator D, which performs image transformation so that the transformed image represents smooth edges, is given by equation (3) below. In equation (3), L D is the loss function of the classifier D, and y represents the CG image.

[0060]

number

[0061] When an edge image is used as each of the transformed image and the CG image, the amount of information contained in the patch image becomes extremely small. Here, each of a plurality of patches cut out from the distance image, which is each of the transformed image and the CG image, is referred to as a patch image. Also, an image in which an object is represented only by edges is referred to as an edge image. The image conversion unit 12 divides the discriminator D into a Local branch that discriminates each patch image and a Global branch that discriminates the entire distance image, and discriminates the transformed image and the CG image. Thereby, the image conversion unit 12 can perform learning to obtain a transformed image that can represent the smooth edges of the CG image while preventing loss of information contained in the transformed image and capturing edge features locally and globally.

[0062] The image conversion unit 12 is the loss function L of the discriminator D D to minimize θ D is obtained. That is, the image conversion unit 12 obtains θ that satisfies the following formula (4) D is obtained.

[0063]

Equation

[0064] Note that each of the configuration of the network used for machine learning, the loss function, and the parameters is not limited to those described in Embodiment 2, and may be appropriately changed as necessary. The machine learning executed by the image conversion unit 12 is not limited to machine learning using GAN, and may be machine learning using a neural network other than GAN.

[0065] The image conversion unit 12 uses the above-described image conversion process using GAN as image conversion parameters, the parameter θ of the generator G G and the parameter θ of the discriminator D DThe image conversion unit 12 learns the following: In addition, the image conversion parameters learned by the image conversion unit 12 may be weight coefficients used in the loss function of the neural network, or weight coefficients set between the units that make up the network.

[0066] The image transformation evaluation unit 15 evaluates whether the image transformation performed by the image transformation unit 12 is an image transformation that allows for the individual separation and recognition of objects. In other words, the image transformation evaluation unit 15 evaluates the validity of the image transformation from the standpoint of object recognition.

[0067] Here, an example of evaluation by the image conversion evaluation unit 15 is described. An image acquired by the image acquisition unit 11 and a labeling image in which the objects contained in the acquired image are assigned correct labels are prepared. Labeling processing is then performed on the converted image obtained by the image conversion unit 12 from the image acquired by the image acquisition unit 11. The image conversion evaluation unit 15 determines the accuracy of individual separation by comparing the result of the labeling processing performed on the converted image with the true values ​​of the labeling image. Based on the accuracy of individual separation, the image conversion evaluation unit 15 evaluates the performance of the image conversion.

[0068] When evaluating the performance of image transformation based on the accuracy of individual separation, the evaluation value is expressed, for example, by the following equation (5). In equation (5), E is the evaluation value, T is the true value of the labeled image, R is the result of performing labeling on the transformed image, g(T) is the centroid position of the ground truth region, which is the region of an object in the labeled image, and g(R) is the centroid position of the region of an object in the transformed image. In equation (5), S(T) is the area of ​​the ground truth region, S(R) is the area of ​​the region of an object in the transformed image, α is the weighting coefficient set for the centroid position, β is the weighting coefficient set for the area of ​​the region of an object, and N is the number of samples, which is the number of objects included in the acquired image. The first term of the calculation formula following the summation symbol on the right side of equation (5) represents the evaluation value for the centroid position. The second term of the calculation formula following the summation symbol on the right side of equation (5) represents the evaluation value for the area of ​​the region.

[0069]

number

[0070] The image conversion evaluation unit 15 evaluates the performance of the image conversion by comparing the calculated evaluation value with a pre-set threshold. In other words, the image conversion evaluation unit 15 evaluates the validity of the image conversion by comparing the evaluation value with a threshold. The image conversion evaluation unit 15 evaluates the performance of the image conversion as high if the evaluation value obtained by equation (5) is less than the threshold. That is, the image conversion evaluation unit 15 evaluates that the image conversion performed by the image conversion unit 12 is an image conversion that allows for the individual separation and recognition of objects.

[0071] On the other hand, the image conversion evaluation unit 15 evaluates the image conversion performance as low if the evaluation value obtained by equation (5) is equal to or greater than a threshold. In other words, the image conversion evaluation unit 15 evaluates that the image conversion performed by the image conversion unit 12 is not an image conversion that allows for the individual separation and recognition of objects. Note that the performance evaluation based on the evaluation value of the image conversion performance may be performed by a method other than comparing the evaluation value with a threshold. The evaluation method by the image conversion evaluation unit 15 is not limited to the method exemplified in Embodiment 2 and is arbitrary.

[0072] If the image transformation performed by the image transformation unit 12 is evaluated as not being an image transformation that allows for the individual separation and recognition of objects, the learning unit 14 learns the image transformation parameters. In this way, the learning unit 14 learns the image transformation parameters based on the evaluation results by the image transformation evaluation unit 15.

[0073] The learning unit 14 reads the image of the object, which is an image acquired by the image acquisition unit 11, and the target transformation image from the storage unit 16. The learning unit 14 learns the image transformation parameters using the image of the object and the target transformation image.

[0074] Furthermore, if an additional sensor or camera used to photograph an object is added, the images captured by the added sensor or camera are added to the storage unit 16. When an image is added to the storage unit 16, the learning unit 14 may learn image transformation parameters based on the images stored in the storage unit 16 and the target transformation image.

[0075] According to Embodiment 2, the image conversion unit 12 converts the image acquired by the image acquisition unit 11 into a converted image based on the results of learning the image acquired by the image acquisition unit 11 and the converted image. The object recognition device 10B can obtain a converted image in which individual differences in the edges of each object are reduced and individual recognition of objects is made easier, regardless of the type of sensor or camera. Furthermore, the object recognition device 10B can obtain a converted image in which individual recognition of objects is made easier, depending on the shape or size of the object to be recognized. In addition, the object recognition device 10B can obtain a converted image that enables easier individual recognition by learning image conversion parameters based on the evaluation results by the image conversion evaluation unit 15. As a result, the object recognition device 10B can improve the accuracy of object recognition. By using the object recognition device 10B for object recognition when picking objects with an industrial robot, the accuracy of object recognition can be improved, and the success rate of picking by the industrial robot can be improved.

[0076] Embodiment 3. Embodiment 3 describes an example of automatically collecting training data used in machine learning through simulation. Figure 6 is a diagram showing the functional configuration of the object recognition device 10C according to Embodiment 3. In Embodiment 3, the same reference numerals are used for components that are the same as those in Embodiment 1 or 2, and the configurations that differ from Embodiment 1 or 2 will be described in detail.

[0077] The object recognition device 10C recognizes objects from images of objects that have been captured. The object recognition device 10C has the same configuration as the object recognition device 10B shown in Figure 3. Furthermore, the object recognition device 10C includes a simulation condition setting unit 17, a scene generation unit 18, and a data set generation unit 19. In Embodiment 3, simulation refers to simulating the scene in which an object is captured.

[0078] The simulation condition setting unit 17 sets simulation conditions for simulating the shooting of an object, which include at least one of the following: shooting equipment information, which is information about the equipment that takes the image, and object information, which is information about the object. The shooting equipment information is information about the camera that takes the image, or information about the sensor that takes the image.

[0079] The scene generation unit 18 simulates a scene in which an object is photographed based on the simulation conditions set by the simulation condition setting unit 17. The dataset generation unit 19 generates a dataset to be used for learning based on the scene generated by the scene generation unit 18. The dataset generated by the dataset generation unit 19 is the training data used for machine learning. The learning unit 14 learns image transformation parameters using the dataset generated by the dataset generation unit 19.

[0080] The memory unit 16 stores a CAD (Computer-Aided Design) model of an object. The CAD model stored in the memory unit 16 is at least one of a 3D CAD model and a 2D CAD model. If the object is irregularly shaped, or if there is no 3D CAD model of the object, a model generated by methods such as mesh generation based on data obtained from measuring the object using sensors, etc., may be stored as a 3D CAD model. In the following description, the CAD model may be simply referred to as a model.

[0081] Next, the procedure for processing performed by the object recognition device 10C will be described. Figure 7 is a flowchart showing the procedure for processing performed by the object recognition device 10C according to Embodiment 3.

[0082] In step S31, the simulation condition setting unit 17 sets the simulation conditions. The simulation condition setting unit 17 sets the simulation conditions which include at least one of the imaging equipment information and the object information. The simulation condition setting unit 17 reads the model from the storage unit 16 and sets the simulation conditions based on the model. The simulation condition setting unit 17 outputs the simulation condition information to the scene generation unit 18.

[0083] In step S32, the scene generation unit 18 simulates a scene in which an object is photographed based on the simulation conditions set by the simulation condition setting unit 17. The scene generation unit 18 reads a model from the storage unit 16. The scene generation unit 18 generates a scene based on the simulation conditions and the model. The scene generation unit 18 outputs information indicating the generated scene to the data set generation unit 19.

[0084] In step S33, the dataset generation unit 19 generates a dataset based on the scenes generated by the scene generation unit 18. The dataset generation unit 19 outputs the generated dataset to the learning unit 14.

[0085] In step S34, the image acquisition unit 11 acquires an image of the object. The image acquisition unit 11 outputs the acquired image to the image conversion unit 12. The image acquisition unit 11 also saves the acquired image to the storage unit 16.

[0086] In step S35, the image conversion unit 12 performs an image conversion process, which is the process of converting the image input to the image conversion unit 12. In the image conversion process, the image conversion unit 12 learns image conversion parameters for obtaining a converted image from an image of an object. The image conversion unit 12 generates a converted image based on the image input to the image conversion unit 12 and the learned image conversion parameters. The image conversion unit 12 outputs the converted image to the recognition unit 13.

[0087] In step S36, the recognition unit 13 performs recognition processing to recognize an object based on the converted image input to the recognition unit 13. The recognition unit 13 outputs the result of object recognition to the outside of the object recognition device 10C. The recognition unit 13 also outputs the result of object recognition to the image conversion evaluation unit 15.

[0088] In step S37, the image conversion evaluation unit 15 performs an evaluation process of the image conversion performed in step S35 based on the result of object recognition by the recognition unit 13. The image conversion evaluation unit 15 determines through evaluation whether the image conversion performed in step S35 is an image conversion that allows for the individual separation and recognition of objects. Specifically, the image conversion evaluation unit 15 evaluates the appropriateness of the image conversion parameters used in the image conversion process.

[0089] If the image transformation performed in step S35 is determined to be an image transformation that allows for the individual separation and recognition of objects (step S38, Yes), the object recognition device 10C terminates the processing according to the procedure shown in Figure 7.

[0090] On the other hand, if the image transformation performed in step S35 is determined not to be an image transformation that allows for the individual separation and recognition of objects (step S38, No), the image transformation evaluation unit 15 outputs an evaluation result to the learning unit 14 indicating that the image transformation is not an image transformation that allows for the individual separation and recognition of objects.

[0091] Next, in step S39, the learning unit 14 performs a learning process to learn image transformation parameters. The learning unit 14 learns image transformation parameters based on the dataset input to the learning unit 14 and the evaluation results from the image transformation evaluation unit 15. The learning unit 14 stores the learned image transformation parameters in the storage unit 16. After completing step S39, the object recognition device 10C returns to step S34. The object recognition device 10C repeats the process from step S34 to step S38.

[0092] Next, we will describe the details of the functions of the simulation condition setting unit 17, the scene generation unit 18, and the dataset generation unit 19. Figure 8 is a diagram showing an example of the configuration of the simulation condition setting unit 17 of the object recognition device 10C according to Embodiment 3.

[0093] The simulation condition setting unit 17 comprises a shooting equipment information setting unit 21, an object information setting unit 22, and an environment information setting unit 23. The simulation condition setting unit 17 sets simulation conditions including shooting equipment information, object information, and environment information.

[0094] The imaging equipment information setting unit 21 sets imaging equipment information that includes at least one of the following: information about the specifications of the equipment used to capture images and information about the installation of the equipment. The information about the equipment specifications includes parameters such as field of view, working distance, installation angle, or resolution. The information about the equipment installation includes parameters that indicate the position of the equipment in the simulation. The imaging equipment information setting unit 21 may also read a model from the storage unit 16 and verify that objects are within the equipment's field of view in the simulation.

[0095] The object information setting unit 22 sets object information. The object information includes information about the object's color or size. The object information setting unit 22 reads a model from the storage unit 16 and sets the object information based on the model. The object information set by the object information setting unit 22 may also include information about the number of models used in the scene generated by the scene generation unit 18.

[0096] The environmental information setting unit 23 sets environmental information. The environmental information includes information about the light source, such as information about diffuse reflection or ambient light. The environmental information may also include information indicating the size of the box in which the objects are piled up, or information indicating the method of supplying the objects to the location where the objects are picked up. Here, the supply method refers to the state in which the objects are arranged, such as a piled-up state or an aligned state. The environmental information setting unit 23 may also read a model from the storage unit 16 and verify the supply method in a simulation based on the model.

[0097] The simulation condition setting unit 17 outputs simulation condition information, including imaging equipment information, object information, and environmental information, to the scene generation unit 18. Note that the simulation conditions set by the simulation condition setting unit 17 are not limited to those including imaging equipment information, object information, and environmental information. The simulation conditions only need to include at least one of either imaging equipment information or object information. The simulation condition setting unit 17 is not limited to comprising all three components: the imaging equipment information setting unit 21, the object information setting unit 22, and the environmental information setting unit 23. The simulation condition setting unit 17 only needs to be equipped with at least one of either the imaging equipment information setting unit 21 or the object information setting unit 22.

[0098] The scene generated by the scene generation unit 18 can be described as the situation in which the objects are photographed. For example, in a case where objects are picked one by one from a pile of scattered objects, the scene is a situation in which the objects are arranged in a random state of position and orientation. The scene generation unit 18 reads a model from the memory unit 16 and simulates the scene using the simulation conditions and the model.

[0099] The scene generation unit 18 may generate a scene using multiple models. If the object is irregularly shaped, the multiple models used to generate the scene may include models of the same shape, or models of different shapes. For example, the scene generation unit 18 may read 10 models of different shapes from the storage unit 16 and generate 5 copies of each model, thereby using 50 models. In this way, by using models of different shapes to generate a scene, the object recognition device 10C can perform image transformations for each of the multiple objects with irregular shapes.

[0100] Figure 9 shows an example of the configuration of the dataset generation unit 19 of the object recognition device 10C according to Embodiment 3. The dataset generation unit 19 comprises a 2D data generation unit 24, a 3D data generation unit 25, an annotation data generation unit 26, and a target transformation image generation unit 27.

[0101] The 2D data generation unit 24 and the 3D data generation unit 25 function as image data generation units that generate image data. The 2D data generation unit 24 generates 2D data, which is image data representing a scene generated by the scene generation unit 18. The 2D data generated by the 2D data generation unit 24 is image data, such as a grayscale image or a color image. The 2D data generation unit 24 outputs the generated 2D data to the learning unit 14. The learning unit 14 may use the 2D data input to the learning unit 14 for learning image transformation parameters.

[0102] The 3D data generation unit 25 generates 3D data, which is image data representing the scene generated by the scene generation unit 18. The 3D data generated by the 3D data generation unit 25 is image data such as a 3D point cloud or a depth image. The 3D data generation unit 25 outputs the generated 3D data to the learning unit 14. The learning unit 14 may use the 3D data input to the learning unit 14 for learning image transformation parameters.

[0103] The 2D data generation unit 24 may store the generated 2D data in the storage unit 16. The 3D data generation unit 25 may store the generated 3D data in the storage unit 16. The image conversion unit 12 may use the 2D data or 3D data, which are image data stored in the storage unit 16, for machine learning in the image conversion process.

[0104] The image data generated by the image data generation unit may be at least one of the following: 2D data representing a scene generated by the scene generation unit 18, and 3D data representing a scene generated by the scene generation unit 18. The dataset generation unit 19 is not limited to having both a 2D data generation unit 24 and a 3D data generation unit 25. The dataset generation unit 19 may be equipped with at least one of the 2D data generation unit 24 and the 3D data generation unit 25.

[0105] The annotation data generation unit 26 generates annotation data to be attached to the image data generated by the image data generation unit. The annotation data includes labeling data representing the color when coloring objects, information indicating the position and orientation of objects, or information on the type of object. Coloring objects refers to cases such as when multiple types of objects are included in multiple objects, and the objects are color-coded according to their type. The type of object is the name that represents the type of object. The annotation data generation unit 26 outputs the generated annotation data to the learning unit 14. The learning unit 14 may use the annotation data input to the learning unit 14 for learning image transformation parameters.

[0106] The target transformation image generation unit 27 generates a target transformation image, which is the target image for image transformation from the image data. The target transformation image is grayscale. Ru The target transformation image is an image, color image, or depth image in which the edges of an object are emphasized. The target transformation image may also be a CG image in which the edges of an object are simplified based on features common to the edges of each of multiple objects. The target transformation image may also be a CG image in which the edges of an object are replaced with edges in which predetermined features are emphasized. The target transformation image may also be an image similar to these CG images. In the target transformation image, the shape of the object may be approximated to a primitive shape. The target transformation image generation unit 27 stores the generated target transformation image in the storage unit 16.

[0107] The learning unit 14 may read the target transformation image generated by the target transformation image generation unit 27 from the storage unit 16 and learn image transformation parameters using the image data input to the learning unit 14 and the target transformation image. The image transformation unit 12 may read the image data generated by the image data generation unit and the target transformation image generated by the target transformation image generation unit 27 from the storage unit 16 and perform image transformation processing using the image data and the target transformation image.

[0108] According to the above description, the dataset generated by the dataset generation unit 19 includes image data, annotation data, and target transformation images. However, the dataset generated by the dataset generation unit 19 is not limited to including image data, annotation data, and target transformation images. The dataset generated by the dataset generation unit 19 does not need to include at least one of the image data, annotation data, and target transformation images.

[0109] The data generated by the dataset generation unit 19 is data generated through simulation. Therefore, differences in noise levels and other characteristics may occur between the image data generated by the dataset generation unit 19 and the images actually captured. The dataset generation unit 19 may be equipped with a function to reproduce pre-measured noise in the image data it generates. The dataset generation unit 19 may also reproduce measured noise according to the sensor's measurement method, such as structured illumination or ToF.

[0110] The image conversion unit 12 may input the distance image, which is image data generated by the image data generation unit, and the CG image, which is the target conversion image generated by the target conversion image generation unit 27, into the neural network. This allows the image conversion unit 12 to learn image conversion parameters such that the smooth edges of the CG image are represented in the converted image.

[0111] The image conversion evaluation unit 15 may calculate an evaluation value using annotation data, which is labeling data attached to the ground truth region, which is the region of an object in the image data generated by the image data generation unit. This evaluation value is the evaluation value used when evaluating the performance of image conversion based on the accuracy of individual separation. The image conversion unit 12 may learn the image conversion parameters by reflecting this evaluation value in the loss function used for machine learning in the image conversion unit 12.

[0112] Furthermore, the method for utilizing the data generated by the dataset generation unit 19 for training is not limited to the method described in Embodiment 3. The object recognition device 10C may appropriately increase the variations in the method for utilizing the data generated by the dataset generation unit 19, in addition to the method described in Embodiment 3.

[0113] According to Embodiment 3, the object recognition device 10C includes a simulation condition setting unit 17 for setting simulation conditions, a scene generation unit 18 for simulating the generation of a scene, and a dataset generation unit 19 for generating a dataset used for learning image conversion parameters. The object recognition device 10C can automatically collect training data by simulating the actual environment in which an object is photographed. The object recognition device 10C can reduce the number of personnel and costs required for collecting training data, and can obtain training data that enables highly accurate learning. As a result, the object recognition device 10C achieves the effects of labor savings, cost reduction, and improved object recognition accuracy.

[0114] By using the object recognition device 10C for object recognition in picking operations by industrial robots, it becomes possible to reduce manpower at the site where the industrial robots are operated, shorten the time required for pre-operation verification of the industrial robots, and reduce the costs associated with pre-operation verification. By simulating the actual environment in which objects are photographed using the object recognition device 10C, the reproducibility of the simulation process before the introduction of industrial robots can be improved. This allows for accurate determination of the placement of industrial robots and verification of the operation of industrial robots.

[0115] Embodiment 4. Embodiment 4 describes an example in which parameters related to picking are automatically adjusted by a control device equipped with an object recognition device. Figure 10 is a diagram showing the functional configuration of the control device 60 according to Embodiment 4. In Embodiment 4, the same reference numerals are used for components that are the same as those in Embodiments 1 to 3, and the configurations that differ from Embodiments 1 to 3 are mainly described. In the following description, industrial robots will be simply referred to as robots.

[0116] The robot comprises a robot body 62 and a control device 60 that controls the robot body 62. Figure 10 shows the control device 60 and the robot body 62. The control device 60 comprises an object recognition device 10D and a robot control unit 61. The object recognition device 10D recognizes objects from images of objects that have been captured. The object recognition device 10D outputs the result of object recognition to the robot control unit 61. The robot control unit 61 controls the robot body 62 based on the result of object recognition by the object recognition device 10D.

[0117] The robot body 62 has a hand for grasping objects. The hand is not shown in the illustration. The robot control unit 61 causes the robot body 62 to perform an object grasping operation based on the grasping position detected by the object recognition device 10D. The robot body 62 grasps the object and picks it up. The robot picks up one object at a time from multiple objects by repeating this operation.

[0118] As an example, the robot control unit 61 picks up each of the multiple objects based on the object recognition result obtained by the object recognition device 10D. The method for determining the picking order for each of the multiple objects is not limited to that described in Embodiment 4. The robot control unit 61 may pick up each of the multiple objects stacked in the height direction, starting with the object at the highest position. Alternatively, the picking order may be determined based on the likelihood level and the height of the object. That is, the robot control unit 61 may pick up each of the multiple objects in order from the object with the highest likelihood and at the highest position. The criteria for determining the picking order may be changed as needed.

[0119] The robot body 62 may be equipped with force sensors, proximity sensors, or tactile sensors. Multiple objects may be picked by a single robot or by multiple robots.

[0120] The object recognition device 10D has the same configuration as the object recognition device 10C shown in Figure 6. Furthermore, the object recognition device 10D includes a parameter adjustment unit 31 for adjusting parameters related to picking. The parameter adjustment unit 31 adjusts the gripping parameters, which are parameters for the hand used for picking. The parameter adjustment unit 31 also adjusts the recognition parameters, which are parameters used in the recognition process during picking. The parameter adjustment unit 31 stores the adjusted gripping parameters and the adjusted recognition parameters in the storage unit 16. The object recognition device 10D automatically adjusts the parameters related to picking through simulation.

[0121] Figure 11 shows an example of the configuration of the parameter adjustment unit 31 of the control device 60 according to Embodiment 4. Figure 11 shows the storage unit 16, the data set generation unit 19, and the parameter adjustment unit 31. The parameter adjustment unit 31 comprises a gripping parameter adjustment unit 32 and a recognition parameter adjustment unit 33. The gripping parameter adjustment unit 32 adjusts the gripping parameters, which are parameters for the hand of the robot body 62 that grips objects. The recognition parameter adjustment unit 33 adjusts the recognition parameters, which are parameters used in the recognition of objects by the recognition unit 13.

[0122] The memory unit 16 stores the CAD model of the object. The gripping parameter adjustment unit 32 reads the CAD model of the object from the memory unit 16 and adjusts the gripping parameters using the CAD model of the object. The gripping parameter adjustment unit 32 stores the adjusted gripping parameters in the memory unit 16.

[0123] The robot body 62 has hands such as tweezers, parallel hands, or suction pads. The gripping parameter adjustment unit 32 adjusts the gripping parameters for these various hands. The gripping parameters include parameters such as the opening width of the hand or the width of the hand's claws. The gripping parameter adjustment unit 32 adjusts the values ​​of parameters such as the opening width or claw width to values ​​that make it easy to grip objects. The type of hand used to grip objects may be determined by the robot user.

[0124] The gripping parameters include several types of parameters related to the hand, such as the opening width parameter or the claw width parameter. In the following description, "gripping parameter" refers to each of these several types of parameters. In the following description, the several types of gripping parameters will also be referred to as "multiple gripping parameters". In Embodiment 4, the gripping parameter adjustment unit 32 adjusts each of the multiple gripping parameters by determining the value of each of the multiple gripping parameters.

[0125] The gripping parameter adjustment unit 32 may adjust all of the multiple gripping parameters simultaneously, or it may adjust the multiple gripping parameters individually in a predetermined order. In the following description, the gripping parameter adjustment unit 32 will be assumed to adjust the multiple gripping parameters individually in a predetermined order.

[0126] The recognition parameter adjustment unit 33 reads the object's CAD model and the gripping parameters adjusted by the gripping parameter adjustment unit 32 from the storage unit 16. The recognition parameter adjustment unit 33 receives the dataset generated by the dataset generation unit 19 as input. The recognition parameter adjustment unit 33 adjusts the recognition parameters using the CAD model, gripping parameters, and dataset. The recognition parameter adjustment unit 33 saves the adjusted recognition parameters to the storage unit 16.

[0127] The recognition parameter adjustment unit 33 adjusts, for example, recognition parameters that are parameters for detecting candidate gripping positions or the position and orientation of each object from 2D or 3D data. Alternatively, the recognition parameter adjustment unit 33 may adjust the parameters of a learning model for detecting gripping positions using machine learning. The recognition parameter adjustment unit 33 may also adjust the parameters of a learning model in machine learning used for auxiliary processing to improve the recognition rate. An example of auxiliary processing to improve the recognition rate is segmentation. In the following description, the recognition parameter adjustment unit 33 will mainly adjust parameters for detecting candidate gripping positions from 2D data.

[0128] The recognition parameters include several types of parameters used in object recognition by the recognition unit 13. In the following description, "recognition parameter" refers to each of these several types of parameters. In the following description, the several types of recognition parameters will also be referred to as "multiple recognition parameters". In Embodiment 4, the recognition parameter adjustment unit 33 adjusts each of the multiple recognition parameters by determining the value of each of the multiple recognition parameters.

[0129] The recognition parameter adjustment unit 33 may adjust all of the recognition parameters simultaneously, or it may adjust the recognition parameters individually in a predetermined order. In the following description, the recognition parameter adjustment unit 33 will be assumed to adjust the recognition parameters individually in a predetermined order.

[0130] Next, the functions of the gripping parameter adjustment unit 32 will be described in detail. Figure 12 is a diagram showing an example of the configuration of the gripping parameter adjustment unit 32 in the control device 60 according to Embodiment 4. Figure 12 shows the storage unit 16, the data set generation unit 19, and the gripping parameter adjustment unit 32 and recognition parameter adjustment unit 33 of the parameter adjustment unit 31.

[0131] The gripping parameter adjustment unit 32 includes a gripping parameter adjustment range determination unit 41, a gripping parameter change unit 42, a model rotation unit 43, a gripping evaluation unit 44, a gripping parameter value determination unit 45, and a gripping parameter adjustment completion determination unit 46.

[0132] The gripping parameter adjustment range determination unit 41 determines the adjustment range for each of the multiple gripping parameters. The gripping parameter adjustment range determination unit 41 determines the adjustment range based on a model or the like stored in the storage unit 16. The gripping parameter adjustment unit 32 adjusts the value of each of the multiple gripping parameters within the range of values ​​determined by the gripping parameter adjustment range determination unit 41. In other words, the adjustment range of the gripping parameters refers to the range within which the value of the gripping parameters is adjusted. The gripping parameter adjustment range determination unit 41 outputs information indicating the determined adjustment range to the gripping parameter change unit 42.

[0133] The gripping parameter changing unit 42 changes the gripping parameter to be adjusted from among multiple gripping parameters. The gripping parameter adjustment unit 32 sequentially switches each of the multiple gripping parameters in the gripping parameter changing unit 42 and adjusts each gripping parameter. In this way, the gripping parameter adjustment unit 32 adjusts multiple gripping parameters individually in a predetermined order. The gripping parameter changing unit 42 outputs information indicating the gripping parameter to be adjusted and information indicating the adjustment range of the gripping parameter to be adjusted to the gripping evaluation unit 44 via the model rotation unit 43.

[0134] The model rotation unit 43 rotates the model of the object. The model rotation unit 43 reads the model stored in the memory unit 16 and performs the process of rotating the read model.

[0135] The gripping evaluation unit 44 evaluates the quality of gripping when the hand grips an object, assuming that the object assumes the same posture as the model rotated by the model rotation unit 43. Based on information indicating the gripping parameters to be adjusted and information indicating the adjustment range of the gripping parameters to be adjusted, the gripping evaluation unit 44 evaluates the quality of gripping when the object is gripped while adjusting the values ​​of the gripping parameters within the adjustment range. The gripping evaluation unit 44 outputs the results of the gripping quality evaluation to the gripping parameter value determination unit 45.

[0136] The gripping parameter value determination unit 45 determines the value of each of the multiple gripping parameters based on the evaluation results from the gripping evaluation unit 44. Once the gripping parameter values ​​are determined, the gripping parameter value determination unit 45 outputs the determined values ​​to the gripping parameter adjustment completion determination unit 46.

[0137] The gripping parameter adjustment completion determination unit 46 determines whether the adjustment has been completed for all of the multiple gripping parameters. The gripping parameter adjustment completion determination unit 46 determines that the adjustment has been completed for all of the multiple gripping parameters because the values ​​of the gripping parameters have been determined for all of the multiple gripping parameters.

[0138] Figure 13 is a flowchart showing the procedure of processing performed by the gripping parameter adjustment unit 32 of the control device 60 according to Embodiment 4.

[0139] In step S41, the gripping parameter adjustment range determination unit 41 determines the adjustment range of the gripping parameter. The gripping parameter adjustment range determination unit 41 reads the model stored in the storage unit 16 and determines the adjustment range of the gripping parameter based on the size of the model, etc. The gripping parameter adjustment range determination unit 41 may also determine the initial value of the gripping parameter based on the size of the model, etc. The initial value of the gripping parameter is the value used as a reference when adjusting the value of the gripping parameter.

[0140] In step S42, the gripping parameter changing unit 42 changes the type of gripping parameter to be adjusted. That is, the gripping parameter changing unit 42 changes the gripping parameter to be adjusted from among a plurality of gripping parameters.

[0141] In step S43, the model rotation unit 43 rotates the model of the object. The model rotation unit 43 may, for example, rotate the model randomly. The model rotation unit 43 may also rotate the model in a predetermined order.

[0142] In step S44, the gripping evaluation unit 44 evaluates the quality of gripping when an object is gripped by the hand. The gripping evaluation unit 44 evaluates the quality of gripping by simulating a situation where an object is gripped by the hand while it is in the same posture as the model, by fitting the hand model to the model rotated by the model rotation unit 43. The gripping evaluation unit 44 evaluates the quality of gripping by, for example, calculating an evaluation function for gripping parameters. The evaluation function for gripping parameters is, for example, the slip occurrence frequency F Ang and interference frequency F Col It consists of two evaluation values: the frequency of deviation F. Ang This represents the frequency of misalignment between the object's principal axis and the gripping direction. Interference frequency F Col This represents the frequency of interference between the object and the hand.

[0143] If the model rotation by the model rotation unit 43 is not finished (step S45, No), the gripping parameter adjustment unit 32 returns to step S43. On the other hand, if the model rotation by the model rotation unit 43 is finished (step S45, Yes), in step S46, the gripping parameter value determination unit 45 determines the value of the gripping parameter. The gripping parameter value determination unit 45 determines the value of the gripping parameter based on the evaluation result in step S44. For example, the gripping parameter value determination unit 45 searches for a gripping parameter value that minimizes the evaluation function of the gripping parameter. The gripping parameter value determination unit 45 determines the value of the gripping parameter that minimizes the evaluation function of the gripping parameter as the adjusted gripping parameter value.

[0144] Next, the gripping parameter adjustment completion determination unit 46 determines whether the adjustment of all gripping parameters to be adjusted by the gripping parameter adjustment unit 32 has been completed. If the adjustment of all gripping parameters to be adjusted has not been completed (step S47, No), the gripping parameter adjustment unit 32 returns to step S42. The gripping parameter adjustment unit 32 executes the processes from step S42 to step S47 for the gripping parameters that have not been adjusted. On the other hand, if the adjustment of all gripping parameters to be adjusted has been completed (step S47, Yes), the gripping parameter adjustment unit 32 terminates the process according to the procedure shown in Figure 13.

[0145] Next, the functions of the recognition parameter adjustment unit 33 will be described in detail. Figure 14 is a diagram showing an example of the configuration of the recognition parameter adjustment unit 33 in the control device 60 according to Embodiment 4. Figure 14 shows the storage unit 16, the data set generation unit 19, and the gripping parameter adjustment unit 32 and the recognition parameter adjustment unit 33 of the parameter adjustment unit 31.

[0146] The recognition parameter adjustment unit 33 includes a recognition parameter adjustment range determination unit 51, a recognition parameter modification unit 52, a recognition trial unit 53, a recognition evaluation unit 54, a recognition parameter value determination unit 55, and a recognition parameter adjustment completion determination unit 56.

[0147] The recognition parameter adjustment range determination unit 51 determines the adjustment range for each of the multiple recognition parameters. The recognition parameter adjustment range determination unit 51 determines the adjustment range based on a model stored in the storage unit 16. The recognition parameter adjustment unit 33 adjusts the value of each of the multiple recognition parameters within the range of values ​​determined by the recognition parameter adjustment range determination unit 51. In other words, the adjustment range of the recognition parameters refers to the range within which the value of the recognition parameters is adjusted. The recognition parameter adjustment range determination unit 51 outputs information indicating the determined adjustment range to the recognition parameter change unit 52.

[0148] The recognition parameter modification unit 52 changes the recognition parameter to be adjusted from among multiple recognition parameters. The recognition parameter adjustment unit 33 sequentially switches each of the multiple recognition parameters in the recognition parameter modification unit 52 to adjust each recognition parameter. In this way, the recognition parameter adjustment unit 33 adjusts multiple recognition parameters individually in a predetermined order. The recognition parameter modification unit 52 outputs information indicating the recognition parameter to be adjusted and information indicating the adjustment range of the recognition parameter to be adjusted to the recognition trial unit 53.

[0149] The recognition trial unit 53 attempts recognition processing using the gripping parameters adjusted by the gripping parameter adjustment unit 32. The recognition trial unit 53 reads the gripping parameters adjusted by the gripping parameter adjustment unit 32 from the storage unit 16. The recognition trial unit 53 obtains a dataset from the dataset generation unit 19. The recognition trial unit 53 detects candidate gripping positions by attempting recognition processing using the gripping parameters on the image data contained in the dataset. Based on information indicating the recognition parameters to be adjusted and information indicating the adjustment range of the recognition parameters to be adjusted, the recognition trial unit 53 detects candidate gripping positions while adjusting the values ​​of the recognition parameters within the adjustment range. The recognition trial unit 53 outputs the detection result of candidate gripping positions, which is the result of attempting the recognition processing, to the recognition evaluation unit 54.

[0150] The recognition evaluation unit 54 evaluates the quality of the results obtained from the recognition trial unit 53. The recognition evaluation unit 54 evaluates the quality of the results obtained from the recognition trial unit 53, for example, by evaluating whether the candidate gripping positions detected by the recognition trial unit 53 are suitable as gripping positions. The recognition evaluation unit 54 outputs the results of the evaluation of the quality of the results obtained from the recognition trial to the recognition parameter value determination unit 55.

[0151] The recognition parameter value determination unit 55 determines the value of each of the multiple recognition parameters based on the evaluation results from the recognition evaluation unit 54. Once the values ​​of the recognition parameters are determined, the recognition parameter value determination unit 55 outputs the determined values ​​to the recognition parameter adjustment completion determination unit 56.

[0152] The recognition parameter adjustment completion determination unit 56 determines whether the adjustment has been completed for all of the multiple recognition parameters. The recognition parameter adjustment completion determination unit 56 determines that the adjustment has been completed for all of the multiple recognition parameters because the values ​​of the recognition parameters have been determined for all of them.

[0153] Figure 15 is a flowchart showing the procedure of processing performed by the recognition parameter adjustment unit 33 of the control device 60 according to Embodiment 4.

[0154] In step S51, the recognition parameter adjustment range determination unit 51 determines the adjustment range of the recognition parameters. The recognition parameter adjustment range determination unit 51 reads the model stored in the storage unit 16 and determines the adjustment range of the recognition parameters based on the size of the model, etc. The recognition parameter adjustment range determination unit 51 may also determine the initial value of the recognition parameters based on the size of the model, etc. The initial value of the recognition parameters is the value used as a reference when adjusting the value of the recognition parameters.

[0155] In step S52, the recognition parameter modification unit 52 changes the type of recognition parameter to be adjusted. That is, the recognition parameter modification unit 52 changes the recognition parameter to be adjusted among a group of recognition parameters.

[0156] In step S53, the recognition trial unit 53 attempts to perform recognition processing using the gripping parameters adjusted by the gripping parameter adjustment unit 32. The recognition trial unit 53 adjusts the values ​​of the recognition parameters within the adjustment range and detects candidate gripping positions for each value of the recognition parameters. If the recognition processing trial has not been completed for each value of the recognition parameters within the adjustment range (step S54, No), the recognition trial unit 53 returns to step S53 and continues adjusting the values ​​of the recognition parameters and attempting the recognition processing.

[0157] On the other hand, if the recognition process has been completed for each value of the recognition parameter within the adjustment range (step S54, Yes), in step S55, the recognition evaluation unit 54 evaluates the quality of the results of the recognition process trials in step S53. The recognition evaluation unit 54 evaluates the quality of the results of the recognition process trials, for example, by calculating an evaluation function for the recognition parameter.

[0158] The evaluation function for the recognition parameter is, for example, the amount of deviation E between the candidate gripping position detected in step S53 and the optimal gripping position of the object. Pos And the amount of displacement E between the principal axis of the object and the gripping direction. Ang And the evaluation value E for the number of recognized items. Num It consists of three evaluation values. The number of recognized objects is the number of objects recognized by the object recognition device 10D. Evaluation value E of the number of recognized objects Num This is an evaluation value regarding the number of recognized objects. In robotic picking, it is required that no objects are left behind. If the number of recognized objects is small, there are many objects that are not recognized among the objects to be picked, making it easy for objects to be left behind, and therefore the evaluation value of the number of recognized objects is E. Num The value will be small. On the other hand, if the number of recognized items is large, it means that more objects that are the target of picking have been recognized, and it is less likely that objects will be left behind, so the evaluation value of the number of recognized items E Num This value will be large. The optimal gripping position for an object is, for example, a position determined based on the object's center of gravity. The optimal gripping position for an object may also be a position arbitrarily determined by the user.

[0159] Next, in step S56, the recognition parameter value determination unit 55 determines the value of the recognition parameter. The recognition parameter value determination unit 55 determines the value of the recognition parameter based on the evaluation result in step S54. For example, the recognition parameter value determination unit 55 searches for the value of the recognition parameter that minimizes the evaluation function of the recognition parameter. The recognition parameter value determination unit 55 determines the value of the recognition parameter that minimizes the evaluation function of the recognition parameter as the adjusted value of the recognition parameter.

[0160] Next, the recognition parameter adjustment completion determination unit 56 determines whether the adjustment of all recognition parameters targeted for adjustment by the recognition parameter adjustment unit 33 has been completed. If the adjustment of all recognition parameters targeted for adjustment has not been completed (step S57, No), the recognition parameter adjustment unit 33 returns to step S52. The recognition parameter adjustment unit 33 executes the processes from step S52 to step S57 for the recognition parameters that have not been adjusted. On the other hand, if the adjustment of all recognition parameters targeted for adjustment has been completed (step S57, Yes), the recognition parameter adjustment unit 33 terminates the process according to the procedure shown in Figure 15.

[0161] Furthermore, the recognition unit 13 shown in Figure 10 may use recognition parameters adjusted by the recognition parameter adjustment unit 33 when detecting candidate gripping positions from the converted image obtained by converting from an image of an object. The recognition unit 13 can read the recognition parameters adjusted by the recognition parameter adjustment unit 33 from the storage unit 16 and use the read recognition parameters to detect candidate gripping positions.

[0162] The image conversion evaluation unit 15 shown in Figure 10 may evaluate the performance of image conversion using image conversion parameters based on the operation results of the robot body 62. The image conversion evaluation unit 15 calculates evaluation values ​​for evaluating the performance of image conversion based on the operation results when the robot body 62 is actually operated.

[0163] For example, by having the robot body 62 pick up an object, the probability that the robot body 62 successfully grasped the object, the operation time during which the robot body 62 performed the grasping action, and the reason why the robot body 62 failed to grasp the object are observed. The image conversion evaluation unit 15 evaluates the performance of the image conversion based on the operation results, which include information on the probability that the robot body 62 successfully grasped the object, information on the operation time during which the robot body 62 performed the grasping action, and information indicating the reason why the robot body 62 failed to grasp the object.

[0164] The image conversion evaluation unit 15 calculates an evaluation value for evaluating the performance of image conversion, for example, using the following equation (6). In equation (6), E is an evaluation value for evaluating the performance of image conversion, P is the gripping success rate, which is the probability of successfully gripping an object, T is the operation time during which the robot body 62 performed the action of gripping an object, and f1, f2, ... each represents the reason why the robot body 62 failed to grip the object. In equation (6), W1 is a weighting coefficient set for the gripping success rate, W2 is a weighting coefficient set for the operation time, W 31 ,W 32 ,··· represent the weight coefficients set for f1, f2,··· respectively.

[0165]

number

[0166] W is a weighting coefficient set for the cause of gripping failure. 31 ,W 32 The values ​​for ,... may be set to be small for hardware-related causes such as dropping during transport, and large for causes related to the image conversion result.

[0167] The operation results used for evaluation in the image conversion evaluation unit 15 only need to include at least one of the following: information on the probability that the robot body 62 succeeded in grasping the object, information on the operation time during which the robot body 62 performed the operation to grasp the object, and information indicating the cause of the robot body 62's failure to grasp the object. In addition, the operation results used for evaluation in the image conversion evaluation unit 15 may include information other than the above. The method for calculating the evaluation value for evaluating the performance of image conversion is not limited to the method described in Embodiment 4. The image conversion evaluation unit 15 may evaluate the performance of image conversion based on a combination of the operation results of the robot body 62 and the accuracy of individual separation. The image conversion evaluation unit 15 may also evaluate the performance of image conversion based on the results of the recognition processing by the recognition unit 13.

[0168] The image conversion evaluation unit 15 may evaluate the performance of the image conversion, i.e., the validity of the image conversion, by comparing the evaluation value calculated as described above with a pre-set threshold. The image conversion evaluation unit 15 evaluates the performance of the image conversion as high if the evaluation value obtained by equation (6) is less than the threshold. In other words, the image conversion evaluation unit 15 evaluates that the image conversion performed by the image conversion unit 12 is an image conversion that allows for the individual separation and recognition of objects.

[0169] On the other hand, the image conversion evaluation unit 15 evaluates the image conversion performance as low if the evaluation value obtained by equation (6) is equal to or greater than a threshold. In other words, the image conversion evaluation unit 15 evaluates that the image conversion performed by the image conversion unit 12 is not an image conversion that allows for the individual separation and recognition of objects. Note that the performance evaluation based on the evaluation value of the image conversion performance may be performed by a method other than comparing the evaluation value with a threshold. The evaluation method by the image conversion evaluation unit 15 is not limited to the method exemplified in Embodiment 4 and is arbitrary.

[0170] If the image transformation performed by the image transformation unit 12 is evaluated as not being an image transformation that allows for the individual separation and recognition of objects, the learning unit 14 may learn the image transformation parameters. The learning unit 14 learns the image transformation parameters using the image data stored in the storage unit 16 and the dataset generated by the dataset generation unit 19. In this way, the learning unit 14 learns the image transformation parameters based on the evaluation results by the image transformation evaluation unit 15.

[0171] According to Embodiment 4, the control device 60 includes a gripping parameter adjustment unit 32 that adjusts gripping parameters, which are parameters for the hand that grips an object, and a recognition parameter adjustment unit 33 that adjusts recognition parameters, which are parameters used in the recognition of an object by the recognition unit 13. The control device 60 can automatically adjust the gripping parameters and recognition parameters by simulating the actual environment in which the object is photographed. The control device 60 can reduce the personnel and costs required to adjust the gripping parameters or recognition parameters, and can improve the success rate of robot picking. As a result, the control device 60 can reduce manpower, reduce costs, and improve the success rate of picking. of This has the effect of enabling improvement.

[0172] The control device 60 includes an image conversion evaluation unit 15 that evaluates the performance of image conversion using image conversion parameters based on the operation results of the robot body 62. The control device 60 can verify the validity of the image conversion parameters based on operation results such as the probability of successfully grasping an object, the operation time for grasping the object, or the reason for failure to grasp the object. As a result, the control device 60 can obtain converted images in which objects can be easily separated and recognized individually, further improving the success rate of picking.

[0173] The use of the control device 60 in controlling industrial robots makes it possible to reduce the number of people working at the site where the industrial robots are operated, shorten the time required for pre-operation verification of the industrial robots, and reduce the costs required for pre-operation verification.

[0174] In the above description, the control device 60 is assumed to include an object recognition device 10D, but it is not limited to this. The control device 60 may also include an object recognition device 10A as shown in Figure 1, an object recognition device 10B as shown in Figure 3, or an object recognition device 10C as shown in Figure 6.

[0175] Next, the hardware configurations for realizing the object recognition devices 10A, 10B, 10C, and 10D according to Embodiments 1 to 4 will be described. Each of the object recognition devices 10A, 10B, 10C, and 10D has a processing unit which is realized by a processing circuit. The processing unit of object recognition device 10A includes an image acquisition unit 11, an image conversion unit 12, and a recognition unit 13 as shown in Figure 1. The processing unit of object recognition device 10B includes an image acquisition unit 11, an image conversion unit 12, a recognition unit 13, a learning unit 14, and an image conversion evaluation unit 15 as shown in Figure 3. The processing unit of object recognition device 10C includes an image acquisition unit 11, an image conversion unit 12, a recognition unit 13, a learning unit 14, an image conversion evaluation unit 15, a simulation condition setting unit, a scene generation unit 18, and a dataset generation unit 19 as shown in Figure 6. The processing unit of the object recognition device 10D includes an image acquisition unit 11, an image conversion unit 12, a recognition unit 13, a learning unit 14, an image conversion evaluation unit 15, a simulation condition setting unit 17, a scene generation unit 18, a dataset generation unit 19, and a parameter adjustment unit 31, as shown in Figure 10. The processing circuit may be a circuit in which a processor executes software, or it may be a dedicated circuit.

[0176] Figure 16 shows a first example of a hardware configuration for realizing the object recognition devices 10A, 10B, 10C, and 10D according to Embodiments 1 to 4. In the first example, the object recognition devices 10A, 10B, 10C, and 10D include a processing circuit 70 having a processor 72 and a memory 73, an input unit 71, and an output unit 74.

[0177] The input unit 71 is an interface circuit that receives data transmitted from outside the object recognition devices 10A, 10B, 10C, and 10D to the object recognition devices 10A, 10B, 10C, and 10D and provides it to the processor 72. The output unit 74 is an interface circuit that outputs data from the processor 72 or memory 73 to the outside of the object recognition devices 10A, 10B, 10C, and 10D.

[0178] Each of the object recognition devices 10A, 10B, 10C, and 10D has a processing unit that is implemented by software, firmware, or a combination of software and firmware. The software or firmware is written as a program and stored in memory 73. In the processing circuit 70, each function of the processing unit is realized when the processor 72 reads and executes the program stored in memory 73. In other words, the processing circuit 70 is equipped with memory 73 for storing a program that will result in the execution of processing by each function of the processing unit. This program can be said to be a program that instructs the computer to execute the procedures and methods of the object recognition devices 10A, 10B, 10C, and 10D. Memory 73 is also used as temporary memory when the processor 72 executes various processes. The storage unit 16 of the object recognition devices 10B, 10C, and 10D is implemented by memory 73.

[0179] The processor 72 includes, for example, one or more of the following: CPU (Central Processing Unit), DSP (Digital Signal Processor), and system LSI (Large Scale Integration). The memory 73 includes one or more of the following: RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), and EEPROM® (Electrically Erasable Programmable Read Only Memory). The memory 73 also includes a recording medium on which a computer-readable program is recorded. Such a recording medium includes one or more of the following: non-volatile or volatile semiconductor memory, magnetic disk, flexible memory, optical disk, compact disk, and DVD (Digital Versatile Disc).

[0180] Figure 16 shows an example of hardware in which a processing unit is implemented using a general-purpose processor 72 and memory 73, but the processing unit may also be implemented using dedicated hardware circuits. Figure 17 shows a second example of a hardware configuration for implementing the object recognition devices 10A, 10B, 10C, and 10D according to embodiments 1 to 4. In the second example, the object recognition devices 10A, 10B, 10C, and 10D include a processing circuit 75, which is a dedicated hardware circuit, an input unit 71, and an output unit 74.

[0181] The processing circuit 75 may be a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or a circuit combining these. Each function of the processing unit may be implemented by the processing circuit 75 individually, or all functions may be implemented together by the processing circuit 75. Each function of the processing unit may also be implemented by a combination of the processing circuit 70 and the processing circuit 75.

[0182] The control device 60 according to Embodiment 4 is implemented with a hardware configuration similar to the hardware configuration illustrated in Figure 16, or the hardware configuration similar to the hardware configuration illustrated in Figure 17. Of the control device 60, the processing unit of the object recognition device 10D and the robot control unit 61 are implemented by a processing circuit 70 or a processing circuit 75. The output unit 74 transmits data to the robot body 62.

[0183] The configurations shown in each of the embodiments described above are examples of the content of this disclosure. The configurations of each embodiment can be combined with other known technologies. The configurations of each embodiment may be combined with each other as appropriate. It is possible to omit or modify parts of the configurations of each embodiment without departing from the gist of this disclosure. [Explanation of Symbols]

[0184] 10A, 10B, 10C, 10D Object recognition device, 11 Image acquisition unit, 12 Image conversion unit, 13 Recognition unit, 14 Learning unit, 15 Image conversion evaluation unit, 16 Memory unit, 17 Simulation condition setting unit, 18 Scene generation unit, 19 Data set generation unit, 21 Shooting equipment information setting unit, 22 Object information setting unit, 23 Environment information setting unit, 24 2D data generation unit, 25 3D data generation unit, 26 Annotation data generation unit, 27 Target conversion image generation unit, 31 Parameter adjustment unit, 32 Grasping parameter adjustment unit, 33 Recognition parameter adjustment unit, 41 Grasping parameter adjustment range determination unit, 42 Grasping parameter change unit, 43 Model rotation unit, 44 Grasping evaluation unit, 45 Grasping parameter value determination unit, 46 Grasping parameter adjustment completion determination unit, 51 Recognition parameter adjustment range determination unit, 52 Recognition parameter change unit, 53 Recognition trial unit, 54 Recognition evaluation unit, 55 Recognition parameter value determination unit, 56 Recognition parameter adjustment completion determination unit, 60 Control device, 61 Robot control unit, 62 Robot body, 70, 75 Processing circuits, 71 Input unit, 72 Processor, 73 Memory, 74 Output unit.

Claims

1. An image acquisition unit that acquires images of objects, An image conversion unit that converts the image into a converted image in which the edges of the multiple objects are replaced with edges that have been enhanced to highlight features common to the edges of each of the objects, The system comprises a recognition unit that recognizes the object based on the converted image, The image conversion unit determines the shape obtained by eliminating random elements that actually exist at the edges of each of the multiple objects as a common feature of each of the multiple objects, and converts the image to the converted image by replacing the shape of each object with the shape obtained as a feature. An object recognition device characterized by the following features.

2. The image conversion unit converts the image to the converted image based on the results of learning the original image and the converted image. The object recognition device according to feature 1.

3. The image conversion unit converts the image to the converted image based on image conversion parameters that are learned from the original image and the converted image. The object recognition device according to feature 2.

4. The system includes a learning unit that learns the image conversion parameters used for converting the aforementioned image to the converted image, The image conversion unit converts the image to the converted image based on the image conversion parameters, which are the result of learning by the learning unit. The object recognition device according to feature 3.

5. The system includes an image conversion evaluation unit that evaluates the performance of image conversion using the aforementioned image conversion parameters. The object recognition device according to feature 4.

6. The learning unit learns the image conversion parameters based on the evaluation results from the image conversion evaluation unit. The object recognition device according to feature 5.

7. A simulation condition setting unit sets simulation conditions for simulating the imaging of the aforementioned object, which include at least one of imaging equipment information, which is information about the equipment used to capture the image, and object information, which is information about the object. A scene generation unit that simulates and generates a scene in which the object is photographed based on the aforementioned simulation conditions, The system comprises a dataset generation unit that generates a dataset used for learning the image transformation parameters based on the scene generated by the scene generation unit. The object recognition device according to any one of claims 4 to 6.

8. The aforementioned simulation condition setting unit is: A camera equipment information setting unit sets the camera equipment information, which includes at least one of the specifications of the equipment and the installation of the equipment. The object information setting unit sets the aforementioned object information, The system includes an environmental information setting unit that sets environmental information, which is information about the environment during shooting in the simulation, and sets the simulation conditions, which include the shooting equipment information, the object information and the environmental information. The object recognition device according to feature 7.

9. The aforementioned dataset generation unit, An image data generation unit generates image data which is at least one of two-dimensional data representing the scene generated by the scene generation unit and three-dimensional data representing the scene generated by the scene generation unit. An annotation data generation unit that generates annotation data to be attached to the aforementioned image data, The system includes a target transformation image generation unit that generates a target transformation image, which is the target image for image transformation from the aforementioned image data. The object recognition device according to feature 7.

10. The learning unit learns the image transformation parameters based on the target transformation image and the annotation data. The object recognition device according to feature 9.

11. An image acquisition unit that acquires images of objects, An image conversion unit that converts the image into a converted image in which the edges of the multiple objects are replaced with edges that have been enhanced to highlight features common to the edges of each of the objects, A recognition unit that recognizes the object based on the converted image, The system comprises a robot control unit that controls the robot body that grasps the object recognized by the recognition unit, The image conversion unit determines the shape obtained by eliminating random elements that actually exist at the edges of each of the multiple objects as a common feature of each of the multiple objects, and converts the image to the converted image by replacing the shape of each object with the shape obtained as a feature. A control device characterized by the following features.

12. A simulation condition setting unit sets simulation conditions for simulating the imaging of the aforementioned object, which include at least one of imaging equipment information, which is information about the equipment used to capture the image, and object information, which is information about the object. A scene generation unit that simulates and generates a scene in which the object is photographed based on the aforementioned simulation conditions, The dataset generation unit generates a dataset used to learn image transformation parameters used for transforming an image to a transformed image, based on the scene generated by the scene generation unit. The control device according to feature 11.

13. The robot body includes a gripping parameter adjustment unit for adjusting the gripping parameters, which are parameters for the hand that grips the object. The control device according to claim 12.

14. The gripping parameter adjustment unit is, A gripping parameter adjustment range determination unit that determines the adjustment range for each of the multiple gripping parameters, A gripping parameter changing unit that changes the gripping parameter to be adjusted among a plurality of gripping parameters, A model rotation unit that rotates the model of the aforementioned object, A gripping evaluation unit that evaluates the quality of gripping when the hand grips the object, assuming that the object is made to assume the same posture as the model rotated by the model rotation unit, A grip parameter value determination unit that determines the value of each of the plurality of grip parameters based on the evaluation results from the grip evaluation unit, The system includes a gripping parameter adjustment completion determination unit that determines whether or not adjustment has been completed for all of the multiple gripping parameters, and adjusts each of the multiple gripping parameters by determining the value of each of the multiple gripping parameters. The control device according to claim 13.

15. The recognition parameter adjustment unit is provided to adjust the recognition parameters, which are parameters used in the recognition of the object by the recognition unit. The control device according to claim 12.

16. The recognition parameter adjustment unit is provided to adjust the recognition parameters, which are parameters used in the recognition of the object by the recognition unit. The aforementioned recognition parameter adjustment unit is A recognition parameter adjustment range determination unit that determines the adjustment range for each of the multiple recognition parameters, A recognition parameter changing unit that changes the recognition parameter to be adjusted among a plurality of gripping parameters, A recognition trial unit attempts to perform recognition processing using the gripping parameters adjusted by the gripping parameter adjustment unit, A recognition evaluation unit that evaluates the quality of the results of the aforementioned recognition process, A recognition parameter value determination unit determines the value of each of the plurality of recognition parameters based on the evaluation result by the recognition evaluation unit, The system includes a recognition parameter adjustment completion determination unit that determines whether or not adjustment has been completed for all of the multiple recognition parameters, and adjusts each of the multiple recognition parameters by determining the value of each of the multiple recognition parameters. The control device according to claim 13 or 14.

17. The system includes an image conversion evaluation unit that evaluates the performance of image conversion using the aforementioned image conversion parameters based on the operating results of the robot body. A control device according to any one of 12 to 15, characterized by the above.

18. The operation result includes at least one of the following: information on the probability that the robot body successfully grasped the object; information on the duration of the operation in which the robot body performed the action of grasping the object; and information indicating the reason why the robot body failed to grasp the object. The control device according to feature 17.

19. A method for recognizing objects using a computer, The steps include: acquiring an image of the aforementioned object, The steps include transforming the image into a transformed image in which the edges of the multiple objects are replaced with edges that have been enhanced to highlight features common to the edges of each of the multiple objects, The step of recognizing the object based on the converted image includes, In the step of converting the aforementioned image to the converted image, the shape obtained by eliminating the random elements actually present at each edge of the multiple objects is determined as a feature common to each edge of the multiple objects, and the shape of each object is replaced with the shape determined as the feature, thereby converting the aforementioned image to the converted image. A method for recognizing objects characterized by the following features.