Generation device and generation method

The generation device and method address the challenge of inaccurate defect position estimation in appearance inspection by generating teacher data through simulated 3DCG images, which reduces visual differences and improves defect detection accuracy.

WO2025110014A1PCT designated stage expired Publication Date: 2025-05-30IHI CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/039562
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing appearance inspection methods using deep learning models face challenges in accurately determining the position of defects due to significant visual differences between simulated and actual images, primarily caused by mismatched surface optical characteristics.

Method used

A generation device and method that create teacher data by generating a virtual object with a defect, simulating various light sources, and producing 3DCG images. These images are then used to derive direction vectors, which are utilized to generate normal map images as teacher data, improving defect position estimation accuracy.

Benefits of technology

The proposed solution enhances the accuracy of defect position estimation in appearance inspection by reducing the visual difference between simulated and actual images, thereby improving the performance of deep learning models in defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039562_30052025_PF_FP_ABST
    Figure JP2024039562_30052025_PF_FP_ABST
Patent Text Reader

Abstract

An image processing device (generation device) 500 comprises: a 3D graphic model generation unit (virtual body generation unit) 500a that generates a virtual body representing the three-dimensional shape of an object for inspection in a virtual space on the basis of three-dimensional shape data of the object for inspection; a defect imparting unit 500b that imparts a defect to the virtual body; a 3D CG image generation unit 500c that generates at least three 3D CG images having different virtual light sources on the basis of the virtual body to which the defect has been imparted, at least three virtual light sources that illuminate the virtual body to which the defect has been imparted, and a virtual imaging device that captures an image of the illuminated virtual body; and a first image generation unit 500d that generates a first image as training data on the basis of a direction vector of the virtual body's surface derived on the basis of the at least three 3D CG images.
Need to check novelty before this filing date? Find Prior Art

Description

Generating device and generating method

[0001] This application claims the benefit of priority from Japanese Patent Application No. 2023-198454, filed on November 22, 2023, the contents of which are incorporated herein by reference.

[0002] Conventionally, visual inspection of an object to be inspected is performed using an imaging device. In visual inspection, the image captured by the imaging device is input into a pre-trained deep learning model (hereinafter also referred to as an inspection model), and the location of defects, which are abnormalities in the object to be inspected, can be obtained as an output.

[0003] Patent Document 1 discloses a technology in which a light source and an object to be inspected are virtually arranged in a virtual space, and images generated by computer simulation are used as training data. The technology described in Patent Document 1 uses the training data to train an inspection model for visual inspection of the object to be inspected, and by providing a captured image of the object to the trained inspection model, it is possible to automatically determine whether the appearance of the object to be inspected is good or bad.

[0004] JP 2019-215240 A

[0005] However, when comparing an image generated by simulation with an image captured by a real imaging device, there is a large visual difference between the two images. This is because it is difficult to match the surface optical properties of the actual object to those of the simulated object, which determine the appearance of the object, such as the specular reflectance and diffuse reflectance of the material of the object.

[0006] Therefore, even if images generated by simulation are used as training data as in Patent Document 1, there is a problem in that the accuracy of the position of defects in the object to be inspected estimated by the inspection model trained using the training data is poor.

[0007] In consideration of the above-mentioned problems, the present disclosure aims to provide a device and method for generating teacher data that can improve the accuracy of the estimated position of defects in an object to be inspected.

[0008] In order to solve the above problem, a generation device according to one aspect of the present disclosure includes a virtual body generation unit that generates a virtual body that represents the three-dimensional shape of an object to be inspected in a virtual space based on three-dimensional shape data of the object to be inspected; a defect imparting unit that imparts defects to the virtual body; a 3DCG image generation unit that generates at least three 3DCG images with different virtual light sources based on the virtual body imparted with the defects, at least three virtual light sources that illuminate the virtual body imparted with the defects, and a virtual imaging device that images the illuminated virtual body; and a first image generation unit that generates a first image as training data based on a direction vector of the surface of the virtual body derived based on the at least three 3DCG images.

[0009] The inspection system may further include an image generation unit that generates at least three captured images with different light sources based on an object to be inspected, at least three light sources that illuminate the object to be inspected, and an imaging device that images the illuminated object to be inspected, and a second image generation unit that generates a second image as training data based on a direction vector of the surface of the object to be inspected derived based on the at least three captured images, and the three light sources may be positioned at positions corresponding to the positions of the three virtual light sources, respectively.

[0010] The training data may be data in which correct answer data indicating the position of the defect is associated with the first image.

[0011] In order to solve the above problem, a generation method according to one aspect of the present disclosure includes the steps of generating a virtual body that represents the three-dimensional shape of an object to be inspected in a virtual space based on three-dimensional shape data of the object to be inspected; adding defects to the virtual body; generating at least three 3DCG images with different virtual light sources based on the virtual body with the defects added, at least three virtual light sources that illuminate the virtual body with the defects added, and a virtual imaging device that images the illuminated virtual body; and generating a first image as training data based on the direction vector of the surface of the virtual body derived based on the at least three 3DCG images.

[0012] According to the present disclosure, it is possible to improve the accuracy of the estimated position of a defect in an inspection object.

[0013] FIG. 1 is a schematic configuration diagram of an inspection system according to the present embodiment. FIG. 2 is a schematic block diagram of the inspection system according to the present embodiment. FIG. 3 is a block diagram showing an example of the functional configuration of an image processing device according to the present embodiment. FIG. 4 is a schematic configuration diagram showing an example of a 3D graphic model generated in a virtual space. FIG. 5 is a diagram showing an example of a first 3DCG image. FIG. 6 is a diagram showing an example of a second 3DCG image. FIG. 7 is a diagram showing an example of a third 3DCG image. FIG. 8 is a diagram showing an example of a first image. FIG. 9 is an explanatory diagram for explaining a directional vector of a surface of an inspection object illuminated by light from a first virtual light source. FIG. 10 is an explanatory diagram for explaining a directional vector of a surface of an inspection object illuminated by light from a second virtual light source. FIG. 11 is an explanatory diagram for explaining a directional vector of a surface of an inspection object illuminated by light from a third virtual light source. FIG. 12 is a diagram showing an example of an inspection model. FIG. 13 is a flowchart showing an example of a method for generating training data according to the present embodiment. Fig. 14 is a diagram showing an example of a captured image of an inspection object captured with all of a plurality of light sources turned on, and a 3DCG image of a virtual body captured with all of a plurality of virtual light sources turned on. Fig. 15 is a diagram showing an example of a normal map image generated by sequentially turning on a plurality of light sources for an inspection object, and a normal map image generated by sequentially turning on a plurality of virtual light sources for a virtual body.

[0014] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Specific dimensions, materials, numerical values, etc. shown in the embodiments are merely examples for ease of understanding and do not limit the present disclosure unless otherwise specified. In this specification and drawings, elements having substantially the same functions and configurations are designated by the same reference numerals to avoid redundant explanation, and elements not directly related to the present disclosure are not shown.

[0015] Fig. 1 is a schematic configuration diagram of an inspection system 100 according to this embodiment. The inspection system 100 is a system for performing a visual inspection of an inspection target object 200. As shown in Fig. 1, the inspection system 100 includes an imaging device 300, a plurality of light sources 400, and an image processing device 500.

[0016] The inspection object 200 is an actual product that is placed in real space and undergoes visual inspection by the inspection system 100. In this embodiment, the inspection object 200 is, for example, a general industrial product. However, the inspection object 200 is not limited to an industrial product as long as it can be imaged by the imaging device 300.

[0017] In this embodiment, a defect 210, which is an abnormal portion, may be included in a portion of the inspection object 200. The defect 210 is formed in a portion of the surface of the inspection object 200, and is, for example, a protrusion or depression formed on the surface of the inspection object 200. The defect 210 in the inspection object 200 is generated, for example, by a reduction in thickness of the inspection object 200 due to corrosion, a foreign object being caught during press working of the inspection object 200, or a tool colliding with the surface of the inspection object 200.

[0018] The imaging device 300 is a real camera placed in real space. The camera viewpoint of the imaging device 300 is, for example, located above the inspection object 200, and the imaging direction of the imaging device 300 is, for example, downward toward the inspection object 200. In this embodiment, the imaging device 300 images the inspection object 200 from above. However, without being limited to this, the imaging device 300 may image the inspection object 200 from below or from the side.

[0019] The imaging device 300 captures an image of the inspection object 200 and generates a captured image of the inspection object 200. The imaging device 300 outputs the generated captured image of the inspection object 200 to the image processing device 500. In this embodiment, one imaging device 300 is arranged for one inspection object 200. One imaging device 300 is arranged facing the inspection object 200 so that the inspection object 200 is included in the imaging range. With the position of the imaging device 300 fixed, multiple captured images of the inspection object 200 are captured.

[0020] The multiple light sources 400 include at least three light sources. The reason for using at least three light sources is to capture images of the inspection object 200 using a photometric stereo method. The photometric stereo method is a three-dimensional measurement technique in which images of the inspection object 200 illuminated from a plurality of different illumination directions are captured and a normal vector, which is a directional vector of the surface of the inspection object 200, is obtained from the shading information. Here, the normal vector is a vector perpendicular to the surface of the inspection object 200. The photometric stereo method is also called photometric stereo.

[0021] The multiple light sources 400 include a first light source 400A, a second light source 400B, and a third light source 400C. The first light source 400A illuminates the inspection object 200 from a first direction L1. The second light source 400B illuminates the inspection object 200 from a second direction L2 different from the first direction L1. The third light source 400C illuminates the inspection object 200 from a third direction L3 different from the first direction L1 and the second direction L2. With the positions of the first light source 400A, the second light source 400B, and the third light source 400C fixed, the inspection object 200 is illuminated by light emitted from any one of the first light source 400A, the second light source 400B, and the third light source 400C.

[0022] The first light source 400A, the second light source 400B, and the third light source 400C are actual light sources arranged in real space. The first light source 400A, the second light source 400B, and the third light source 400C are arranged at 90° intervals, for example, 0°, 90°, or 180°, in the circumferential direction of the inspection object 200. However, without being limited thereto, the first light source 400A, the second light source 400B, and the third light source 400C may be arranged at equal 120° intervals, for example, 0°, 120°, or 240°, in the circumferential direction of the inspection object 200. Note that when the multiple light sources 400 are composed of four light sources, the four light sources may be arranged at equal 90° intervals, for example, 0°, 90°, 180°, or 270°, in the circumferential direction of the inspection object 200.

[0023] In this way, the first light source 400A, the second light source 400B, and the third light source 400C illuminate the inspection object 200 from different directions. In this embodiment, the number of the plurality of light sources 400 is three. However, the number of the plurality of light sources 400 is not limited to three as long as it is at least three or more.

[0024] The image processing device 500 is electrically connected to the image capturing device 300 and the plurality of light sources 400, and controls the image capturing device 300 and the plurality of light sources 400. The image processing device 500 controls the image capturing device 300 and the plurality of light sources 400 so as to illuminate the inspection object 200 from at least three different illumination directions at at least three different timings in order to capture an image of the inspection object 200 by the photometric stereo method.

[0025] Specifically, the image processing device 500 controls the first light source 400A to irradiate the inspection object 200 with light in a first direction L1 at a first timing, thereby acquiring a first captured image of the illuminated inspection object 200. The image processing device 500 also controls the second light source 400B to irradiate the inspection object 200 with light in a second direction L2 at a second timing different from the first timing, thereby acquiring a second captured image of the illuminated inspection object 200. The second timing is, for example, a timing after the first timing. The image processing device 500 also controls the third light source 400C to irradiate the inspection object 200 with light in a third direction L3 at a third timing different from the first and second timings, thereby acquiring a third captured image of the illuminated inspection object 200. The third timing is, for example, a timing after the second timing.

[0026] 2 is a schematic block diagram of the inspection system 100 according to this embodiment. As shown in FIG. 2, the image processing device 500 includes an I / F 510, a storage device 520, a system bus 530, one or more processors 540, and one or more memories 550. The I / F 510 is an interface for communicating with the imaging device 300, the first light source 400A, the second light source 400B, and the third light source 400C. Note that the image processing device 500 according to this embodiment also functions as a generation device that generates training data, as will be described in detail later.

[0027] The storage device 520 is composed of RAM, flash memory, HDD, etc., and holds various information necessary for processing by the processor 540, which will be described below. Specifically, the storage device 520 stores three-dimensional shape data of the inspection object 200 and a defect 610 to be assigned to a virtual body 600, which will be described later. The three-dimensional shape data represents the three-dimensional shapes of the inspection object 200 and the defect 610, is expressed using a polygon mesh, and has coordinate information in three-dimensional space. The three-dimensional shape data in this embodiment is, for example, 3D CAD data. However, the three-dimensional shape data is not limited to this, and may also be 3D data obtained by measuring the inspection object 200 and the defect 210 using a three-dimensional optical measuring device such as a laser displacement meter or a stereo imaging device, an X-ray CT device, or the like.

[0028] The system bus 530 is a transmission path that electrically connects the I / F 510, the storage device 520, the processor 540, and the memory 550 and transmits data among them.

[0029] The processor 540 includes, for example, a CPU (Central Processing Unit). The memory 550 includes, for example, a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM is a storage element that stores programs and calculation parameters used by the CPU. The RAM is a storage element that temporarily stores data such as variables and parameters used in processing executed by the CPU.

[0030] 3 is a block diagram showing an example of the functional configuration of an image processing device 500 according to this embodiment. For example, as shown in FIG. 3, the image processing device 500 includes a 3D graphic model generation unit (virtual body generation unit) 500a, a defect adding unit 500b, a 3DCG image generation unit 500c, a first image generation unit 500d, a captured image generation unit 500e, a second image generation unit 500f, a teacher data generation unit 500g, a model learning unit 500h, and an estimation detection unit 500i.

[0031] 2 cooperates with a program stored in memory 550 and executes the program stored in memory 550. This realizes various processes including the processes described below that are performed by the 3D graphic model generation unit 500a, the defect adding unit 500b, the 3DCG image generation unit 500c, the first image generation unit 500d, the captured image generation unit 500e, the second image generation unit 500f, the teacher data generation unit 500g, the model learning unit 500h, and the estimation detection unit 500i.

[0032] The 3D graphic model generating unit 500a generates a virtual body 600, which is a 3D graphic model that represents the three-dimensional shape of the inspection object 200 in a virtual space S, based on the three-dimensional shape data of the inspection object 200 stored in the storage device 520. The position and orientation of the virtual body 600 that imitates the inspection object 200 and is provided in the virtual space S is set to be the same as the position and orientation of the actual inspection object 200 shown in FIG.

[0033] Furthermore, the 3D graphic model generation unit 500a sets the position of the camera viewpoint and the imaging direction of the virtual imaging device 700 in the virtual space S. Here, the position of the camera viewpoint and the imaging direction of the virtual imaging device 700 are set to be the same as the camera viewpoint and the imaging direction of the actual imaging device 300 shown in FIG.

[0034] Furthermore, 3D graphic model generation unit 500a sets the position and illumination direction of first virtual light source 800A in virtual space S. Here, the position and illumination direction of first virtual light source 800A are set to be the same as the position and illumination direction of real first light source 400A shown in FIG. 1 . 3D graphic model generation unit 500a sets the position and illumination direction of second virtual light source 800B in virtual space S. Here, the position and illumination direction of second virtual light source 800B are set to be the same as the position and illumination direction of real second light source 400B shown in FIG. 1 . 3D graphic model generation unit 500a sets the position and illumination direction of third virtual light source 800C in virtual space S. Here, the position and illumination direction of third virtual light source 800C are set to be the same as the position and illumination direction of real third light source 400C shown in FIG. 1 . First virtual light source 800A, second virtual light source 800B, and third virtual light source 800C are also collectively referred to as multiple virtual light sources 800.

[0035] FIG. 4 is a schematic diagram showing an example of a 3D graphic model generated in a virtual space S. In FIG. 4, the X, Y, and Z directions are perpendicular to one another. The X and Y directions are horizontal directions, and the Z direction is an up-down direction. As shown in FIG. 4, a virtual body 600 simulating the inspection object 200, a virtual imaging device 700 simulating the imaging device 300, and a plurality of virtual light sources 800 simulating the plurality of light sources 400 are arranged in the virtual space S. Note that the plurality of virtual light sources 800 includes at least three virtual light sources in order to capture an image of the virtual body 600 simulating the inspection object 200 by a photometric stereo method.

[0036] In the XYZ coordinate system within virtual space S, virtual body 600 is disposed at a position corresponding to inspection target object 200 shown in Fig. 1 . Furthermore, in the XYZ coordinate system within virtual space S, virtual imaging device 700 is disposed at a position corresponding to imaging device 300 shown in Fig. 1 . In the XYZ coordinate system within virtual space S, first virtual light source 800A is disposed at a position corresponding to first light source 400A shown in Fig. 1 . In the XYZ coordinate system within virtual space S, second virtual light source 800B is disposed at a position corresponding to second light source 400B shown in Fig. 1 . In the XYZ coordinate system within virtual space S, third virtual light source 800C is disposed at a position corresponding to third light source 400C shown in Fig. 1 .

[0037] In this embodiment, virtual imaging device 700 is disposed on an extension line in the +Z direction with respect to virtual body 600. Furthermore, first virtual light source 800A is disposed on the +Y direction side with respect to virtual body 600. In this case, first direction L1, in which light from first virtual light source 800A is emitted, is the −Y direction and a direction tilted at −45 degrees toward the −Z side with respect to the horizontal when the position of first virtual light source 800A is taken as the reference position. In other words, first direction L1 is the −Y direction with respect to the imaging direction of virtual imaging device 700 and a direction tilted at −45 degrees toward the −Z side with respect to the horizontal. Second virtual light source 800B is disposed on the +X direction side with respect to virtual body 600. In this case, second direction L2, in which light from second virtual light source 800B is emitted, is the −X direction and a direction tilted at −45 degrees toward the −Z side with respect to the horizontal when the position of second virtual light source 800B is taken as the reference position. In other words, second direction L2 is in the −X direction with respect to the imaging direction of virtual imaging device 700 and is inclined at −45 degrees to the −Z side with respect to the horizontal. Third virtual light source 800C is disposed on the −Y direction side with respect to virtual body 600. In this case, third direction L3 in which light from third virtual light source 800C is irradiated is in the +Y direction with respect to the horizontal and is inclined at −45 degrees to the −Z side with respect to the horizontal when the position of third virtual light source 800C is taken as the reference position. In other words, third direction L3 is in the +Y direction with respect to the imaging direction of virtual imaging device 700 and is inclined at −45 degrees to the −Z side with respect to the horizontal.

[0038] In this way, the virtual body 600 of this embodiment is positioned so that it can be imaged from directly above by the virtual imaging device 700, and so that it can be illuminated by the first virtual light source 800A, the second virtual light source 800B, and the third virtual light source 800C at 90° intervals around the circumference from a diagonal 45° angle above.

[0039] Note that the imaging direction of imaging device 300 and the first direction L1, second direction L2, and third direction L3 of multiple light sources 400 shown in Figure 1 are the same as the imaging direction of virtual imaging device 700 and the first direction L1, second direction L2, and third direction L3 of multiple virtual light sources 800 shown in Figure 4.

[0040] The defect adding unit 500b adds a defect 610 as an abnormal portion to the virtual body 600 generated by the 3D graphic model generating unit 500a. Specifically, the defect adding unit 500b forms the defect 610, which is expressed as a three-dimensional shape as a depression, on the surface of the virtual body 600.

[0041] The three-dimensional shape data of the depression defect 610 is generated, for example, by randomly generating a depth distribution using uniform random numbers and smoothing the depth distribution to prevent sudden changes such as edges. By applying a window function when generating the data for the defect 610, the boundary between the defect portion and the non-defect portion on the surface of the virtual body 600 can be smoothed. By randomly varying the depth distribution of the defect 610, three-dimensional shape data for defects 610 with a wide variety of patterns can be easily generated.

[0042] Specifically, the defect assignment unit 500b varies the value corresponding to the depth of each coordinate of the defect 610 in three-dimensional space based on a probability distribution, randomly generates different three-dimensional shape data for the defect 610, and assigns a wide variety of defects 610 to the virtual body 600. For example, the defect assignment unit 500b uses a value corresponding to the depth of each coordinate of the defect 610 in three-dimensional space as a variable, and changes the variable by adding one of a value of +a, +0.5a, -0.5a, or -a to generate three-dimensional shape data for each of the coordinates. Here, a is a predetermined value, and the defect assignment unit 500b varies the value corresponding to the depth of the defect 610 using a probability distribution in which a random variable of 1 / 4 is assigned to each of +a, +0.5a, -0.5a, and -a. Note that while an example of changing the depth of the defect 610 has been shown here, it is also possible to change not only the depth of the defect 610 but also its position. That is, the defect adding unit 500b changes the shape parameters of the defects 610 based on the probability distribution to generate a wide variety of types of virtual bodies 600. In this way, a wide variety of types of virtual bodies 600 having the wide variety of defects 610 are generated.

[0043] In this way, the defect adding unit 500b changes the parameters of the defect 610, thereby generating a plurality of types of virtual bodies 600 having a wide variety of defects 610. The parameters of the defect 610 are the values ​​of the coordinates of the defect 610 in three-dimensional space, and changing the parameters of the defect 610 makes it possible to change, for example, the position, depth, shape, etc. of the defect 610.

[0044] The three-dimensional shape data of the defect 610 may be obtained by measuring the defect 210, such as a dent or a scratch, that has occurred in the actual inspection object 200 using a three-dimensional optical measuring device such as a laser displacement meter or a stereo imaging device, an X-ray CT device, etc. The three-dimensional shape data of the defect 610 that has been generated or obtained in this manner is stored in the storage device 520.

[0045] The 3DCG image generation unit 500c emits light from one of the multiple virtual light sources 800, virtually captures an image of the virtual body 600 illuminated by the light from the one virtual light source using the virtual imaging device 700, and generates one 3DCG image. Specifically, the 3DCG image generation unit 500c generates a 3DCG image by performing physically based rendering on the virtual body 600. Physically based rendering is a method of calculating an image captured by the virtual imaging device 700 based on real-world optical laws such as reflection, transmission, and diffusion. In this way, the 3DCG image generation unit 500c generates a 3DCG image based on the virtual body 600 to which the defect 610 has been added.

[0046] 3DCG image generation unit 500c virtually captures an image of virtual body 600 using virtual imaging device 700 while changing the irradiation direction of light from multiple virtual light sources 800, and generates multiple 3DCG images. Specifically, 3DCG image generation unit 500c first emits light from first virtual light source 800A of multiple virtual light sources 800, and virtually captures an image of virtual body 600 illuminated by the light from first virtual light source 800A using virtual imaging device 700, thereby generating first 3DCG image 710. Next, 3DCG image generation unit 500c emits light from second virtual light source 800B of multiple virtual light sources 800, and virtually captures an image of virtual body 600 illuminated by the light from second virtual light source 800B using virtual imaging device 700, thereby generating second 3DCG image 720. Finally, the 3DCG image generation unit 500c emits light from a third virtual light source 800C among the multiple virtual light sources 800, virtually captures an image of the virtual body 600 illuminated by the light from the third virtual light source 800C using the virtual imaging device 700, and generates a third 3DCG image 730.

[0047] The first 3DCG image 710, the second 3DCG image 720, and the third 3DCG image 730 are images virtually captured at different times. Note that the first 3DCG image 710, the second 3DCG image 720, and the third 3DCG image 730 are images virtually captured using a photometric stereo method under the same or similar capturing conditions as the first captured image, the second captured image, and the third captured image.

[0048] Fig. 5 is a diagram showing an example of a first 3DCG image 710. Fig. 6 is a diagram showing an example of a second 3DCG image 720. Fig. 7 is a diagram showing an example of a third 3DCG image 730. As shown in Figs. 5 to 7, the first 3DCG image 710, the second 3DCG image 720, and the third 3DCG image 730 include images of the virtual body 600 and the defect 610 imparted to the surface of the virtual body 600.

[0049] As shown in Fig. 5 , virtual body 600 is illuminated by light emitted from first virtual light source 800A in first direction L1. As a result, as shown in Fig. 5 , a shadow region SH, indicated by hatching, is formed at the end of virtual body 600 on the first direction L1 side. Furthermore, a shadow region SH, indicated by hatching, is formed inside the depression on the first direction L1 side of defect 610. Furthermore, no shadow region SH is formed inside the depression on the opposite side of defect 610 from first direction L1 due to the illumination of light.

[0050] As shown in Fig. 6 , virtual body 600 is illuminated by light emitted from second virtual light source 800B in second direction L2. As a result, as shown in Fig. 6 , a shadow region SH, indicated by hatching, is formed at the end of virtual body 600 on the second direction L2 side. Furthermore, a shadow region SH, indicated by hatching, is formed inside the depression on the second direction L2 side of defect 610. Furthermore, no shadow region SH is formed inside the depression on the opposite side of defect 610 from second direction L2 due to the illumination of light.

[0051] As shown in Fig. 7 , virtual body 600 is illuminated by light emitted from third virtual light source 800C in third direction L3. As a result, as shown in Fig. 7 , a shadow region SH, indicated by hatching, is formed at the end of virtual body 600 on the third direction L3 side. Furthermore, a shadow region SH, indicated by hatching, is formed inside the depression on the third direction L3 side of defect 610. Furthermore, no shadow region SH is formed inside the depression on the opposite side of defect 610 from third direction L3 due to the illumination of light.

[0052] The first image generation unit 500d generates the first image 740 based on the direction vector of the surface of the virtual body 600 derived based on the first 3DCG image 710, the second 3DCG image 720, and the third 3DCG image 730. A method for generating the first image 740 will be described in detail below.

[0053] FIG. 8 is a diagram showing an example of a first image 740. The first image 740 is a so-called normal map image. The normal map image is an image in which a direction vector n of the surface of the inspection object 200 or the virtual body 600 is derived for each pixel of a captured image of the inspection object 200 or a 3DCG image of the virtual body 600, and visualized as RGB pixel values. The direction vector n of the surface of the inspection object 200 or the virtual body 600 is derived using the following equation (1). Note that the first method of deriving the direction vector n of the surface of the virtual body 600 and generating the first image 740, which is a normal map image, is similar to the second method of deriving the direction vector n of the surface of the inspection object 200 and generating the second image, which is a normal map image. Therefore, the first method will be described in detail below, and a detailed description of the second method will be omitted.

[0054] Fig. 9 is an explanatory diagram for explaining direction vector n of the surface of virtual body 600 illuminated by light from first virtual light source 800A. Fig. 10 is an explanatory diagram for explaining direction vector n of the surface of virtual body 600 illuminated by light from second virtual light source 800B. Fig. 11 is an explanatory diagram for explaining direction vector n of the surface of virtual body 600 illuminated by light from third virtual light source 800C.

[0055] 9, l = (a1, a2, a3) indicates the direction vector of first virtual light source 800A. i = A indicates a luminance value, which is the pixel value of each pixel of first 3DCG image 710. n = (x, y, z) indicates the direction vector of the surface. The relationship between direction vector l of first virtual light source 800A, luminance value i, and direction vector n of the surface is i = n l. In other words, luminance value i is expressed as the dot product of direction vector l of first virtual light source 800A and direction vector n of the surface.

[0056] 10, l = (b1, b2, b3) indicates the direction vector of second virtual light source 800B. i = B indicates the luminance value, which is the pixel value of each pixel of second 3DCG image 720. n = (x, y, z) indicates the direction vector of the surface. The relationship between direction vector l, luminance value i, and direction vector n of the surface of second virtual light source 800B is i = n l. In other words, the luminance value i is expressed as the dot product of direction vector l of second virtual light source 800B and direction vector n of the surface.

[0057] 11 , l = (c1, c2, c3) indicates the direction vector of third virtual light source 800C. i = C indicates a luminance value, which is the pixel value of each pixel of third 3DCG image 730. n = (x, y, z) indicates the direction vector of the surface. The relationship between direction vector l, luminance value i, and direction vector n of the surface of third virtual light source 800C is i = n · l. In other words, the luminance value i is expressed as the dot product of direction vector l of third virtual light source 800C and direction vector n of the surface.

[0058] Since the direction vector l and brightness value i of the first virtual light source 800A, the second virtual light source 800B, and the third virtual light source 800C are known, the direction vector n of the surface can be derived by solving the simultaneous equations of the above formula (1).

[0059] To generate a normal map image, it is necessary to visualize the direction vector n of the surface derived for each pixel as RGB pixel values. In this embodiment, the normal map image is derived using the following formulas (2) to (5).

[0060] The following formula (2) expresses the pixels of the first 3DCG image 710, the second 3DCG image 720, and the third 3DCG image 730 arranged in a row as a matrix I. In formula (2), img1 receives the luminance values ​​of pixels 1 to p of the first 3DCG image 710. img2 receives the luminance values ​​of pixels 1 to p of the second 3DCG image 720. img3 receives the luminance values ​​of pixels 1 to p of the third 3DCG image 730. The luminance values ​​input to pixels 1 to p are values ​​in the range of 0 to 255 normalized to values ​​in the range of 0.0 to 1.0.

[0061] The following formula (3) is a pseudo-inverse matrix L obtained by converting a matrix L representing the direction vectors of first virtual light source 800A, second virtual light source 800B, and third virtual light source 800C. -1 Here, a case will be described in which first direction L1 of first virtual light source 800A, second direction L2 of second virtual light source 800B, and third direction L3 of third virtual light source 800C are all inclined at 45 degrees to the +Z side with respect to the horizontal direction with virtual body 600 as the reference. However, this is not limiting, and first direction L1, second direction L2, and third direction L3 may be inclined to the -Z side with respect to the horizontal direction, or may be inclined at an angle other than 45 degrees. Furthermore, in equation (3), matrix L is converted into a pseudo-inverse matrix L -1 However, the matrix L may be converted into an inverse matrix.

[0062] The following formula (4) is the matrix I of formula (2) and the pseudo-inverse matrix L of formula (3). -1 The matrix product L -1 I. Furthermore, the following formula (5) represents the magnitude of the vector at each pixel in the matrix product of formula (4) converted to 1.

[0063] The normal map image is obtained by normalizing the values ​​in the range from the minimum value to the maximum value of the matrix in the above formula (5) to values ​​in the range from 0 to 255, so that (R, G, B) = (x, y, z). In this way, the first image generation unit 500d can generate the first image 740, which is the normal map image shown in Fig. 8, based on the first 3DCG image 710, second 3DCG image 720, and third 3DCG image 730 shown in Figs. 5 to 7.

[0064] The captured image generating unit 500e generates at least three captured images using different light sources based on at least a first light source 400A, a second light source 400B, and a third light source 400C that illuminate the inspection object 200 and an imaging device 300 that captures an image of the illuminated inspection object 200. The at least three captured images include a first captured image, a second captured image, and a third captured image. The first captured image is an image captured by the imaging device 300 of the inspection object 200 illuminated by light irradiated from the first light source 400A in a first direction L1. The second captured image is an image captured by the imaging device 300 of the inspection object 200 illuminated by light irradiated from the second light source 400B in a second direction L2. The third captured image is an image captured by the imaging device 300 of the inspection object 200 illuminated by light irradiated from the third light source 400C in a third direction L3.

[0065] The second image generation unit 500f generates a second image based on a direction vector n of the surface of the inspection object 200 derived based on at least three captured images generated by the captured image generation unit 500e. Specifically, the second image generation unit 500f generates a second image, which is a normal map image, based on the direction vector n of the surface of the inspection object 200 derived based on the first captured image, the second captured image, and the third captured image. The normal map image, which is the second image, is also derived using the above formulas (2) to (5), similar to the normal map image, which is the first image 740. In generating the second image, the terms used in generating the first image 740 are replaced as follows: "virtual body 600" is replaced with "inspection object 200." "first 3DCG image 710" is replaced with "first captured image." "second 3DCG image 720" is replaced with "second captured image." "Third 3DCG image 730" is to be read as "third captured image." "First virtual light source 800A" is to be read as "first light source 400A." "Second virtual light source 800B" is to be read as "second light source 400B." "Third virtual light source 800C" is to be read as "third light source 400C." The method for generating the second image is the same as the method for generating first image 740, and therefore detailed description thereof will be omitted.

[0066] The teacher data generation unit 500g generates correct answer data indicating the position of the defect 610 in the first image 740 based on the defect information of the defect 610 assigned by the defect assignment unit 500b. Here, the three-dimensional shape data of the defect 610 includes information on coordinates constituting the defect 610 in three-dimensional space, and also includes position information indicating the position of the defect 610 and depth information indicating its depth. Information about the defect 610, including the position information indicating the position of the defect 610 and the depth information indicating its depth, is also referred to as defect information. The defect assignment unit 500b assigns the defect 610 to the surface of the virtual body 600 using the three-dimensional shape data including the defect information of the defect 610 stored in the storage device 520. Because the position and depth distribution of the defect 610 assigned by the defect assignment unit 500b are known, the teacher data generation unit 500g can generate correct answer data based on the defect information of the defect 610.

[0067] The correct answer data is associated with a pixel corresponding to the position of the defect 610 in the first image 740 as an abnormal portion, and for example, a numerical value "1" is associated with it. Also, the correct answer data is associated with a pixel corresponding to a position other than the defect 610 in the first image 740 as a normal portion, and for example, a numerical value "0" is associated with it.

[0068] Furthermore, the correct answer data is associated with each pixel by a numerical value between "0" and "1" other than "0" based on the depth distribution at the position of the defect 610 in the first image 740. For example, the correct answer data is such that, for a pixel corresponding to the position of the defect 610, the deeper the defect 610, the closer to "1" the correct answer data is associated, and the shallower the defect 610, the closer to "0" the correct answer data is associated.

[0069] In the first image 740, pixel values ​​change depending on the distance from the imaging surface of the virtual imaging device 700 to the virtual body 600. For example, the pixel values ​​of the first image 740 are expressed in 256 gradations ranging from "0" to "255." Here, for example, the closer the distance from the imaging surface of the virtual imaging device 700, the closer the pixel value is to "0," and the farther the distance from the imaging surface, the closer the pixel value is to "255." Because the pixel values ​​of the first image 740 include distance information related to the distance from the imaging surface of the virtual imaging device 700, the depth distribution of the defect 610 can be expressed by converting the pixel values.

[0070] For example, the deeper the defect 610, i.e., the closer the pixel value is to 255, the closer to "1" the correct answer data is associated with the pixel corresponding to the position of the defect 610. Also, the shallower the defect 610, i.e., the closer the pixel value is to 0, the closer to "0" the correct answer data is associated with the pixel corresponding to the position of the defect 610.

[0071] The teacher data generation unit 500g associates the first image 740 with the correct answer data, collectively defines the teacher data, and stores the teacher data in the storage device 520. In this embodiment, the teacher data generation unit 500g generates teacher data that associates defect information about the defect 610, which includes at least the position information and depth information of the defect 610, with the first image 740. The teacher data generation unit 500g generates a plurality of such teacher data.

[0072] Specifically, the 3DCG image generation unit 500c generates a plurality of types of first 3DCG images 710, second 3DCG images 720, and third 3DCG images 730 based on a plurality of types of virtual bodies 600 having a wide variety of defects 610. At this time, the 3DCG image generation unit 500c randomly changes parameters of the imaging conditions of the virtual imaging device 700 to generate a wide variety of first 3DCG images 710, second 3DCG images 720, and third 3DCG images 730. The imaging conditions include the surface optical characteristics of the inspection object 200 in the virtual body 600, the orientation of the imaging surface of the virtual imaging device 700, the orientation of the virtual light source 800, the light intensity of the virtual light source 800, etc.

[0073] The surface optical characteristics of the inspection object 200 include, for example, surface reflectance, diffuse reflectance, surface roughness, etc. However, changing the parameters of the imaging conditions is not a necessary condition, and the 3DCG image generation unit 500c may generate multiple types of first 3DCG images 710, second 3DCG images 720, and third 3DCG images 730 based on multiple types of virtual bodies 600 in which only the parameters of the defects 610 have been changed.

[0074] A plurality of pieces of teacher data can be generated by changing at least the parameters of the defect 610. The teacher data generating unit 500g stores the generated plurality of types of teacher data in the storage device 520.

[0075] The model learning unit 500h inputs multiple types of teacher data stored in the storage device 520 into the test model M, and trains the test model M so that output data that is close to the correct answer data contained in the teacher data is obtained.

[0076] The inspection model M of this embodiment includes, for example, a neural network (NN). The neural network is a convolutional neural network (CNN) trained by supervised learning. However, a neural network other than a convolutional neural network may also be used. Furthermore, a learning model other than a neural network may also be used.

[0077] The test model M is, for example, a deep learning model in image analysis. In this embodiment, the test model M is configured by, for example, a combination of a neural network structure and parameters that represent the strength of the connections between each neuron. Each connection between neurons is provided with a parameter that is a coefficient. Each parameter is configured to be adjustable. The internal state of the test model M is represented by a set of numerical values ​​that are a combination of the neural network structure, called internal variables, and the parameters between each neuron.

[0078] 12 is a diagram showing an example of the test model M. As shown in FIG. 12, training data is input to the test model M, and each neuron N 1 , N 2 , N 3 , N 4 , N 5 , N M-1 , N M By passing through the above, output data is output in which the output node value 0 to 1 is associated with each pixel as described above.

[0079] The inspection model M is trained based on training data including the first image 740 to which correct answer data is linked in advance. The training data is input to the inspection model M, and the values ​​of the internal variables of the inspection model M are set so as to reduce the error between the output data of the inspection model M and the correct answer data linked to the first image 740 of the training data. In this way, the inspection model M is trained based on multiple types of training data.

[0080] The model learning unit 500h may use the first image 740 to which no supervised data is linked as training data. That is, the model learning unit 500h may input the first image 740 to which no supervised data is linked to the inspection model M, and train the inspection model M so that the error between the information indicating the position of the defect 610 output from the inspection model M and the supervised data is reduced. In this case, the first image generation unit 500d functions as a training data generation unit.

[0081] Furthermore, the model learning unit 500h may use the second image as training data in addition to the first image 740. That is, the model learning unit 500h may input the first image 740 and the second image to the inspection model M, and train the inspection model M so that the error between the information indicating the position of the defect 610 output from the inspection model M and the ground truth data is reduced. In this case, the first image generation unit 500d and the second image generation unit 500f function as training data generation units.

[0082] The estimation detection unit 500i inputs the second image generated by the second image generation unit 500f based on the first captured image, the second captured image, and the third captured image to the inspection model M. At this time, the inspection model M classifies each pixel of the second image, and a value greater than "0" among the numerical values ​​"0" to "1" is associated as an output node with the defect 210 in the second image, i.e., the pixel corresponding to the defect 210. Furthermore, a numerical value "0" is associated as an output node with the pixel corresponding to a normal portion of the second image. Thus, in this embodiment, the second image input to the inspection model M is output in the output format of semantic segmentation.

[0083] The estimation detection unit 500i estimates the position of the defect 210 in the inspection object 200 included in the second image based on the output data of the inspection model M, i.e., the numerical values ​​"0" to "1" as the output nodes, and further estimates the depth of the defect 210. Specifically, the estimation detection unit 500i estimates the position of a pixel associated with a numerical value of an output node other than "0" as the position of the defect 210. Furthermore, the estimation detection unit 500i estimates the depth of the defect 210 according to the magnitude of the numerical value of the output node other than "0". In this way, the estimation detection unit 500i estimates the position and depth distribution of the defect 210 according to the position of a pixel associated with a numerical value of an output node other than "0" and the magnitude of the numerical value of the output node other than "0".

[0084] The estimation detection unit 500i displays information indicating the estimated position of the defect 210 in the inspection object 200 and information indicating the estimated depth of the defect 210 (hereinafter also referred to as estimation result data) on a display (not shown). For example, the estimation detection unit 500i adds a predetermined color to the position of the defect 210 and superimposes it on the second image displayed on the display. At this time, the estimation detection unit 500i may superimpose and display the second image displayed on the display by changing the color depending on the depth of the defect 210. For example, the estimation detection unit 500i may superimpose and display color information on the second image so that the deeper the defect 210, the darker the color, and the shallower the defect 210, the lighter the color.

[0085] 13 is a flowchart showing an example of a method for generating training data according to this embodiment. As shown in FIG. 13, the 3D graphic model generating unit 500a generates a virtual body 600 that represents the three-dimensional shape of the inspection object 200 in a virtual space S based on the three-dimensional shape data of the inspection object 200 (step S100).

[0086] The defect adding unit 500b adds a defect 610 to the virtual body 600 (step S102). The 3DCG image generating unit 500c generates at least three 3DCG images with different virtual light sources (step S104). The first image generating unit 500d generates a first image 740, which is a normal map image, based on a direction vector n of the surface of the virtual body 600 derived from the at least three 3DCG images (step S106).

[0087] The training data generation unit 500g generates training data in which the first image 740 is associated with the supervised data indicating the position of the defect 610 (step S108). The model learning unit 500h inputs only the training data generated by the training data generation unit 500g into the inspection model M, and trains the inspection model M (step S110).

[0088] The captured image generation unit generates at least three captured images with different light sources (step S112). The second image generation unit 500f generates a second image, which is a normal map image, based on the direction vector of the surface of the inspection object 200 derived based on the at least three captured images (step S114).

[0089] The estimation detection unit 500i inputs the second image, which is a normal map image, into the trained inspection model M, and estimates the position of the defect 210 of the inspection object 200 contained in the second image (step S116).

[0090] As described above, the image processing device 500 of this embodiment includes a first image generation unit 500d. The first image generation unit 500d generates a first image 740, which is a normal map image, based on a direction vector n of the surface of the virtual body 600 derived from at least three 3DCG images captured using a photometric stereo method and having different virtual light sources. The image processing device 500 also includes a second image generation unit 500f. The second image generation unit 500f generates a second image, which is a normal map image, based on a direction vector n of the surface of the inspection object 200 derived from at least three captured images captured using a photometric stereo method and having different light sources. The model learning unit 500h inputs the first image 740 as training data into the inspection model M and causes it to learn. The estimation detection unit 500i inputs a second image into the trained inspection model M and estimates the position of the defect 210 in the inspection object 200 included in the second image.

[0091] As described above, according to this embodiment, a virtual body 600 simulating the inspection object 200 is imaged in the virtual space S using a photometric stereo method to generate a first image 740, which is a normal map image. The generated first image 740 is then used as training data to train the inspection model M. The actual inspection object 200 is then imaged in the real space using a photometric stereo method to generate a second image, which is a normal map image. The generated second image is input to the trained inspection model M, which estimates the position of the defect 210 in the inspection object 200 contained in the second image. In this way, the first image 740 and the second image input to the inspection model M are both normal map images of the same type.

[0092] A large amount of training data, for example, hundreds to tens of thousands of images, is required to train the inspection model M used in the visual inspection of the inspection object 200. However, when the inspection object 200 is a general industrial product, it is difficult to obtain a large amount of training data for each inspection object 200 due to the variety of product shapes.

[0093] In this embodiment, the defect adding unit 500b adds defects 610 to a virtual body 600 that imitates the inspection target 200. The defect adding unit 500b changes the parameters of the defects 610 and adds them to the virtual body 600, thereby making it possible to generate a large amount of training data through simulation. Therefore, even when there is a small amount of training data, a large amount of training data can be generated through simulation, allowing the inspection model M to learn with high accuracy.

[0094] Furthermore, even if a large amount of training data can be generated by simulation, there may be a large difference in appearance between the images generated by the simulation that serve as training data and the images captured by an imaging device in real space. In such cases, even if the images generated by the simulation are used as training data, there is a problem in that the accuracy of the defect positions of the inspection object estimated by the inspection model trained using the training data is low.

[0095] In this embodiment, the first image 740 and the second image input to the inspection model M are both normal map images of the same type. Therefore, the difference in appearance between the first image 740 and the second image can be reduced, and as a result, the accuracy of the position of the defect 210 in the inspection object 200 estimated by the trained inspection model M can be improved.

[0096] 14 is a diagram showing an example of a captured image of the inspection object 200 captured with all of the plurality of light sources 400 turned on, and a 3DCG image of the virtual body 600 captured with all of the plurality of virtual light sources 800 turned on. Fig. 15 is a diagram showing an example of a normal map image generated by sequentially turning on the plurality of light sources 400 for the inspection object 200, and a normal map image generated by sequentially turning on the plurality of virtual light sources 800 for the virtual body 600.

[0097] 14, the image on the left is a captured image of inspection object 200 captured with all of the plurality of light sources 400 turned on, and the image on the right is a 3DCG image of virtual body 600 captured with all of the plurality of virtual light sources 800 turned on. Also, in Fig. 15, the image on the left is a normal map image generated by sequentially turning on the plurality of light sources 400 for inspection object 200, and the image on the right is a normal map image generated by sequentially turning on the plurality of virtual light sources 800 for virtual body 600.

[0098] The result of deriving the similarity between the captured image and the 3DCG image shown in Fig. 14 using SSIM (Structural Similarity) was 0.715758. In contrast, the result of deriving the similarity between the two normal map images shown in Fig. 15 using SSIM was 0.806890. In other words, it can be seen that the similarity between the two normal map images shown in Fig. 15 is higher than the similarity between the captured image and the 3DCG image shown in Fig. 14. Therefore, according to this embodiment, it is possible to reduce the appearance difference between normal map images of the same type compared to when the captured image captured by the imaging device 300 and the 3DCG image captured by the virtual imaging device 700 are used as is.

[0099] Furthermore, according to this embodiment, when training the inspection model M, a second image generated based on at least three captured images captured by the imaging device 300 can be used as training data in parallel with the training data of the first image 740. Since the second image is generated from a captured image of the actual inspection object 200 in real space, the shape of the defect 210 actually formed in the inspection object 200 can be trained by the inspection model M. As a result, the accuracy of the position of the defect 210 in the inspection object 200 estimated by the trained inspection model M can be improved.

[0100] Furthermore, according to this embodiment, when training the inspection model M, training data in which ground truth data indicating the position of the defect 610 is associated with the first image 740 is used. This makes it possible to improve the accuracy of the position of the defect 610 estimated by the inspection model M.

[0101] Although the embodiments have been described above with reference to the accompanying drawings, the present disclosure is not limited to the above-described embodiments. It is clear that a person skilled in the art can conceive of various modifications or alterations within the scope of the claims, and it is understood that such modifications also fall within the technical scope of the present disclosure.

[0102] The present disclosure can contribute, for example, to Goal 12 of the Sustainable Development Goals (SDGs), "Ensure sustainable consumption and production patterns."

[0103] REFERENCE SIGNS LIST 100 Inspection system 200 Inspection object 300 Imaging device 400 Multiple light sources 400A First light source 400B Second light source 400C Third light source 500 Image processing device 600 Virtual object 700 Virtual imaging device 800 Multiple virtual light sources 800A First virtual light source 800B Second virtual light source 800C Third virtual light source

Claims

1. A generation device comprising: a virtual body generation unit that generates a virtual body that represents the three-dimensional shape of an inspection object in a virtual space based on three-dimensional shape data of the inspection object; a defect imparting unit that imparts defects to the virtual body; a 3DCG image generation unit that generates at least three 3DCG images with different virtual light sources based on the virtual body with the defects imparted, at least three virtual light sources that illuminate the virtual body with the defects imparted, and a virtual imaging device that images the illuminated virtual body; and a first image generation unit that generates a first image as teacher data based on a directional vector of the surface of the virtual body derived based on the at least three 3DCG images.

2. The generating device described in claim 1, further comprising: an image generating unit that generates at least three captured images with different light sources based on the object to be inspected, at least three light sources that illuminate the object to be inspected, and an imaging device that images the illuminated object to be inspected; and a second image generating unit that generates a second image as the teacher data based on a directional vector of the surface of the object to be inspected derived based on the at least three captured images, wherein the three light sources are positioned at positions corresponding to the positions of the three virtual light sources, respectively.

3. The generating device according to claim 1 or 2, wherein the training data is data in which correct answer data indicating the position of the defect is associated with the first image.

4. A generation method comprising the steps of: generating a virtual body that represents the three-dimensional shape of an object to be inspected in a virtual space based on three-dimensional shape data of the object to be inspected; adding a defect to the virtual body; generating at least three 3DCG images having different virtual light sources based on the virtual body with the defect added, at least three virtual light sources that illuminate the virtual body with the defect added, and a virtual imaging device that images the illuminated virtual body; and generating a first image as teacher data based on a directional vector of the surface of the virtual body derived based on the at least three 3DCG images.

Citation Information

Patent Citations

  • Image processing method, image processing device, image processing system, imaging device, program, and storage medium

    JP2020119333A

  • Imaging control device, evaluation system, imaging control method, and program

    JP2021093712A

  • Computer vision method and system

    JP2022032937A

  • Information processing device and information processing method

    WO2021095672A1

  • Image processing device, image processing method, and program

    WO2021229984A1