Training data generation device and training data generation method
By synchronizing gradients and Z coordinates in the synthesis of 3D image portions, the method prevents unintended features at image boundaries, improving prediction accuracy in training data generation for machine learning.
Patent Information
- Application Number
- JP2024141020
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
Existing methods for synthesizing 3D images in training data generation for machine learning can result in unintended features at the boundary between images, leading to decreased prediction accuracy of the model.
The method involves generating training image data by synthesizing portions of 3D images in the XY coordinate plane while ensuring the gradient and Z coordinates of the objects in the synthesis area match those of the original images, thereby preventing the formation of steps at the image boundary.
This approach suppresses the deterioration of prediction accuracy by maintaining consistent gradients and coordinates, thus enhancing the model's performance.
Smart Images

Figure 2026037765000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a training data generation device and a training data generation method for generating image data used in machine learning. [Background technology]
[0002] Patent Document 1 discloses a training data generation device that uses first image data representing a first 3D image and second image data representing a second 3D image to generate training image data to be used in machine learning to create a model by synthesizing the second 3D image with the first 3D image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2021 / 193347 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in Patent Document 1, depending on the method of synthesizing the first 3D image and the second 3D image, unintended features such as a step may appear at the boundary between the first and second 3D images in the synthesized image shown by the training image data, which may deteriorate the accuracy of predictions made using the model.
[0005] The present invention has been made in consideration of these points, and its purpose is to suppress deterioration in the prediction accuracy of the model caused by the appearance of unintended features at the boundary between the first and second 3D images in the composite image. [Means for solving the problem]
[0006] In order to achieve the above object, the present disclosure provides a training data generation process for generating training image data to be used in machine learning to create a model by using first image data representing a first 3D image and second image data representing a second 3D image to synthesize at least a portion of an extracted image in the XY coordinate plane of the second 3D image with a portion of an area in the XY coordinate plane of the first 3D image, wherein the training image data is generated based on the first image data, the second image data, and segmentation information that identifies the position of the extracted image in the second 3D image, so that the gradient of the object included in the synthesis area of the extracted image in the synthesis image represented by the training image data is equal to the gradient of the corresponding object in the extracted image, and the Z coordinate of each point of the object in the boundary area adjacent to the synthesis area in the XY coordinate plane in the synthesis image is equal to the Z coordinate of the corresponding point in the first 3D image.
[0007] This makes the gradient of the object included in the synthesis area in the synthetic image equal to the gradient of the corresponding object in the second 3D image, thereby allowing the shape of the object in the second 3D image to be reflected in the synthesis area in the synthetic image.
[0008] Furthermore, since the Z coordinates of each point of the photographed object in the boundary region adjacent to the synthesis region in the XY coordinate plane in the synthesis image are set equal to the Z coordinates of the corresponding points in the first 3D image, a step as an unintended feature is not formed at the boundary between the first and second 3D images in the synthesis image, thereby suppressing deterioration in the prediction accuracy of the model due to the occurrence of an unintended feature at the boundary between the first and second 3D images in the synthesis image. [Effects of the Invention]
[0009] According to the present disclosure, deterioration of the prediction accuracy of the model can be suppressed. [Brief explanation of the drawings]
[0010] [Figure 1]FIG. 1 is a block diagram showing the configuration of a welding system including an AI model generation device as a learning data generation device according to a first embodiment of the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating the first 3D image. [Figure 3] FIG. 3 is a diagram illustrating a second 3D image. [Figure 4] FIG. 4 is a view corresponding to FIG. 2 in which the region where the weld bead is formed is painted black. [Figure 5] FIG. 5 is a view corresponding to FIG. 3 showing a welding defect. [Figure 6] FIG. 6 is a flowchart illustrating the operation of the welding system. [Figure 7] FIG. 7 is a flowchart illustrating the data extension process executed in S103. [Figure 8] FIG. 8 is a diagram illustrating a source image. [Figure 9] FIG. 9 is a diagram illustrating an example of a mask image. [Figure 10] FIG. 10 is a flowchart illustrating the process of generating learning image data executed in S210. [Figure 11A] FIG. 11A is an explanatory diagram illustrating a method for determining whether each pixel on the XY coordinate plane belongs to a synthesis area, a boundary area, or another area. [Figure 11B] FIG. 11B is an explanatory diagram illustrating a method for determining whether each pixel on the XY coordinate plane belongs to a synthesis area, a boundary area, or another area. [Figure 12] FIG. 12 is an explanatory diagram illustrating a Laplacian filter. [Figure 13] FIG. 13 is an explanatory diagram showing the Z coordinates of each pixel represented by the source image data as a to y. [Figure 14] FIG. 14 is an explanatory diagram illustrating a method for acquiring gradient image data. [Figure 15] FIG. 15 is a diagram illustrating the first 3D image. [Figure 16] FIG. 16 is a diagram illustrating a composite image corresponding to FIG. [Figure 17] FIG. 17 is a diagram illustrating a composite image generated by a conventional method (alpha blending) and a composite image shown by learning image data generated by the AI model generation device according to the first embodiment. [Figure 18] FIG. 18 is a view corresponding to FIG. 16 of the second embodiment. [Figure 19] FIG. 19 is a view equivalent to FIG. 7 of the third embodiment. [Figure 20] FIG. 20 is a diagram illustrating a composite image in which hole and sputter regions are designated by annotation labels. [Figure 21] FIG. 21 is a diagram illustrating an example of a mask image. [Figure 22] FIG. 22 is an explanatory diagram illustrating an annotation label indicating the defective weld U in FIG. [Figure 23] FIG. 23 is an explanatory diagram illustrating an annotation label indicating the defective weld L in FIG. [Figure 24] FIG. 24 is a diagram illustrating the first and second 3D images in the fifth embodiment. [Figure 25] FIG. 25 is a diagram corresponding to FIG. 24 in which the areas where welding defects were formed are painted black. [Figure 26] FIG. 26 is a diagram corresponding to FIG. 24 in which normal regions are painted black. [Figure 27] FIG. 27 is a diagram illustrating the first and second 3D images in the fifth embodiment. [Figure 28] FIG. 28 is a diagram corresponding to FIG. 27, in which an image of a normal region is superimposed on an area where one welding defect has been formed. [Figure 29] FIG. 29 is a diagram corresponding to FIG. 27, in which an image of a normal region is superimposed on an image of two regions where welding defects have been formed. [Figure 30] FIG. 30 is a view equivalent to FIG. 25 of the sixth embodiment. [Figure 31] FIG. 31 is a view corresponding to FIG. 26 of the sixth embodiment. [Figure 32] FIG. 32 is a diagram corresponding to FIG. 30 and showing a composite image. [Figure 33]FIG. 33 is a diagram illustrating a first 3D image in the seventh embodiment. [Figure 34] FIG. 34 is a diagram illustrating a second 3D image in the seventh embodiment. [Figure 35] FIG. 35 is a diagram corresponding to FIG. 33 and showing a composite image. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The following description of the preferred embodiments is merely exemplary in nature and is not intended to limit the present invention, its applications, or its uses.
[0012] (Embodiment 1) A first embodiment of the present invention will be described below with reference to the drawings.
[0013] 1 shows a welding system 1. This welding system 1 includes an AI model generation device 2 as a learning data generation device according to the first embodiment of the present disclosure, and an appearance inspection device 3 that performs welding and inspects the appearance of a weld bead WB (see FIG. 2) as a weld mark.
[0014] The AI model generation device 2 generates learning image data that constitutes learning data used in machine learning to create a defect detection model.
[0015] The defect detection model is an object detection model that predicts and outputs welding defect information based on input image data. The defect detection model is represented by a CNN (Convolutional Neural Network). The welding defect information includes category information indicating the type of welding defect in the 3D image represented by the image data, and position information indicating the position of the welding defect. In detail, the category information indicates whether the welding defect is a hole HL, a pit PT, spatter SP, undercut UC, or a protrusion (not shown), etc. A pit PT is an opening on the surface of a weld bead WB. The position information indicates the position of a bounding box containing the welding defect. The bounding box is a rectangular boundary area that surrounds an object such as a welding defect.
[0016] Furthermore, the AI model generation device 2 acquires annotation data of welding defect information, and generates a defect detection model by machine learning using a plurality of data sets of the generated learning image data and the annotation data as learning data.
[0017] Specifically, the AI model generation device 2 includes a storage device 21, a processor 22 as a processing unit, a first output device 23, and a first input device 24.
[0018] The storage device 21 has a first image data storage unit 211, a second image data storage unit 212, a first segmentation information storage unit 213, an annotation data storage unit 214, a learning image data storage unit 215, and a parameter storage unit 216.
[0019] The first image data storage unit 211 stores, as first image data, image data of a first 3D image (target image) acquired by capturing an image of a welding point using a 3D sensor 36 (described later). The first 3D image is an image captured of the area around the weld bead WB, as exemplified in FIG. 2, and does not include welding defects.
[0020] The second image data storage unit 212 stores, as second image data, image data of a second 3D image acquired by capturing an image of the welding point using the 3D sensor 36 (described later). The second 3D image is an image captured of the area around the weld bead WB, as exemplified in FIG. 3, and includes welding defects.
[0021] First segmentation information storage unit 213 stores first segmentation information that identifies the formation region of the weld bead WB on the XY coordinate plane of the first 3D image. The first segmentation information is identified by a user's input to first input device 24 while the first 3D image is being output to first output device 23. Figure 4 is a diagram equivalent to Figure 2 in which the formation region of the weld bead WB (the position of the weld bead WB) is shaded in black.
[0022] The annotation data storage unit 214 stores annotation data specified by a user's input to the first input device 24 while the second 3D image is being output to the first output device 23. This annotation data includes category information indicating the type of welding defect in the second 3D image and second segmentation information specifying the position (segmentation area) of an extracted image in the second 3D image. The extracted image is an image of an area in which the welding defect is captured. Figure 5 explains the welding defect captured in the second 3D image shown in Figure 3. In Figure 5, spatter SP, pit PT, and hole HL are specified as welding defects.
[0023] The learning image data storage unit 215 stores learning image data and its annotation data generated by the processor 22 (described later). The learning image data represents a 3D composite image obtained by combining a partial area of the XY coordinate plane of the first 3D image (a partial area as viewed from the Z-axis direction) with a partial extracted image of the XY coordinate plane of the second 3D image (an image of a partial area of the second 3D image as viewed from the Z-axis direction). The first 3D image and the composite image have the same image size (number of pixels in the horizontal direction and the vertical direction).
[0024] The parameter storage unit 216 stores parameters that specify a fault detection model generated by the processor 22, which will be described later.
[0025] The processor 22 has an image data receiving unit 221, a first segmentation information receiving unit 222, an annotation data receiving unit 223, a learning image data generating unit 224, and an AI model generating unit 225.
[0026] The image data receiving unit 221 receives image data of a 3D image acquired by photographing the welding point with the 3D sensor 36 described later from the appearance inspection device 3, and stores the image data as first image data in the first image data storage unit 211 or as second image data in the second image data storage unit 212. Whether the image data received from the appearance inspection device 3 is stored as the first image data in the first image data storage unit 211 or as the second image data in the second image data storage unit 212 is determined by, for example, a user's input to the first input device 24.
[0027] First segmentation information receiving unit 222 outputs a first 3D image based on the first image data stored in first image data storage unit 211 to first output device 23. In this state, first input device 24 receives an input signal corresponding to a user's input operation. Here, first input device 24 receives a user's input operation specifying a formation region of a weld bead WB in the image. Here, the user's input operation is, for example, an operation of clicking and tracing the outer periphery of the weld bead WB with a mouse. When the user's input operation ends, first segmentation information receiving unit 222 stores first segmentation information indicating the formation region of the weld bead WB specified based on the user's input operation in first segmentation information storage unit 213.
[0028] The annotation data receiving unit 223 outputs a second 3D image based on the second image data stored in the second image data storage unit 212 to the first output device 23. In this state, an input signal corresponding to a user's input operation is received from the first input device 24. Here, the first input device 24 receives a user's input operation specifying the type and position of a welding defect in the image. Here, the input operation specifying the position of the welding defect is, for example, an operation of clicking and tracing the outer periphery of the welding defect with a mouse. When the user's input operation is completed, the annotation data receiving unit 223 stores, in the annotation data storage unit 214, category information indicating the type of welding defect identified based on the user's input operation and second segmentation information specifying the area containing the welding defect as the position of an extracted image in the second 3D image.
[0029] The training image data generation unit 224 generates training image data using the first image data stored in the first image data storage unit 211 and the second image data stored in the second image data storage unit 212. The training image data generation unit 224 then stores the generated training image data in the training image data storage unit 215. Details of how the training image data generation unit 224 generates training image data will be described later.
[0030] The AI model generation unit 225 creates a defect detection model by performing machine learning using training data consisting of multiple sets of training image data stored in the training image data storage unit 215 and its annotation data. The annotation data of the training image data is information that identifies the type and position of a welding defect in a composite image represented by the training image data. The annotation data of the training image data can be acquired, for example, by receiving user input via the first input device 24 while the composite image represented by the training image data is output to the first output device 23. The method of acquiring the annotation data of the training image data is not limited to this. The AI model generation unit 225 identifies the weights and biases of each node constituting the CNN that represents the defect detection model as parameters that identify the defect detection model. The AI model generation unit 225 then evaluates the accuracy of the defect detection model and stores the identified parameters in the parameter storage unit 216.
[0031] In this way, the processor 22 generates learning image data representing a composite image obtained by combining the first and second 3D images as image data to be used for machine learning by the AI model generation unit 225.
[0032] The first output device 23 is configured by, for example, a liquid crystal monitor.
[0033] The first input device 24 receives an input operation from the user specifying the type and position of a welding defect in the image, and outputs an input signal corresponding to the received input operation.
[0034] Visual inspection device 3 has a welding torch 31, a wire feeder (not shown), a welding power source 32, an output control unit 33, a robot arm 34, a robot control unit 35, a 3D sensor 36, a computer 37, a second output device 38, and a second input device 39. When power is supplied from welding power source 32 to a welding wire WI held by welding torch 31, an arc is generated between the tip of the welding wire WI and a workpiece (base metal) W, and the workpiece W is heated to perform arc welding. Note that visual inspection device 3 has other components and equipment such as piping and gas cylinders for supplying shielding gas to welding torch 31, but for convenience of explanation, these are not shown and will not be described.
[0035] Output control unit 33 is connected to welding power source 32 and a wire feeder (not shown) and controls the welding output of welding torch 31, in other words, the power supplied to welding wire WI and the power supply time, in accordance with predetermined welding conditions. Output control unit 33 also controls the feed speed and feed amount of welding wire WI fed from a wire feeder (not shown) to welding torch 31. Note that the welding conditions may be input directly to output control unit 33 via an input unit (not shown), or may be selected from a welding program separately read out from a recording medium or the like.
[0036] Robot arm 34 is a known articulated robot that holds welding torch 31 at its tip and is connected to robot control unit 35. Robot control unit 35 controls the operation of robot arm 34 so that the tip of welding torch 31, in other words, the tip of welding wire WI held by welding torch 31, moves to a desired position while tracing a predetermined welding trajectory.
[0037] The 3D sensor 36 is attached to the welding torch 31 and measures the shape of the welded portion PW of the workpiece W. The 3D sensor 36 is a three-dimensional shape measurement sensor that includes, for example, a laser light source (not shown) configured to scan the surface of the workpiece W and a camera (not shown) that captures the reflection trajectory of the laser light projected onto the surface of the workpiece W (hereinafter, sometimes referred to as a shape line). The 3D sensor 36 scans the entire welded portion PW of the workpiece W with a laser beam, and the camera captures the laser beam reflected by the welded portion PW, thereby measuring the shape of the welded portion PW. The 3D sensor 36 is configured to measure the shape of not only the welded portion PW but also a predetermined area around it. This is to evaluate the presence or absence of spatter SP, etc. The camera has a CCD or CMOS image sensor as an imaging element. The configuration of the 3D sensor 36 is not limited to the above, and other configurations are possible. For example, an optical interferometer may be used instead of the camera.
[0038] The computer 37 executes software implemented on a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) to realize the functions of multiple functional blocks within the computer 37. The computer 37 has an image processing unit 371, a welding defect information calculation unit 372, and a welding defect display image generation unit 373.
[0039] The image processing unit 371 receives the shape data acquired by the 3D sensor 36 and converts it into image data of a 3D image including the welding point PW. For example, the image processing unit 371 acquires point cloud data of the shape lines captured by the 3D sensor 36. The image processing unit 371 also corrects the inclination and distortion of the base portion of the welding point PW relative to a predetermined reference plane, for example, the installation surface of the workpiece W, by statistically processing the point cloud data, and acquires image data including the welding point PW.
[0040] The defective welding information calculation unit 372 calculates (predicts) defective welding information based on the image data generated by the image processing unit 371, using the defect detection model generated by the AI model generation device 2. The calculation of the defective welding information may be performed when a predetermined input operation is performed on the second input device 39.
[0041] The defective welding display image generating unit 373 calculates the pixel value of each pixel of the defective welding display image displaying the defective welding information, based on the defective welding information calculated by the defective welding information calculating unit 372. The information specifying the defective welding display image may be compressed using a JPEG format or the like.
[0042] The second output device 38 displays the poor welding display image based on the pixel value of each pixel calculated by the poor welding display image generating unit 373. The second output device 38 is configured by a liquid crystal monitor or the like.
[0043] The second input device 39 accepts a predetermined input operation by the user to cause the computer 37 to start calculating the welding defect information.
[0044] Next, the operation of the welding system 1 configured as described above will be described with reference to the flowchart of FIG.
[0045] First, in S101, the 3D sensor 36 captures an image of a non-defective bead WB, which is a weld bead WB that does not include a welding defect, and a defective bead WB, which is a weld bead WB that includes a welding defect. The 3D sensor 36 captures an image of the surface of the workpiece W (the object to be captured) around the weld bead WB. The learning image data storage unit 215 pre-stores multiple learning image data and corresponding annotation data. The image processing unit 371 receives the shape data acquired by the 3D sensor 36 and acquires image data of the 3D image of the non-defective bead and image data of the 3D image of the defective bead. The image data receiving unit 221 of the AI model generation device 2 then receives the image data of the 3D image of the non-defective bead from the appearance inspection device 3 and stores the image data in the first image data storage unit 211 as first image data. The image data receiving unit 221 then receives the image data of the 3D image of the defective bead from the appearance inspection device 3 and stores the image data in the second image data storage unit 212 as second image data.
[0046] Next, in S102, first segmentation information receiving unit 222 outputs a first 3D image based on the first image data stored in first image data storage unit 211 to first output device 23. In this state, first segmentation information receiving unit 222 accepts an input signal corresponding to a user's input to first input device 24. Then, first segmentation information receiving unit 222 stores, in first segmentation information storage unit 213, first segmentation information indicating the formation region of weld bead WB identified based on the user's input operation.
[0047] Furthermore, the annotation data receiving unit 223 outputs a second 3D image based on the second image data stored in the second image data storage unit 212 to the first output device 23. In this state, the annotation data receiving unit 223 accepts an input signal corresponding to a user's input to the first input device 24. Then, the annotation data receiving unit 223 stores, in the annotation data storage unit 214, the category information and the second segmentation information identified based on the user's input operation.
[0048] Next, in S103, the learning image data generation unit 224 performs data extension processing. Specifically, the learning image data generation unit 224 generates learning image data using the first image data stored in the first image data storage unit 211 and the second image data stored in the second image data storage unit 212, and adds (stores) the learning image data to the learning image data storage unit 215.
[0049] Next, in S104, the AI model generation unit 225 generates a fault detection model by performing machine learning using the training image data and its annotation data stored in the training image data storage unit 215. The annotation data of the training image data can be acquired, for example, by outputting a composite image represented by the training image data to the first output device 23 and accepting user input via the first input device 24 in this state, but the acquisition method is not limited to this.
[0050] Next, in S105, the AI model generation unit 225 evaluates the accuracy of the fault detection model generated in S104.
[0051] Next, in S106, the AI model generation unit 225 stores parameters for specifying the defect detection model generated in S104 in the parameter storage unit 216. The parameters stored in the parameter storage unit 216 are sent to the defective welding information calculation unit 372 of the visual inspection device 3. The defective welding information calculation unit 372 stores the parameters.
[0052] Next, in S107, when the 3D sensor 36 captures an image of the welded portion PW, the image processing unit 371 receives the shape data acquired by the 3D sensor 36 and converts it into image data of a welding image including the welded portion PW. The defective welding information calculation unit 372 performs an appearance inspection to calculate (predict) defective welding information based on the image data generated by the image processing unit 371, using the defect detection model generated by the AI model generation device 2. The defective welding display image generation unit 373 calculates a pixel value of each pixel of the defective welding display image that displays the defective welding information, based on the defective welding information calculated by the defective welding information calculation unit 372. Then, the second output device 38 displays the defective welding display image based on the pixel value of each pixel calculated by the defective welding display image generation unit 373.
[0053] Next, the procedure in S103 in which the learning image data generating unit 224 generates learning image data will be described with reference to the flowchart in FIG.
[0054] First, in S201, the learning image data generation unit 224 reads the first image data from the first image data storage unit 211 and also reads the first segmentation information from the first segmentation information storage unit 213.
[0055] Next, in S202, the learning image data generation unit 224 selects at least one selected coordinate from a plurality of coordinates on the XY coordinate plane included in the formation region of the weld bead WB indicated by the first segmentation information.
[0056] The learning image data generation unit 224 executes the operations of S203 to S207 in parallel with the processes of S201 and S202.
[0057] In S203, the learning image data generation unit 224 reads the second image data from the second image data storage unit 212 and also reads the second segmentation information from the annotation data storage unit 214.
[0058] Next, in S204, the learning image data generation unit 224 extracts extracted image data indicating the extracted image based on the second image data and second segmentation information read in S203.
[0059] Next, in S205, the learning image data generation unit 224 calculates the XY coordinates (points) of the circumscribing rectangle (a quadrangle that contacts the outer edge of the extracted image) of the extracted image on the XY coordinate plane.
[0060] Next, in S206, the learning image data generation unit 224 calculates the center coordinates of the circumscribing rectangle identified in S205 for the extracted image.
[0061] Next, in S207, the learning image data generation unit 224 extracts rectangular area image data from the second image data that indicates the image within the circumscribing rectangle identified in S205 for the extracted image (the surface shape around the weld bead WB within the circumscribing rectangle).
[0062] When the processes of S201 to S207 are completed, in S208, the learning image data generation unit 224 generates source image data indicating a source image SI (see FIG. 8) in which the image within the circumscribing rectangle is pasted onto the first 3D image so that the center coordinates calculated in S206 correspond to the selected coordinates selected in S202. This source image data is generated using the first image data, the selected coordinates selected in S202, the center coordinates calculated in S206, and the rectangular area image data extracted in S207. FIG. 8 illustrates an example of the source image SI. In FIG. 8, the pasted image (the image within the circumscribing rectangle) is indicated by the symbol RE. The image sizes of the source image SI, the first 3D image, and the composite image are the same.
[0063] Furthermore, the learning image data generation unit 224 executes the operation of S209 in parallel with the processing of S204 to S208.
[0064] In S209, the learning image data generation unit 224 performs a predetermined conversion process on the second image data based on the second image data and the second segmentation information to generate mask image data that sets pixel values of each pixel included in the extracted image to values different from pixel values of each pixel in the second 3D image that are not included in the extracted image. Specifically, the conversion process converts, for example, the Z coordinate of each point included in the extracted image to 1 and the Z coordinate of each point in the second 3D image that is not included in the extracted image to 0. Figure 9 illustrates an example of a mask image MI. In Figure 9, the region of the extracted image (poor welding) is indicated by the symbol EX.
[0065] Next, in S210, the learning image data generation unit 224 generates learning image data representing a composite image using the first image data, the source image data generated in S208, and the mask image data generated in S209.
[0066] Next, in S211, the learning image data generation unit 224 stores the learning image data generated in S210 in the learning image data storage unit 215.
[0067] Here, the details of the process executed by the learning image data generation unit 224 in S210 will be described with reference to the flowchart of FIG.
[0068] In S210, first, in S301, the learning image data generation unit 224 determines whether each pixel on the XY coordinate plane in the mask image MI belongs to a synthesis region, a boundary region, or other region, based on the mask image data.
[0069] Specifically, as shown in FIG. 11A, the training image data generation unit 224 first regards pixels in the mask image MI whose Z coordinate is 1 as part of the synthesis region, and regards pixels whose Z coordinate is 0 as outside the synthesis region. Next, it determines whether the four pixels adjacent to each pixel in the synthesis region on the XY coordinate plane, above, below, left, and right, are 0 or 1. For example, in FIG. 11A, it determines whether the pixels P2 to P5 adjacent to the pixel indicated by P1 on the XY coordinate plane, above, below, left, and right, are 0 or 1. Then, the training image data generation unit 224 determines that, among the four pixels adjacent to each pixel in the synthesis region on the XY coordinate plane, those whose Z coordinate is 0 belong to the boundary region. Furthermore, the training image data generation unit 224 determines that, among the four pixels adjacent to each pixel in the synthesis region on the XY coordinate plane, those whose Z coordinate is 1 belong to the synthesis region. In the example of FIG. 11A, as shown in FIG. 11B, the pixels with the codes P2 and P3 are determined to belong to the boundary region, and the pixels with the codes P4 and P5 are determined to belong to the composite region.
[0070] Next, in S302, the learning image data generation unit 224 generates a Laplacian matrix A (shown in equation (4) described later) which is a sparse matrix.
[0071] Next, in S303, the learning image data generation unit 224 sets i=0.
[0072] Next, in S304, the learning image data generation unit 224 sets i=i+1.
[0073] Next, in S305, the training image data generation unit 224 determines whether the i-th pixel of the n pixels belonging to the combination area and the boundary area belongs to the combination area or the boundary area. If the i-th pixel belongs to the combination area, the training image data generation unit 224 proceeds to S306, whereas if the i-th pixel belongs to the boundary area, the training image data generation unit 224 proceeds to S307.
[0074] Next, in S306, the learning image data generation unit 224 performs filtering using the Laplacian filter shown in FIG. 12 on the Z coordinate of the i-th pixel in the synthesis area represented by the source image data, and sets the value obtained by this filtering using the Laplacian filter as the Z coordinate (pixel value) of the corresponding pixel (i-th pixel) in the gradient image data. Therefore, the gradient image data sets the pixel value of the corresponding pixel to be a value obtained by performing filtering using the Laplacian filter on the Z coordinate of each point of the object included in the extracted image. In FIG. 13, the i-th pixel is pixel TP, and a to y indicate the Z coordinates of the pixels surrounding pixel TP in the source image. Then, the Z coordinate of pixel TP in the gradient image data is expressed by the following equation. Here, the Z coordinate of pixel TP is defined as z(i). The learning image data generation unit 224 then proceeds to S308.
[0075] z(i)=0*g+1*h+0*i+1*l+(―4)*m+1*n+0*q+1*r+0*s=―4m+h+l+n+r In addition, in S307, the learning image data generation unit 224 sets the Z coordinate of the i-th pixel indicated by the first image data as the Z coordinate of the i-th pixel in the gradient image data.
[0076] Next, in S308, the learning image data generation unit 224 determines whether i = n. If i = n is not true, the process returns to S304. On the other hand, if i = n, the process proceeds to S309.
[0077] In this way, the mask image data is used to obtain the gradient image data.
[0078] FIG. 14 shows a method for acquiring gradient image data by performing a Laplacian filter process on the Z coordinate of each pixel of the source image SI.
[0079] In this example, gradient image data representing the lower right gradient image GI is obtained by performing Laplacian filtering on the Z coordinates of pixels TP1 to TP3 using the Z coordinates of the area surrounded by a thick rectangle. TP1 is the kth pixel, TP2 is the k+1th pixel, and TP3 is the k+2th pixel. In the lower right gradient image GI, the pixel value of each pixel in the area surrounded by a thick rectangle indicates the amount of change in height.
[0080] In S309, the learning image data generation unit 224 specifies a composite image I(x,y) so that the following formulas (1) and (2) are satisfied. In the following formulas (1) and (2), S(x,y) is the source image SI indicated by the source image data, and T(x,y) is the first 3D image indicated by the first image data. I(x,y), S(x,y), and T(x,y) are functions with the x-coordinate value and the y-coordinate value as variables, and indicate the z-coordinate. Furthermore, Ωin indicates the composite area, and ∂Ω indicates the boundary area. Satisfying the following formulas (1) and (2) corresponds to the composite condition.
[0081]
number
[0082] Here, Δ is the Laplacian operator, which represents the second derivative.
[0083]
number
[0084] Specifically, the pixel values of the composite image are calculated so as to satisfy the following equation (3) using the Laplacian matrix A and a matrix B whose components are the Z coordinates of the 1st to nth pixels in the gradient image data: In equation (3), the matrix x has as its components the Z coordinates of the 1st to nth pixels belonging to the composite region or boundary region in the composite image represented by the training image data.
[0085]
number
[0086] If equation (3) is rewritten more specifically, it becomes as shown in equation (4) below.
[0087]
number
[0088] By calculating the pixel values of the composite image in this manner, the gradient of the photographed object (the surface of the weld bead WB and the workpiece W) at each point included in the composite area of the abstract image in the composite image becomes equal to the gradient of the corresponding photographed object in the extracted image. Here, the gradient is a value obtained by performing a Laplacian filter process on the Z coordinate of each point (each pixel) of the photographed object. Furthermore, the Z coordinate of each point of the photographed object in the boundary area adjacent to the composite area in the composite image on the XY coordinate plane becomes equal to the Z coordinate of the corresponding point of the photographed object in the first 3D image.
[0089] The image shown in Fig. 16 is an example of a composite image when the first 3D image is the image shown in Fig. 15. In the first embodiment, in S202, selected coordinates are selected from a plurality of coordinates included within the area of the weld bead WB. Therefore, in Fig. 16, the center coordinates of the circumscribed rectangles of the hole HL and the pit PT are located only within the area of the weld bead WB, and the center coordinates of the weld defect are not located outside the area of the weld bead WB.
[0090] Therefore, according to this embodiment 1, the gradient of the object to be photographed included in the synthesis area in the composite image is made equal to the gradient of the corresponding object to be photographed in the second 3D image, so that the shape of the object to be photographed in the second 3D image can be reflected in the synthesis area of the composite image.
[0091] Furthermore, since the Z coordinates of each point of the photographed object in the boundary area adjacent to the composite area in the XY coordinate plane in the composite image are set equal to the Z coordinates of the corresponding points in the first 3D image, a step as an unintended feature is not formed at the boundary between the first and second 3D images in the composite image, thereby suppressing deterioration in the prediction accuracy of the defect detection model due to the occurrence of an unintended feature at the boundary between the first and second 3D images in the composite image.
[0092] The image on the left in Fig. 17 is an example of a composite image generated by a conventional method (α blending) as a comparative example. The image on the right in Fig. 17 is an example of a composite image displayed by learning image data generated by the AI model generation device 2 according to the present embodiment 1. Fig. 17 shows an example in which the welding defect in the second 3D image is a hole HL.
[0093] 17, in the composite image generated by the conventional method, a step St is formed at the boundary between the first and second 3D images. In contrast, in the composite image according to the first embodiment, the step St is not formed as an unintended feature at the boundary between the first and second 3D images in the composite image.
[0094] (Embodiment 2) Figure 18 is a diagram equivalent to Figure 17 of the second embodiment. In this second embodiment, in S202, learning image data generation unit 224 selects selected coordinates from a plurality of coordinates included within the region of weld bead WB, and further selects selected coordinates from a plurality of coordinates outside the region of weld bead WB. In the example of Figure 18, spatters SP are formed both within the region of weld bead WB and outside the region of weld bead WB in the composite image.
[0095] The rest of the configuration and operation of the welding system 1 is the same as in the first embodiment, so detailed description thereof will be omitted.
[0096] (Embodiment 3) 19 is a diagram equivalent to FIG. 7 of the third embodiment. In the third embodiment, after S204, in S212, the learning image data generation unit 224 deforms the extracted image data. Also in S212, the learning image data generation unit 224 deforms the second segmentation information according to the shape of the extracted image represented by the deformed extracted image data.
[0097] After executing S212, the learning image data generation unit 224 then generates mask image data in S208 based on the second segmentation information after the transformation.
[0098] The rest of the configuration and operation of the welding system 1 is the same as in the first embodiment, so detailed description thereof will be omitted.
[0099] According to the third embodiment, learning image data representing a composite image obtained by transforming an extracted image and combining it with the first 3D image can be generated so as to satisfy the composition conditions.
[0100] (Embodiment 4) In embodiment 4, in S205, the learning image data generation unit 224 stores the position of the annotation label as annotation data in the learning image data memory unit 215 together with the type of welding defect contained in the extracted image based on the XY coordinates (points) of the circumscribing rectangle of the extracted image.
[0101] 20 shows annotation labels indicated by annotation data for each type of welding defect included in the circumscribing rectangle. In FIG. 20, hole labels are indicated by the symbol HLL, and spatter labels are indicated by the symbol SPL.
[0102] Next, a method for identifying annotation labels will be described. Here, it is assumed that the mask image MI is an image of M pixels (X direction) * N pixels (Y direction) shown in FIG. 21 and includes a defective weld (extracted image) U and a defective weld (extracted image) L. In this case, as shown in FIG. 22, lines obtained by shifting both sides of the circumscribing rectangle of the defective weld U outward by ΔM in the X direction are defined as both sides of the annotation label corresponding to the defective weld U in the X direction. Furthermore, lines obtained by shifting both sides of the circumscribing rectangle of the defective weld U outward by ΔN are defined as both sides of the annotation label corresponding to the defective weld U in the Y direction. Similarly, as shown in FIG. 23, lines obtained by shifting both sides of the circumscribing rectangle of the defective weld L outward by ΔM in the X direction are defined as both sides of the annotation label corresponding to the defective weld L in the X direction. Furthermore, lines obtained by shifting both sides of the circumscribing rectangle of the defective weld L outward by ΔN are defined as both sides of the annotation label corresponding to the defective weld L in the Y direction. ΔM and ΔN are set to predetermined integers between 0 and 100 so that the welding defect does not protrude from the composite image.
[0103] The rest of the configuration and operation of the welding system 1 is the same as in the first embodiment, so detailed description thereof will be omitted.
[0104] (Embodiment 5) FIG. 24 illustrates first and second 3D images in the fifth embodiment. In the fifth embodiment, the first and second 3D images are common. That is, in the fifth embodiment, the first 3D image represented by the first image data includes the welding defect IW. Furthermore, the first segmentation information identifies the region where the welding defect IW occurs. FIG. 25 is a diagram equivalent to FIG. 24 in which the region where the welding defect IW occurs is shaded in black. Furthermore, the annotation data does not include category information. Furthermore, the extracted image is an image of a normal region PR in the weld bead WB that has no welding defect. In FIG. 26, the extracted image (segmentation region), i.e., the normal region PR, is shaded in black.
[0105] The images shown in Figures 28 and 29 are examples of composite images when the first and second 3D images are the images shown in Figure 27. In Figure 28, an extracted image of the normal region PR is not composited into the formation area of the welding defect IW1, but an extracted image of the normal region PR is composited into the formation area of the welding defect IW2. In Figure 29, an extracted image of the normal region PR is composited into the formation areas of the welding defect IW1 and the welding defect IW2.
[0106] The rest of the configuration and operation of the welding system 1 is the same as in the first embodiment, so detailed description thereof will be omitted.
[0107] (Embodiment 6) Fig. 30 is a diagram corresponding to Fig. 25 of the sixth embodiment, and Fig. 31 is a diagram corresponding to Fig. 26 of the sixth embodiment. In the sixth embodiment, the first segmentation information identifies the region where a welding defect IW occurs around the root RO (the line where two surfaces to be welded intersect) of the lap joint. The extracted image is an image of a normal region PR that includes the root RO and has no welding defect.
[0108] The image shown in FIG. 32 is a composite image obtained by combining the image of the normal region PR in FIG. 26 with the region where the defective weld IW in FIG. 30 has been formed.
[0109] The rest of the configuration and operation of the welding system 1 is the same as in the fifth embodiment, so detailed description thereof will be omitted.
[0110] (Embodiment 7) Fig. 33 illustrates a first 3D image in embodiment 7. Fig. 34 illustrates a second 3D image in embodiment 7. In embodiment 7, the first segmentation information identifies the formation area of spatter SP1, which is a welding defect. The second segmentation information identifies the formation area of spatter SP2, which is a welding defect. In other words, the extracted image is an image of spatter SP2.
[0111] Fig. 35 illustrates a composite image when the first 3D image is Fig. 33 and the second 3D image is Fig. 34. In this example, an extracted image including sputter SP2 is composited with the first 3D image so that sputter SP2 overlaps a portion of sputter SP1.
[0112] The rest of the configuration and operation of the welding system 1 is the same as in the fifth embodiment, so detailed description thereof will be omitted.
[0113] The functions of the processor 22 in the first to seventh embodiments may be realized by a plurality of processors.
[0114] Furthermore, in the above embodiments 1 to 7, the present invention was applied to generate a defect detection model that predicts welding defect information based on image data, but the present invention can also be applied to detecting objects other than welding defects or generating other models that make predetermined judgments about images. [Industrial Applicability]
[0115] The training data generation device and training data generation method disclosed herein can suppress deterioration in the prediction accuracy of a model, and are useful as a training data generation device and training data generation method for acquiring image data to be used in machine learning. [Explanation of symbols]
[0116] 2. AI model generation device (learning data acquisition device) 22 Processor (processing unit) 36 3D sensors 21 Storage device (storage unit) EX Extracted Image WB Weld bead (weld mark) H hole (poor welding) PT pit (poor welding) SP, SP1, SP2 spatter (poor welding) UC Undercut (defective welding) PR normal area IW, IW1, IW2 Welding defects
Claims
1. A learning data generation device that generates learning image data to be used in machine learning for creating a model by using first image data representing a first 3D image and second image data representing a second 3D image to synthesize at least a portion of an extracted image in an XY coordinate plane of the second 3D image with a portion of an area in an XY coordinate plane of the first 3D image, a storage unit configured to store the first image data, the second image data, and segmentation information that identifies a position of the extracted image in the second 3D image; and a processing unit that generates the training image data based on the first image data, the second image data, and the segmentation information stored in the memory unit so as to satisfy the following synthesis conditions: a gradient of the object included in a synthesis area of the extracted image in the synthetic image shown by the training image data is equal to a gradient of the corresponding object in the extracted image; and a Z coordinate of each point of the object in a boundary area adjacent to the synthesis area in the synthetic image on the XY coordinate plane is equal to a Z coordinate of a corresponding point in the first 3D image.
2. 2. The training data generation device according to claim 1, The learning data generation device is characterized in that the first 3D image is an image of the area around a weld mark, and the extracted image includes a welding defect.
3. 2. The training data generation device according to claim 1, The learning data generating device is characterized in that the gradient is a value obtained by performing a Laplacian filter process on the Z coordinate of each point of the object to be photographed.
4. 4. The training data generation device according to claim 3, The processing unit A learning data generation device characterized by acquiring gradient image data in which the pixel value of each point is the value obtained by performing Laplacian filter processing on the Z coordinate of each point of the object to be photographed included in the extracted image based on the second image data and the segmentation information.
5. 5. The training data generation device according to claim 4, The processing unit Based on the second image data and the segmentation information, mask image data is generated that sets a pixel value of each pixel included in the extracted image to a value different from a pixel value of each point in the second 3D image that is not included in the extracted image; The learning data generating device is characterized in that the mask image data is used to obtain the gradient image data.
6. 2. The training data generation device according to claim 1, The processing unit A training data generation device characterized by being able to generate training image data representing a composite image obtained by transforming the extracted image and combining it with the first 3D image so as to satisfy the synthesis conditions.
7. A learning data generation method for generating learning image data to be used in machine learning for creating a model by using first image data representing a first 3D image and second image data representing a second 3D image to synthesize at least a portion of an extracted image in an XY coordinate plane of the second 3D image with a partial region in an XY coordinate plane of the first 3D image, the method comprising: a training data generation method for generating the training image data based on the first image data, the second image data, and segmentation information that identifies the position of the extracted image in the second 3D image, such that the gradient of the object included in the synthesis area of the extracted image in the synthetic image represented by the training image data is equal to the gradient of the corresponding object in the extracted image, and the Z coordinate of each point of the object in the boundary area adjacent to the synthesis area in the synthetic image on the XY coordinate plane is equal to the Z coordinate of the corresponding point in the first 3D image.
Citation Information
Patent Citations
Data generation system, data generation method, data generation device, and additional learning requirement assessment device
WO2021193347A1