Learning model generation method, image processing method, image processing system, and welding system
By using feature position change information from multiple input images in the learning model, the feature extraction process is optimized, solving the problem of insufficient feature extraction accuracy in existing technologies and achieving higher precision welding feature extraction and welding effects.
Patent Information
- Application Number
- CN202110725305.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-02
- Filing Date
- 2021-06-29
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2041-06-29
AI Technical Summary
In existing technologies, the learning models have low feature extraction accuracy and cannot effectively utilize feature position change information in multiple input images, resulting in insufficient welding accuracy and strength.
A learning model generation method is adopted, which learns from multiple input images in the teacher data to ensure that the change in feature position is less than the convolution kernel size. The feature extraction process is optimized by combining the training of the generator and the recognizer, and the U-NET structure is used for feature map generation and recognition.
This improved the accuracy of feature extraction and welding position, enhanced the strength and precision of welding, and improved the overall performance of the welding system.
Smart Images

Figure CN114387491B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment relates to a learning model generation method, an image processing method, an image processing system, and a welding system. BACKGROUND
[0002] Conventionally, a technique of estimating an extracted image of a feature from an input image using a learning model in which learning is completed is known.
[0003] PRIOR ART DOCUMENT
[0004] Patent Document 1: Japanese Patent Application Publication No. 2019-141902 SUMMARY
[0005] The embodiment aims to provide a learning model generation method, an image processing method, an image processing system, and a welding system that have high feature extraction accuracy.
[0006] The learning model generation method of the embodiment includes a process of obtaining teacher data including a plurality of learning input images and a learning extracted image of a feature extracted from one of the plurality of learning input images, and a process of causing a learning model to learn using the teacher data, the learning model outputting an extracted image of the feature estimated from a plurality of input images. The learning model includes an input layer that performs convolution. The positions of the feature in the respective plurality of learning input images are different from each other. The amount of change in the positions of the feature in the plurality of learning input images is smaller than the size of a convolution kernel of a filter of the input layer.
[0007] The image processing method of the embodiment includes a process of obtaining a plurality of input images, and a process of outputting an extracted image of a feature estimated from the plurality of input images using a learning completed model including an input layer that performs convolution and learned using teacher data including a plurality of learning input images and a learning extracted image of the feature extracted from one of the plurality of learning input images, the positions of the feature in the respective plurality of learning input images being different from each other, and the amount of change in the positions of the feature in the plurality of learning input images being smaller than the size of a convolution kernel of a filter of the input layer.
[0008] The image processing system of the embodiment includes an image processing section that outputs an extracted image of a feature estimated from a plurality of input images using a learned model, the learned model including an input layer that performs convolution and being completed using teacher data including a plurality of input images for learning and a feature extracted image for learning obtained by extracting the feature from one of the plurality of input images for learning, the positions of the feature in the plurality of input images for learning being different from each other, and the amount of change in the positions of the feature in the plurality of input images for learning being smaller than the size of a convolution kernel of a filter of the input layer.
[0009] The welding system of the embodiment includes a welding section that welds a welded member, one or more imaging devices that image a welding position of the welded member, an image processing section that outputs an extracted image of a feature of welding estimated from a plurality of images imaged by the imaging devices using a learned model, and a control section that controls the welding section based on the extracted image of the feature output by the image processing section, the learned model including an input layer that performs convolution and being completed using teacher data including a plurality of input images for learning and a feature extracted image for learning obtained by extracting the feature from one of the plurality of input images for learning, the positions of the feature in the plurality of input images for learning being different from each other, and the amount of change in the positions of the feature in the plurality of input images for learning being smaller than the size of a convolution kernel of a filter of the input layer. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 FIG. 1 is a diagram showing a welding system of a first embodiment.
[0011] Figure 2 (a) of FIG. 2 is a plan view showing a welded member before welding, Figure 2 (b) of FIG. 2 is a plan view showing the welded member during welding.
[0012] Figure 3 FIG. 3 is a block diagram showing a hardware structure of a control device in the welding system of the first embodiment.
[0013] Figure 4 FIG. 4 is a diagram showing a learned model of the first embodiment.
[0014] Figure 5 FIG. 5 is a flowchart showing a method of generating the learned model of the first embodiment.
[0015] Figure 6 FIG. 6 is a diagram showing data used in learning of the learned model of the first embodiment.
[0016] Figure 7 is a diagram illustrating a pre-processing method of an input image for learning in the generation method of the learning model of the first embodiment.
[0017] Figure 8 is a diagram illustrating a generator of the learning model of the first embodiment.
[0018] Figure 9 (a) of FIG. 1 is a diagram illustrating processing of an input layer in the learning model of the first embodiment, Figure 9 (b) of FIG. 1 is a diagram illustrating a convolution method in the input layer.
[0019] Figure 10 is a diagram illustrating that positions of outlines of the molten pool differ from each other in a plurality of input images for learning.
[0020] Figure 11 (a) of FIG. 2 is a diagram illustrating processing of a first intermediate layer in the learning model of the first embodiment, Figure 11 (b) of FIG. 2 is a diagram illustrating processing of a second intermediate layer in the learning model of the first embodiment, Figure 11 (c) of FIG. 2 is a diagram illustrating processing of a third intermediate layer in the learning model of the first embodiment.
[0021] Figure 12 (a) of FIG. 3 is a diagram illustrating processing of a fourth intermediate layer in the learning model of the first embodiment, Figure 12 (b) of FIG. 3 is a diagram illustrating processing of a fifth intermediate layer in the learning model of the first embodiment, Figure 12 (c) of FIG. 3 is a diagram illustrating processing of a sixth intermediate layer in the learning model of the first embodiment.
[0022] Figure 13 is a diagram illustrating processing of an output layer in the learning model of the first embodiment.
[0023] Figure 14 is a diagram illustrating a feature extraction image output by the learning model of the first embodiment.
[0024] Figure 15 is a flowchart illustrating a welding method using the learning model of the first embodiment.
[0025] Figure 16 is a diagram illustrating a part of a welding system of the second embodiment.
[0026] Figure 17 is a diagram illustrating a part of a welding system of the third embodiment.
[0027] Figure 18 is a diagram illustrating a part of a welding system of the fourth embodiment. DETAILED DESCRIPTION
[0028] <First Embodiment>
[0029] First, the first embodiment will be described.
[0030] Figure 1 is a view showing a welding system of the present embodiment.
[0031] Figure 2 (a) of FIG. 1 is a plan view showing the welded members before welding, Figure 2 (b) is a plan view showing the welded members during welding.
[0032] (Welding system)
[0033] The welding system 10 according to the present embodiment welds and integrates two or more welded members. The welding system 10 performs, for example, laser welding or arc welding. Here, an example in which the welding system 10 performs laser welding of two welded members 21, 22 as shown in (a) of FIG. 1 and (b) of FIG. 1 will be described. Hereinafter, the two welded members 21, 22 will also be referred to as a "first welded member 21" and a "second welded member 22". Figure 2 Figure 2
[0034] The first welded member 21 and the second welded member 22 are, for example, plate-shaped members. The first welded member 21 and the second welded member 22 are arranged in a manner facing each other. Hereinafter, a surface of the first welded member 21 facing the second welded member 22 will be referred to as a "first surface 21a", and a surface of the second welded member 22 facing the first welded member 21 will be referred to as a "second surface 22a".
[0035] As shown in FIG. 1, the welding system 10 includes, for example, a welding portion 11, a photographing device 15, an illuminating device 16, and a control device 17. Figure 1
[0036] Hereinafter, for easier understanding of the description, an XYZ orthogonal coordinate system will be used. Here, a direction from the first welded member 21 and the second welded member 22 toward the welding head 13 will be set as a "Z direction". In addition, a direction orthogonal to the Z direction, that is, a direction from the first welded member 21 toward the second welded member 22 will be set as a "Y direction". Furthermore, a direction orthogonal to the Z direction and the Y direction, that is, a direction of travel of the welding head 13 will be set as an "X direction".
[0037] The welding portion 11 includes a light source 12, a welding head 13, and an arm 14. The welding head 13 is connected to the light source 12, and irradiates the laser light L emitted from the light source 12 to the first and second welding members 21 and 22. The arm 14 holds the welding head 13, and moves the welding head 13 with respect to the first and second welding members 21 and 22. The arm 14 can move the welding head 13 in, for example, the X direction, the Y direction, and the Z direction.
[0038] The imaging device 15 is a camera including, for example, a CCD image sensor or a CMOS image sensor. The imaging device 15 is disposed above the first and second welding members 21 and 22. In the present embodiment, the imaging device 15 images a dynamic image D of the welding portion during welding. Hereinafter, the dynamic image D will also be referred to as a "control dynamic image D".
[0039] The illuminating device 16 illuminates the welding portion to obtain a clearer image by the imaging device 15. The illuminating device 16 can not be provided if an image that can be used for image processing of the image processing system described below can be obtained even if the welding portion is not illuminated.
[0040] Figure 3 is a block diagram showing a hardware structure of the control device in the welding system according to the present embodiment.
[0041] In the present embodiment, the control device 17 is a computer including a GPU (Graphics Processing Unit) 17a, a ROM (Read Only Memory) 17b, a RAM (Random Access Memory) 17c, a hard disk 17d, and the like. The GPU 17a, the ROM 17b, the RAM 17c, and the hard disk 17d are connected to each other through a bus 17e. However, the structure of the control device is not limited to the above-described structure. For example, the control device can use a CPU or another processor instead of the GPU. In addition, the control device can include another structure such as an input / output interface.
[0042] In the present embodiment, as shown in Figure 1 , the control device 17 functions as an acquisition section 171, an image processing section 172, a control section 173, and a storage section 174. The functions of the acquisition section 171, the image processing section 172, and the control section 173 are implemented by, for example, the GPU 17a. In addition, the function of the storage section 174 is implemented by, for example, the ROM 17b, the RAM 17c, the hard disk 17d, or the like.
[0043] When welding the first welded component 21 and the second welded component 22, the control unit 173 controls the welding unit 11 to emit a laser L from the welding head 13 towards the first welded component 21 and the second welded component 22, while moving the welding head 13 in the X direction. Additionally, the control unit 173 controls the imaging device 15 to capture a dynamic image D of the welded area during the welding process.
[0044] By irradiating the first welded component 21 and the second welded component 22 with laser L, such as Figure 2 As shown in (b), a portion of the first welded component 21 and a portion of the second welded component 22 melt to form a molten pool 31. In the X-direction of the welding head 13's travel direction, unmelted first surfaces 21a and second surfaces 22a exist in front of the molten pool 31. Furthermore, in areas of high energy density irradiated by the laser L within the molten pool 31, molten metal sometimes evaporates, creating a keyhole 32. The first welded component 21 and the second welded component 22 are integrated through the solidification of the molten pool 31. A weld 33 is formed at the joint between the first welded component 21 and the second welded component 22. Therefore, each image constituting the dynamic image D includes any one of the following: first surface 21a, second surface 22a, molten pool 31, keyhole 32, and weld 33.
[0045] like Figure 1 As shown, the acquisition unit 171 acquires multiple images from the images constituting the dynamic image D at predetermined time intervals during the welding process, serving as multiple control input images IA1, IA2, and IA3. Here, an example with three control input images is described, but the number of control input images is not particularly limited, as long as there are two or more. For example, control input image IA3 is the most recent image, and control input image IA2 is an image captured immediately before control input image IA3. Control input image IA1 is an image captured immediately before control input image IA2. Alternatively, control input image IA2 may not be an image captured immediately before control input image IA3, and control input image IA1 may not be an image captured immediately before control input image IA2.
[0046] The image processing unit 172 uses the completed learning model 200 stored in the storage unit 174 to output a feature extraction image IB estimated based on multiple control input images IA1, IA2, and IA3 at predetermined time intervals during the welding process. Hereinafter, the feature extraction image IB is also referred to as the "control feature extraction image IB". Furthermore, the completed learning model 200 is also referred to as the "completed learning model 200".
[0047] The features extracted by the image processing section 172 are outlines and the like of specific regions in the plurality of control-use input images IA1, IA2, and IA3. Here, an example in which the image processing section 172 extracts a plurality of features will be described. However, the number of features extracted by the image processing section is not particularly limited as long as it is one or more.
[0048] The image processing section 172 extracts, as a line R1, an outline of the molten pool 31, as a line R2, an outline of the keyhole 32, as a line R3, a portion of an outline of the first welded member 21, that is, the first face 21a, and as a line R4, a portion of an outline of the second welded member 22, that is, the second face 22a, from the plurality of control-use input images IA1, IA2, and IA3. That is, the control-use feature extraction image IB is an image in which the outline of the molten pool 31 is expressed as the line R1, the outline of the keyhole 32 is expressed as the line R2, the first face 21a is expressed as the line R3, and the second face 22a is expressed as the line R4. However, the features extracted by the image processing section are not limited to the above. For example, the image processing section can extract an outline of a weld as a feature.
[0049] The control section 173 controls the welding section 11 at a prescribed time interval using the control-use feature extraction image IB. Specifically, the control section 173 calculates, from the control-use feature extraction image IB, a deviation between a center position of the keyhole 32 in the Y direction and a center position of a gap between the first face 21a and the second face 22a in front of the keyhole 32 in the Y direction, and controls the arm 14 so as to eliminate the deviation. In addition, the control section 173 controls the output of the light source 12 so that the position of the outline of the molten pool 31 in the Y direction in the control-use feature extraction image IB is outside the first face 21a and the second face 22a and converges within a certain range. Thereby, it is possible to improve the positional accuracy of the welding of the first welded member 21 and the second welded member 22 and the strength of the welding.
[0050] (Learning Model)
[0051] Next, a learned model 200 used by the welding system 10 will be described.
[0052] Figure 4 is a diagram illustrating the learned model of the present embodiment.
[0053] The learned model 200 for the welding system 10 is completed using the teacher data TD.
[0054] The teacher data TD includes a plurality of learning input images IC1, IC2, and IC3, and a learning feature extraction image ID2 extracted from one of the plurality of learning input images IC1, IC2, and IC3. The number of learning input images IC1, IC2, and IC3 used in the one-time learning of the learning model 200 is the same as the number of control input images IA1, IA2, and IA3 used in the one-time image processing, for example, 3.
[0055] The plurality of learning input images IC1, IC2, and IC3 are, for example, 3 images constituting a learning dynamic image in which a welding portion is imaged. The learning dynamic image is imaged, for example, by the imaging device 15. For example, the learning input image IC1 is an image imaged immediately before the learning input image IC2, and the learning input image IC3 is an image imaged immediately after the learning input image IC2. In addition, the imaging device of the learning dynamic image and the imaging device of the control dynamic image D can be different.
[0056] The learning feature extraction image ID2 is, for example, an image extracted from the learning input image IC2, and is prepared by a user of the generation device 40 before the learning of the learning model 200. Specifically, like the control feature extraction image IB, the learning feature extraction image ID2 is an image in which the outline of the molten pool 31 is expressed as a line R5, the outline of the keyhole 32 is expressed as a line R6, the first surface 21a is expressed as a line R7, and the second surface 22a is expressed as a line R8. The learning feature extraction image ID2 is, for example, produced by drawing lines by the manufacturer on portions of the learning input image IC2 identified as the outline of the molten pool 31, the outline of the keyhole 32, the first surface 21a, and the second surface 22a, and extracting the drawn lines. However, the production method of the learning feature extraction image is not limited to the above-described method. In addition, the learning feature extraction image can be, for example, an image extracted from the learning input image IC1 or the learning input image IC3.
[0057] The algorithm for the learning model 200 is an algorithm for producing an image from an image, for example, pix2pix.
[0058] The learning model 200 has a generator 210 and an identifier 220. The generator 210 outputs an extracted image IE of features estimated from a plurality of learning input images IC1, IC2, and IC3. When a pair of the learning input image IC2 and the learning extracted image ID2, and a pair of the learning input image IC2 and the extracted image IE generated by the generator 210 are input, the identifier 220 identifies which pair is the teacher data TD (i.e., true) and which pair is not the teacher data TD (i.e., false). The learning of the generator 210 is continued so that the pair of the learning input image IC2 and the extracted image IE generated by the generator 210 is identified as true by the identifier 220. Further, the learning of the identifier 220 is continued so that the pair of the learning input image IC2 and the learning extracted image ID2 can be identified as true and the pair of the learning input image IC2 and the extracted image IE generated by the generator 210 can be identified as false. Specific processes performed by the generator 210 and the identifier 220 will be described later.
[0059] In the present embodiment, as shown in Figure 1 The learning model 200 is generated by the generation device 40. The generation device 40 is a computer including a processor such as a GPU or a CPU, a ROM, a RAM, a hard disk, and the like. Further, the control device 17 can also generate the learning model.
[0060] (Method of generating learning model)
[0061] Next, the method of generating the learning model 200 will be described.
[0062] Figure 5 is a flowchart showing the method of generating the learning model of the present embodiment.
[0063] The method of generating the learning model 200 includes a process S11 of obtaining the teacher data TD, a process S12 of performing preprocessing on each of the learning input images IC1, IC2, and IC3, and a step S13 of causing the learning model 200 to learn. Each of the processes will be described in detail below.
[0064] Figure 6 is a diagram showing data used in the learning of the learning model of the present embodiment.
[0065] First, the generation device 40 obtains the teacher data TD prepared in advance by the user (process S11). That is, the generation device 40 obtains a plurality of learning input images IC1, IC2, and IC3, and a learning extracted image ID2 in which features are extracted from the learning input image IC2.
[0066] Further, in the present embodiment, the generating device 40 also acquires a pre-processing feature extraction image ID1 in which features are extracted from the learning input image IC1, and a pre-processing feature extraction image ID3 in which features are extracted from the learning input image IC3. The pre-processing feature extraction images ID1, ID3, like the learning feature extraction image ID2, are images in which the outline of the molten pool 31 is extracted as the line R5, the outline of the small hole 32 is extracted as the line R6, the first face 21a is extracted as the line R7, and the second face 22a is extracted as the line R8, and are prepared in advance by the user. The pre-processing feature extraction images ID1, ID3, like the learning feature extraction image ID2, are made by the manufacturer drawing the portions that are considered to be the outline of the molten pool 31, the outline of the small hole 32, the first face 21a, and the second face 22a with lines in the learning input images IC1, IC3, and extracting the drawn lines.
[0067] Figure 7 FIG. 6 is a diagram showing a pre-processing method of a learning input image in a learning model generation method of the present embodiment.
[0068] Next, the generating device 40 pre-processes the learning input images IC1, IC2, and IC3 (step S12).
[0069] Specifically, the generating device 40 uses the pre-processing feature extraction image ID1 to make a first mask M1 in which the values of the pixels that constitute the lines R5, R6, R7, R8 and the periphery of the lines R5, R6, R7, R8 are set to zero, and the values of the pixels other than these are set to 1. Hereinafter, images are also regarded as matrices, and pixels are also referred to as "elements". Further, the generating device 40 makes a second mask M2 in which the values of the elements that constitute the lines R5, R6, R7, R8 and the periphery of the lines R5, R6, R7, R8 are set to 1, and the values of the elements other than these are set to 0 in the pre-processing feature extraction image ID1. Further, in the present embodiment, the generating device 40 makes a third mask M3 in which the values of the elements that constitute the lines R5, R6, R7, R8 and the periphery of the lines R5, R6, R7, R8 are set to 1, and the values of the elements other than these are set to 0 in the pre-processing feature extraction image ID3. Figure 7 In the present embodiment, in order to make the explanation easier to understand, the elements whose values are zero in the first mask M1 and the second mask M2 are represented in black, and the elements whose values are 1 are represented in white.
[0070] Next, the generating device 40 multiplies the elements of the learning input image IC1 and the elements of the first mask M1 with each other. Here, "multiplying the elements with each other" means that, in two matrices of the learning input image IC1 and the first mask M1, and the like, all the elements are multiplied by the element of the i-th row and the j-th column in one matrix and the element of the i-th row and the j-th column in the other matrix. By this, an image M4 in which the features and the periphery of the features are removed from the learning input image IC1 is made.
[0071] In addition, the generation device 40 creates an image M3 in which the entire input image IC1 is blurred, by applying a filter such as a smoothing filter, a Gaussian filter, or a median filter to the input image IC1 for learning. "Blurring" refers to processing that reduces the variation in the gray scale in the image. Then, the generation device 40 multiplies the elements of the image M3 in which the entire input image IC1 is blurred and the elements of the second mask M2. Thus, an image M5 is created in which the feature and the region around the feature are extracted from the image M3 in which the entire input image IC1 is blurred.
[0072] Next, the generation device 40 adds the elements of the image M4 in which the input image IC1 for learning and the first mask M1 are multiplied, and the elements of the image M5 in which the second mask M2 and the image M3 in which the entire input image IC1 is blurred are multiplied. Here, "adding the elements" refers to processing in which the element of the i-th row and the j-th column of one matrix and the element of the i-th row and the j-th column of the other matrix are added in both matrices. Thus, a pre-processed image IM1 is created.
[0073] By performing the above processing, a pre-processed image IM1 in which the feature and the region around the feature of the input image IC1 for learning are blurred, and other regions are not blurred, can be obtained. The generation device 40 also performs the same processing on the input image IC2 for learning, and creates a pre-processed image IM2 of the input image IC2 for learning. In addition, the generation device 40 also performs the same processing on the input image IC3 for learning, and creates a pre-processed image IM3 of the input image IC3 for learning.
[0074] In the process S12, the degree of blurring of the plurality of input images IC1, IC2, and IC3 for learning can be the same or different from each other. The degree of blurring of each input image IC1, IC2, and IC3 for learning can be adjusted, for example, by the weighting value when a filter such as a smoothing filter, a Gaussian filter, or a median filter is applied. In a case where the degree of blurring in the plurality of pre-processed images IM1, IM2, and IM3 is different from each other, the learning of the learning model 200 is continued so that the feature can be extracted in the pre-processed image in which the degree of blurring is the largest among the plurality of pre-processed images IM1, IM2, and IM3.
[0075] However, an image in which the entire input image for learning is blurred can also be input to the input layer of the learning model described later as a pre-processed image. In addition, an input image for learning that is not pre-processed can also be input to the input layer.
[0076] Figure 8 is a diagram illustrating a generator of a learning model according to the present embodiment.
[0077] Next, the generating device 40 causes the learning model 200 to learn using the plurality of pre-processed images IM1, IM2, and IM3, and the feature extraction image ID2 for learning (step S13).
[0078] In the present embodiment, U-NET is used in the generator 210. Specifically, in the present embodiment, the generator 210 includes an input layer 211, a first intermediate layer 212a, a second intermediate layer 212b, a third intermediate layer 212c, a fourth intermediate layer 213a, a fifth intermediate layer 213b, a sixth intermediate layer 213c, and an output layer 214. In addition, in the present embodiment, the number of layers of the generator 210 is six, but the number of layers is not limited to this. Figure 8 In the present embodiment, although an example in which the number of intermediate layers 212a, 212b, 212c, 213a, 213b, and 213c is six is shown, the number of intermediate layers is not limited to this.
[0079] Figure 9 (a) of FIG. 12 is a diagram showing the processing of the input layer in the learning model of the present embodiment. Figure 9 (b) of FIG. 12 is a diagram showing the convolution method in the input layer.
[0080] Hereinafter, in order to make the explanation easy to understand, in a matrix of an image or a filter, or the like, the arrangement direction of the elements in a row is referred to as "lateral direction x", and the arrangement direction of the elements in a column is referred to as "vertical direction y".
[0081] The plurality of pre-processed images IM1, IM2, and IM3 are input to the input layer 211 as one set of data. In the input layer 211, convolution is performed on the one set of pre-processed images IM1, IM2, and IM3. Hereinafter, an example in which convolution is performed by b filters F11, F12, to F1b in the input layer 211, and the convolution kernel size of each filter F11 to F1b is n1 x n1 will be described.
[0082] First, the generating device 40 extracts a region A1 of the same size as the filter F11 in the pre-processed image IM1. Next, the generating device 40 calculates a value r1(i, j) after multiplying the element im1(i, j) of the i-th row and j-th column of the extracted region A1 and the element f1(i, j) of the i-th row and j-th column of the filter F11. The generating device 40 performs the same processing on all the elements im1(i, j) within the region A1. Next, the generating device 40 calculates a value c1(p, q) after adding all the values r1(i, j) calculated with respect to the region A1.
[0083] Likewise, the generating device 40 extracts a region A2 of the same size as the filter F11 and located in the same position as the region Al from the pre-processed image IM2. Next, the generating device 40 calculates a value r2(i, j) after multiplying an element im2(i, j) of the i-th row and j-th column of the extracted region A2 by an element f1(i, j) of the i-th row and j-th column of the filter F11. The generating device 40 performs the same processing for all elements im2(i, j) within the region A2. Next, the generating device 40 calculates a value c2(p, q) after adding all the values r2(i, j) calculated for the region A2.
[0084] Likewise, the generating device 40 extracts a region A3 of the same size as the filter F11 and located in the same position as the region Al from the pre-processed image IM3. Next, the generating device 40 calculates a value r3(i, j) after multiplying an element im3(i, j) of the i-th row and j-th column of the extracted region A3 by an element f1(i, j) of the i-th row and j-th column of the filter F11. The generating device 40 performs the same processing for all elements im3(i, j) within the region A3. Next, the generating device 40 calculates a value c3(p, q) after adding all the values r3(i, j) calculated for the region A3.
[0085] Then, the generating device 40 calculates a value cs(p, q) after adding the calculated values cl(p, q), c2(p, q), and c3(p, q).
[0086] Then, the generating device 40 shifts the regions Al, A2, and A3 to which the filter F11 is applied in the pre-processed images IM1, IM2, and IM3 in the horizontal direction x in order, and likewise calculates the value cs(p, q). After the regions Al, A2, and A3 are shifted to the last row of each of the pre-processed images IM1, IM2, and IM3, the regions Al, A2, and A3 are returned to the initial row, and the regions Al, A2, and A3 are shifted in the vertical direction y, and the same processing is performed. The above processing is repeated until the regions Al, A2, and A3 are shifted to the element belonging to the last row and last column of each of the pre-processed images IM1, IM2, and IM3.
[0087] In addition, in the present embodiment, in the input layer 211, each of the regions Al, A2, and A3 is moved by one element in the horizontal direction x or the vertical direction y. That is, the stride is 1. When each of the regions Al, A2, and A3 is shifted, if each of the regions Al, A2, and A3 protrudes from the pre-processed images IM1, IM2, and IM3, the value of the element of the protruding portion in each of the regions Al, A2, and A3 is set to zero, that is, zero padding is performed. In addition, each of the regions Al, A2, and A3 can be shifted by every two or more elements. That is, the stride can be 2 or more.
[0088] By the above, as shown in (a) of Figure 9 The first feature map P11 in which the element in the p-th row and the q-th column is the value cs(p, q) is created. As described above, in the present embodiment, the regions A1, A2, and A3 are each moved by one element in the horizontal direction x and the vertical direction y. Therefore, the size of the first feature map P11 is the same as the size of each of the pre-processed images IM1, IM2, and IM3.
[0089] Next, the same processing as that of the filter F11 is performed on the filters F12 to F1b. Thus, a plurality of first feature maps P12 to P1b are created. In this way, in the input layer 211, the three pre-processed images IM1, IM2, and IM3 can be convolved as one set of data.
[0090] Figure 10 is a diagram showing that the positions of the outlines of the molten pools differ from each other in the plurality of input images for learning.
[0091] In Figure 10 , the position of the outline of the molten pool 31 of the input image IC1 for learning is indicated by a line R5a, the position of the outline of the molten pool 31 of the input image IC2 for learning is indicated by a line R5b, and the position of the outline of the molten pool 31 of the input image IC3 for learning is indicated by a line R5c.
[0092] As the plurality of input images IC1, IC2, and IC3 for learning, images in which the positions of the features differ from each other and the amount of change Δx, Δy in the positions of the features of the input images IC1, IC2, and IC3 for learning is smaller than the kernel size n1 of each of the filters F11 to F1b are used.
[0093] For example, when the laser L is continuously applied to a certain region of the first and second members 21 and 22 to be welded, the molten pool 31 gradually expands. At this time, in a case where a dynamic image of the welding site is captured by the imaging device 15, the positions of the outlines of the molten pool 31 differ from each other in the images constituting the dynamic image.
[0094] In the present embodiment, as the input images IC1, IC2, and IC3 for each learning, a combination of images in which the maximum variation Δx in the lateral direction x of the position of the outline of the molten pool 31 and the maximum variation Δy in the longitudinal direction y of the position of the outline of the molten pool 31 are smaller than the convolution kernel size n1 of each filter F11 to F1b is selected from among the images constituting the moving image. In order to be able to make such a selection, the time interval, that is, the frame rate at which the imaging device 15 performs imaging is set in a manner in which the variations Δx, Δy in the position of the features of the plurality of input images IC1, IC2, and IC3 for learning are smaller than the convolution kernel size n1 of each filter F11 to F1b. When the frame rate has been determined, it is also possible to reduce the convolution kernel size n1 in a manner in which the variations Δx, Δy in the position of the features of the plurality of input images IC1, IC2, and IC3 for learning are smaller than the convolution kernel size n1 of each filter F11 to F1b. In addition, likewise, it is also possible to increase the angle of view.
[0095] As to the outline of the small hole 32 and the first face 21a and the second face 22a, which are other features, the input images IC1, IC2, and IC3 for learning are also selected in a manner in which the same requirements are satisfied.
[0096] By selecting the plurality of input images IC1, IC2, and IC3 for learning as described above, for example, it is possible to make the possibility higher that, in a case where the features are contained within the region A1 of the same size as the filter F11 in one input image IC1 for learning, the features are also contained within the regions A2, A3 of the same size as the filter F11 in the other input images IC2, IC3 for learning. Therefore, the learning model 200 is able to perform learning to integrate information related to the variations in the positions of the features in the plurality of input images IC1, IC2, and IC3 for learning and to estimate the extraction image IE of the features from the plurality of input images IC1, IC2, and IC3 for learning. Thereby, even when it is difficult to extract the position of the features in one image, it is possible to capture and extract the position of the features with high precision from the variations in the positions of the features of a plurality of images. As a result, it is possible to improve the extraction precision of the features when the plurality of input images IA1, IA2, and IA3 for control are input to the learning model 200.
[0097] In addition, in the present embodiment, an example in which the positions of the features of the plurality of input images IC1, IC2, and IC3 for learning are positions based on the passage of time is described. That is, in the present embodiment, the variations Δx, Δy are generated due to the passage of time. However, as in the other embodiment described later, the variations can not be generated due to the passage of time.
[0098] Figure 11 Fig. 1 is a diagram showing the processing of the first intermediate layer in the learning model according to the present embodiment. Figure 11(b) is a diagram illustrating the processing of the second intermediate layer in the learning model according to this embodiment. Figure 11 (c) is a diagram illustrating the processing of the third intermediate layer in the learning model according to this embodiment.
[0099] Next, as Figure 11 As shown in (a), multiple first feature maps P11 to P1b generated in the input layer 211 are input into the first intermediate layer 212a.
[0100] In the first intermediate layer 212a, multiple first feature maps P12 to P1b are treated as a set of data and convolved by c filters F21, F22 to F2c. Furthermore, the specific method of convolution is the same as that in the input layer 211, except that regions of the same size as each filter F21 to F2c are shifted in the convolved image by at least two features. Therefore, a detailed description of the convolution in the first intermediate layer 212a is omitted here.
[0101] In the first intermediate layer 212a, multiple second feature maps P21, P22, and P2c are generated by convolving multiple first feature maps P12 to P1b with c filters F21 to F2c. In this embodiment, in each first feature map P11 to P1b, the regions where each filter F21 to F2c is applied are shifted by at least two elements. As a result, the size of the multiple second feature maps P21 to P2c is smaller than the size of the multiple first feature maps P12 to P1b.
[0102] Next, as Figure 11 As shown in (b), in the second intermediate layer 212b, multiple second feature maps P21 to P2c are convolved as a group of data by d filters F31, F32 to F3d. This produces d third feature maps P31, P32 to P3d. In this embodiment, in each second feature map P21 to P2c, the regions where filters F31 to F3d are applied are shifted by at least two elements. Therefore, the size of the multiple third feature maps P31 to P3d is smaller than the size of the multiple second feature maps P21 to P2c.
[0103] Next, as Figure 11 As shown in (c), in the third intermediate layer 212c, multiple third feature maps P31 to P3d are convolved as a group of data by e filters F41, F42 to F4e. This produces e fourth feature maps P41, P42 to P4e. In this embodiment, in each third feature map P31 to P3d, the regions where filters F41 to F4e are applied are shifted by at least two features. Therefore, the size of the multiple fourth feature maps P41 to P4e is smaller than the size of the multiple third feature maps P31 to P3d.
[0104] In addition, in the present embodiment, the amount of change in the position of the features of the plurality of input images IC1, IC2, and IC3 for learning, Δx, Δy, is smaller than the convolution kernel size n2 of each filter F21 to F2c of the first intermediate layer 212a, the convolution kernel size n3 of each filter F31 to F3d of the second intermediate layer 212b, and the convolution kernel size n4 of each filter F41 to F4e of the third intermediate layer 212c. Thus, it is easy to propagate information related to the change in the position of the features included in the plurality of first feature maps P11 to P1b from the first intermediate layer 212a to the third intermediate layer 212c.
[0105] Figure 12 (a) of FIG. 10 is a view illustrating the processing of the fourth intermediate layer in the learning model of the present embodiment. Figure 12 (b) of FIG. 10 is a view illustrating the processing of the fifth intermediate layer in the learning model of the present embodiment. Figure 12 (c) of FIG. 10 is a view illustrating the processing of the sixth intermediate layer in the learning model of the present embodiment.
[0106] Next, the plurality of fourth feature maps P41 to P4e produced by the third intermediate layer 212c are input to the fourth intermediate layer 213a. In the fourth intermediate layer 213a, the plurality of fourth feature maps P41 to P4e are deconvoluted as a group of data. Here, "deconvolution" refers to a process of convoluting a filter corresponding to a transpose matrix of a certain filter on an input feature map, assuming that the input feature map is generated by convoluting a certain image with the certain filter.
[0107] Specifically, first, first enlarged maps K11, K12 to K1e that enlarge the size in the horizontal direction x and the size in the vertical direction y of each of the fourth feature maps P41 to P4e are produced. The enlarged maps K11 to K1e are produced by adding elements having a value of zero to the fourth feature maps P41 to P4e. Next, the plurality of first enlarged maps K11, K12, and K13 to K1e are convoluted as a group of data with f filters F51, F52 to F5f. Thus, f fifth feature maps P51, P52 to P5f are produced. Here, the f filters F51, F52 to F5f correspond to a transpose matrix of a certain filter, assuming that the fourth feature maps P41 to P4e are created by convoluting an image with the certain filter. Thus, it is possible to make the size of the plurality of fifth feature maps P51 to P5f output larger than the size of the plurality of fourth feature maps P41 to P4e input.
[0108] Next, as in the first intermediate layer 212a, the plurality of fifth feature maps P51 to P5f are convoluted as a group of data with the plurality of filters F61 to F6c of the fifth intermediate layer 213a. Thus, the plurality of sixth feature maps P61 to P6c are produced. Figure 12As shown in (b), multiple fifth feature maps P51 to P5f generated by the fourth intermediate layer 213a and third feature maps P31 to P3d generated by the second intermediate layer 212b are input into the fifth intermediate layer 213b. In the fifth intermediate layer 213b, the multiple fifth feature maps P51 to P5f and the third feature maps P31 to P3d are convolved as a set of data.
[0109] Specifically, in the fifth intermediate layer 213b, second magnified images K21 to K2f are created by enlarging the horizontal x-axis and vertical y-axis dimensions of multiple fifth feature maps P51 to P5f, and third magnified images K31 to K3d are created by enlarging the horizontal x-axis and vertical y-axis dimensions of multiple third feature maps P31 to P3d. Then, the multiple second magnified images K21 to K2f and the third magnified images K31 to K3d are treated as a set of data and convolved by g filters F61, F62 to F6g. This produces g sixth feature maps P61, P62 to P6g. The size of the output sixth feature maps P61 to P6g is larger than the size of the input fifth feature maps P51 to P5f.
[0110] Next, as Figure 12 As shown in (c), multiple sixth feature maps P61 to P6g generated by the fifth intermediate layer 213b and second feature maps P21 to P2c generated by the first intermediate layer 212a are input into the sixth intermediate layer 213c. In the sixth intermediate layer 213c, the multiple sixth feature maps P61 to P6g and the second feature maps P21 to P2c are convolved as a set of data.
[0111] Specifically, in the sixth intermediate layer 213c, fourth magnified images K41 to K4g are created by enlarging the horizontal x-axis and vertical y-axis dimensions of multiple sixth feature maps P61 to P6g, and fifth magnified images K51 to K5c are created by enlarging the horizontal x-axis and vertical y-axis dimensions of multiple second feature maps P21 to P2c. Then, the multiple fourth magnified images K41 to K4g and the fifth magnified images K51 to K5c are treated as a set of data and convolved through h filters F71, F72 to F7h. This produces h seventh feature maps P71, P72 to P7h. The size of the output seventh feature maps P71 to P7h is larger than the size of the input sixth feature maps P61 to P6g.
[0112] In the present embodiment, the variation amounts Δx, Δy of the positions of the features of the plurality of input images IC1, IC2, and IC3 for learning are smaller than the convolution kernel sizes n5 of the respective filters F51 to F5f of the fourth intermediate layer 213a, the convolution kernel sizes n6 of the respective filters F61 to F6g of the fifth intermediate layer 213b, and the convolution kernel sizes n7 of the respective filters F71 to F7h of the sixth intermediate layer 213c. Thus, it is easy to propagate information about the variation of the positions of the features included in the plurality of fourth feature maps P41 to P4e from the fourth intermediate layer 213a to the sixth intermediate layer 213c.
[0113] Figure 13 is a view showing processing of an output layer in the learning model according to the present embodiment.
[0114] Next, as shown in Figure 13 , in the output layer 214, the plurality of seventh feature maps P71 to P7h are convolved by the 3 filters F81, F82, and F83 as a group of data. Thereby, 3 eighth feature maps P81, P82, and P83 are created.
[0115] In the present embodiment, the variation amounts of the positions of the features of the plurality of input images IC1, IC2, and IC3 for learning are smaller than the convolution kernel sizes n8 of the filters F81 to F83 of the output layer 214. Thus, the learning model 200 is able to perform learning in a manner of integrating the variations of the positions of the features in the plurality of input images IC1, IC2, and IC3 for learning and estimating the extracted image IE of the features from the plurality of input images IC1, IC2, and IC3 for learning.
[0116] Further, in the learning model 200, for example, the number c of the filters F21 to F2c of the first intermediate layer 212a is larger than the number b of the filters F11 to F1b of the input layer 211. In addition, the number d of the filters F31 to F3d of the second intermediate layer 212b is larger than the number c of the filters F21 to F2c of the first intermediate layer 212a. In addition, the number e of the filters F41 to F4e of the third intermediate layer 212c is larger than the number d of the filters F31 to F3d of the second intermediate layer 212b. In addition, the number f of the filters F51 to F5f of the fourth intermediate layer 213a is the same as the number e of the filters F41 to F4e of the third intermediate layer 212c. In addition, the number g of the filters F61 to F6g of the fifth intermediate layer 213b is the same as the number d of the filters F31 to F3d of the second intermediate layer 212b. In addition, the number h of the filters F71 to F7h of the sixth intermediate layer 213c is the same as the number c of the filters F21 to F2c of the first intermediate layer 212a. However, the magnitude relationship of b to h is not limited to the above.
[0117] In addition, in the learning model 200, for example, the convolution kernel size n1 of the input layer 211 is the same as the convolution kernel size n8 of the output layer 214. In addition, for example, the convolution kernel sizes n2 to n7 of the intermediate layers 212a, 212b, 212c, 213a, 213b, and 213c are the same and larger than the convolution kernel size n1 of the input layer 211. However, the size relationship of the convolution kernel sizes n1 to n8 is not limited to the above.
[0118] Figure 14 is a diagram illustrating a feature extraction image output by the generator of the learning model of the present embodiment.
[0119] In the eighth feature map P81, a portion presumed to be the molten pool 31 profile is extracted as a line R9. In the eighth feature map P82, a portion presumed to be the keyhole 32 profile is extracted as a line R10. In the eighth feature map P83, a portion presumed to be the first face 21a is extracted as a line R11, and a portion presumed to be the second face 22a is extracted as a line R12. The combination of the three eighth feature maps P81, P82, and P83 corresponds to the extraction image IE of the features presumed from the plurality of learning input images IC1, IC2, and IC3.
[0120] Next, the pair of the learning input image IC2 and the learning feature extraction image ID2 and the pair of the learning input image IC2 and the feature extraction image IE output by the generator 210 are input to the discriminator 220. Then, the discriminator 220 identifies which pair is true and which pair is false. The generator 210 learns in such a manner that the discriminator 220 identifies the pair of the learning input image IC2 and the feature extraction image IE output by the generator 210 as true, and the generator 210 determines the values of the elements of the filter at the time of convolution or deconvolution. In addition, the discriminator 220 learns in such a manner that the pair of the learning input image IC2 and the learning feature extraction image ID2 is identified as true and the pair of the learning input image IC2 and the feature extraction image IE output by the generator 210 is identified as false. By simultaneously performing the learning of the generator 210 and the learning of the discriminator 220, the learning of both progresses.
[0121] (Welding method)
[0122] Next, a welding method using the learning model 200 of the present embodiment will be described.
[0123] Figure 15 is a flowchart illustrating a welding method using the learning model of the present embodiment.
[0124] In the following description, during welding, the control section 173 controls the welding section 11 so that the laser L is emitted from the welding head 13 while the welding head 13 is gradually moved in the X direction. In addition, during welding, the control section 173 controls the imaging device 15 to capture a dynamic image D of the welding site during welding.
[0125] In a case where welding has started, first, the acquisition section 171 acquires, as the plurality of control input images IA1, IA2, and IA3, the latest image among images that constitute the dynamic image D captured by the imaging device 15 and two images captured immediately before the latest image (step S21). In the present embodiment, the frame rate and the angle of view of the imaging device 15 are set so that the amount of change in the position of the feature of the plurality of control input images IA1, IA2, and IA3 is smaller than the convolution kernel size n1 of the filters F11 to F1b of the input layer 211.
[0126] Next, the image processing section 172 outputs an extracted image IB of the feature estimated from the three input images IA1, IA2, and IA3 using the learning model 200 stored in the storage section 174 (step S22).
[0127] Next, the control section 173 controls the welding section 11 based on the feature extracted image IB output by the image processing section 172 (step S23). Specifically, the control section 173 calculates the deviation between the center position of the small hole 32 in the Y direction and the center position of the gap between the first face 21a and the second face 22a in front of the small hole 32 in the Y direction based on the control feature extracted image IB, and controls the arm 14 to eliminate the deviation. In addition, the control section 173 controls the output of the light source 12 so that the position of the outline of the molten pool 31 in the Y direction in the control feature extracted image IB is located outside the first face 21a and the second face 22a, and converges within a certain range.
[0128] Next, the control section 173 determines whether or not welding is completed (step S24). In a case where it is determined that welding is completed (step S24: Yes), the control section 173 cuts off the output of the laser, and completes welding. In a case where it is determined that welding is not completed (step S24: No), the processes of steps S21 to S24 are performed again.
[0129] Next, the effects of the present embodiment will be described.
[0130] The generation method of the learning model 200 of the present embodiment includes a process of obtaining teacher data TD including a plurality of input images for learning IC1, IC2, and IC3 and a feature extraction image for learning ID2 extracted from a feature in one of the plurality of input images for learning IC1, IC2, and IC3, and a process of causing the learning model 200 to learn using the teacher data TD, the learning model 200 outputting an extraction image IB of a feature estimated from the plurality of input images IC1, IC2, and IC3. The learning model 200 includes an input layer 211 that performs convolution. The positions of the features in each of the plurality of input images for learning IC1, IC2, and IC3 are different from each other. The amount of change Δx, Δy in the positions of the features in the plurality of input images for learning IC1, IC2, and IC3 is smaller than the kernel size n1 of the filters F11 to F1b of the input layer 211.
[0131] In such a generation method of the learning model 200, the learning model 200 is caused to learn in a manner in which the extraction image IE of the feature is estimated from the plurality of input images for learning IC1, IC2, and IC3 in accordance with information in which the amount of change in the positions of the features in the plurality of input images for learning IC1, IC2, and IC3 is integrated. Therefore, when the plurality of input images IA1, IA2, and IA3 are input, the learning model 200 can extract the feature with high accuracy.
[0132] Further, the learning model 200 includes an output layer 214 that performs convolution. The amount of change Δx, y is smaller than the kernel size n8 of the filters F81, F82, and F83 of the output layer 214. Therefore, the learning model 200 is caused to learn in a manner in which the extraction image IE of the feature is estimated from the plurality of input images for learning IC1, IC2, and IC3 in accordance with information in which the amount of change in the positions of the features in the plurality of input images for learning IC1, IC2, and IC3 is integrated. Therefore, when the plurality of input images IA1, IA2, and IA3 are input, the learning model 200 can extract the feature with high accuracy.
[0133] Further, the learning model 200 includes intermediate layers 212a, 212b, and 212c that perform convolution. The amount of change Δx, Δy is smaller than the kernel sizes n2, n3, n4 of the filters F21 to F2c, F31 to F3d, and F41 to F4e of the intermediate layers 212a, 212b, and 212c. Therefore, the learning model 200 is caused to learn in a manner in which the extraction image IE of the feature is estimated from the plurality of input images for learning IC1, IC2, and IC3 in accordance with information in which the amount of change in the positions of the features in the plurality of input images for learning IC1, IC2, and IC3 is integrated. Therefore, when the plurality of input images IA1, IA2, and IA3 are input, the learning model 200 can extract the feature with high accuracy.
[0134] Further, the learning model 200 includes the intermediate layers 213a, 213b, and 213c that perform convolution. The variation amounts Δx, Δy are smaller than the sizes n5, n6, and n7 of the convolution kernels of the filters F51 to F5f, F61 to F6g, and F71 to F7h of the intermediate layers 213a, 213b, and 213c. Thus, the learning model 200 can be caused to learn in a manner in which the extracted image IE of the feature is estimated from the plurality of input images IC1, IC2, and IC3 for learning, based on information on variations in the positions of the features integrated in the plurality of input images IC1, IC2, and IC3 for learning. Thus, when the plurality of input images IA1, IA2, and IA3 are input, the learning model 200 can extract the feature with high accuracy.
[0135] Further, the U-NET is used in the learning model. That is, the feature maps P21 to 2c and P31 to 3d output from the first intermediate layer 212a and the second intermediate layer 212b and the like are input to the deconvolution layers of the fifth intermediate layer 213b and the sixth intermediate layer 213c and the like. Thus, when the plurality of input images IA1, IA2, and IA3 are input, the learning model 200 can extract the feature with high positional accuracy.
[0136] Further, the method of generating the learning model 200 of the present embodiment further includes a process of creating pre-processed images in which the features of the plurality of input images IC1, IC2, and IC3 for learning are blurred, before the learning process. In the learning process, the pre-processed images IM1, IM2, and IM3 are input to the input layer 211. Thus, the learning model 200 can be caused to learn so as to be able to extract the feature even in a strict condition in which the features are blurred.
[0137] Further, in the process of creating the pre-processed images, the degree to which the features are blurred in one of the plurality of input images IC1, IC2, and IC3 for learning is different from the degree to which the features are blurred in the other input images for learning. Thus, the learning model 200 is caused to learn so as to be able to extract the feature even in a case where the degrees to which the features are blurred are different.
[0138] Further, the plurality of input images IC1, IC2, and IC3 for learning are images that constitute a dynamic image obtained by capturing a welding portion corresponding to a subject portion. Thus, it is possible to easily prepare the plurality of input images IC1, IC2, and IC3 for learning in which the positions of the features are different from each other.
[0139] In addition, one of the plurality of learning input images IC1, IC2, and IC3 is an image taken immediately before or after the other learning input images. Therefore, it is possible to easily prepare the plurality of learning input images IC1, IC2, and IC3 in which the variation Δx, Δy in the position of the feature is smaller than the filter kernel size n1 of the convolutional filters F11 to F1b.
[0140] In addition, the plurality of learning input images IC1, IC2, and IC3 are images taken at the time of welding of a welding site, and the features are at least a part of the outline of the molten pool 31, at least a part of the outline of the keyhole 32, or at least a part of the outline of the welded members 21 and 22. Therefore, it is possible to accurately extract features related to welding.
[0141] In addition, the learning completed model 200 according to the present embodiment includes an input layer 211 that performs convolution, and learning is completed using teacher data TD including the plurality of learning input images IC1, IC2, and IC3 and a learning feature extraction image ID2 in which a feature is extracted from one of the plurality of learning input images IC1, IC2, and IC3. The positions of the features of the plurality of learning input images IC1, IC2, and IC3 are different from each other, and the variation Δx, Δy in the position of the feature in the plurality of learning input images IC1, IC2, and IC3 is smaller than the filter kernel size n1 of the convolutional filters F11 to F1b of the input layer 211. Furthermore, the learning completed model 200 causes a computer to output an extraction image IB of a feature that is estimated from the plurality of input images IA1, IA2, and IA3. Therefore, it is possible to provide a learning completed model 200 that can accurately extract a feature when a plurality of input images IA1, IA2, and IA3 are input.
[0142] Furthermore, the image processing method according to the present embodiment includes a process of obtaining a plurality of input images IA1, IA2, and IA3, and a process of outputting, using the learning completed model 200, an extraction image IB of a feature that is estimated from the plurality of input images IA1, IA2, and IA3. Therefore, it is possible to provide an image processing method that can accurately extract a feature when a plurality of input images IA1, IA2, and IA3 are input.
[0143] In addition, since the variation in the position of the feature in the plurality of input images IA1, IA2, and IA3 is smaller than the filter kernel size n1 of the convolutional filters F11 to F1b of the input layer 211, it is possible to accurately extract a feature when a plurality of input images IA1, IA2, and IA3 are input.
[0144] Further, the image processing system according to the present embodiment is provided with the image processing section 172 that outputs the extracted image IB of the feature of the welding estimated from the plurality of input images IA1, IA2, and IA3 using the learned model 200. Thus, it is possible to provide an image processing system that can extract a feature with high accuracy when a plurality of input images IA1, IA2, and IA3 are input.
[0145] Further, the welding system 10 according to the present embodiment is provided with the welding section 11 that welds a plurality of welded members 21, 22, the imaging device 15 that images the welding site of the plurality of welded members 21, 22, the image processing section 172 that outputs the extracted image IB of the feature of the welding estimated from the plurality of images imaged by the imaging device 15 using the learned model 200, and the control section 173 that controls the welding device based on the extracted image IB of the feature output by the image processing section 173. Thus, it is possible to provide a welding system 10 that can produce the extracted image IB of the feature from a plurality of input images IA1, IA2, and IA3 and can control the welding work with high accuracy.
[0146] <Second Embodiment>
[0147] Next, the second embodiment will be described.
[0148] Figure 16 Fig. 1 is a diagram showing a part of the welding system according to the present embodiment.
[0149] Further, in the following description, only the points different from the first embodiment will be described in principle. Except for the matters described below, the same as the first embodiment.
[0150] In the first embodiment, the example in which the image constituting the dynamic image D imaged by the imaging device 15 is used as the input image IA1, IA2, and IA3 for control and the input image IC1, IC2, and IC3 for learning has been described. In contrast, in the present embodiment, the welding system 310 is provided with the imaging device 315 that can acquire a plurality of images different in wavelength, polarization, or exposure time. In the plurality of images different in wavelength, polarization, or exposure time, the positions of the features are sometimes different from each other. Further, the plurality of images different in wavelength, polarization, or exposure time imaged by the imaging device 315 can also be used as the input image IA1, IA2, and IA3 for control and the input image IC1, IC2, and IC3 for learning.
[0151] The imaging device 315 can include a filter that transmits light of mutually different wavelengths, and the imaging device 315 can acquire images corresponding to the respective filters. In this case, one illumination device 16 can emit a plurality of lights of mutually different wavelengths, can emit a wide-area light including a plurality of lights of mutually different wavelengths, or a plurality of illumination devices 16 can be provided and emit lights of mutually different wavelengths. In addition, the imaging device 315 can include a polarizer that transmits light of mutually different polarizing directions, and the imaging device 315 can acquire images corresponding to the respective polarizers. In addition, the imaging device 315 can acquire a non-polarized image and a polarized image. In these cases, one illumination device 16 can emit a plurality of lights of mutually different polarizing directions, or a plurality of illumination devices 16 can be provided and emit lights of mutually different polarizing directions. In addition, the imaging device 315 can include a shutter that acquires images of mutually different exposure times, and the imaging device 315 can acquire images corresponding to the respective exposure times.
[0152] In this case, the wavelengths or the polarizations of the plurality of imaging devices are set so that the amount of change in the positions of the features of the learning input images IC1, IC2, and IC3 is smaller than the kernel size n1 of the plurality of filters F11 to F1b of the input layer 211.
[0153] <Third Embodiment>
[0154] Next, the third embodiment will be described.
[0155] Figure 17 is a view that shows a part of a welding system according to the present embodiment.
[0156] In the present embodiment, the welding system 410 is provided with a plurality of imaging devices 415a, 415b, and 415c that image the welding site from mutually different positions. Also, the images imaged by the plurality of imaging devices 415a, 415b, and 415c can be used as the control input images IA1, IA2, and IA3 and the learning input images IC1, IC2, and IC3.
[0157] In this case, the positions of the plurality of imaging devices 415a, 415b, and 415c are adjusted so that the amount of change in the positions of the features of the learning input images IC1, IC2, and IC3 is smaller than the kernel size n1 of the plurality of filters F11 to F1b of the input layer 211.
[0158] <Fourth Embodiment>
[0159] Next, the fourth embodiment will be described.
[0160] Figure 18 is a diagram showing a part of the welding system of the present embodiment.
[0161] In the present embodiment, the welding system 510 is provided with a plurality of photographing devices 515a, 515b, and 515c, the photographing angles of which are different from each other. Also, the images photographed by the plurality of photographing devices 515a, 515b, and 515c can be used as the input images IA1, IA2, and IA3 for control and the input images IC1, IC2, and IC3 for learning.
[0162] In this case, the photographing angles of the plurality of photographing devices 515a, 515b, and 515c are adjusted so that the amount of change in the positions of the features of the input images IC1, IC2, and IC3 for learning is smaller than the convolution kernel size n1 of the plurality of filters F11 to F1b of the input layer 211.
[0163] As described above, the plurality of input images for learning are images in which the photographing conditions at the time of photographing the welding site are different from each other. There is no particular limitation on the photographing conditions, but as described above, the time at the time of photographing the welding site, the polarization direction of light, the photographing position, the photographing angle, the wavelength of light, and the exposure time, and the like can be cited. The plurality of input images for control are also images in which the photographing conditions at the time of photographing the welding site are different from each other. In addition, in the above-described embodiment, although only a manner in which one photographing condition is different is described, a plurality of photographing conditions can be different.
[0164] In addition, in the above-described embodiment, although only a manner in which the photographing device photographs the welding site during welding is described, the welding site after welding can be photographed. In the case where the welding site after welding is photographed, the image processing system can extract, for example, a weld or the like as a feature, and use the feature extraction image output from the image processing system for determination of the precision of welding or the like.
[0165] In addition, in the above-described embodiment, although only a manner in which the image processing system is realized by the control device of the welding system is described, the device that realizes the image processing system is not limited to the above. The image processing system can also be realized by an edge device attached to the photographing device. In addition, the image processing system can also be realized by a computer that processes images uploaded to the cloud. In addition, the image processing system can also be realized by a plurality of computers.
[0166] Further, the image processing system can also be applied to systems other than the welding system.
[0167] The above describes embodiments of the present application, but these embodiments are presented as examples and are not intended to limit the scope of the application. These new embodiments can be implemented in other various ways, and various omissions, substitutions, and changes can be made without departing from the scope of the application. These embodiments and modifications thereof are included in the scope and spirit of the application, and are included in the scope of the application and equivalents thereof recited in the claims.
[0168] Explanation of Reference Signs
[0169] 10, 310, 410, 510: welding system
[0170] 11: weld
[0171] 12: light source
[0172] 13: welding head
[0173] 14: arm
[0174] 15, 315, 415a, 415b, 415c, 515a, 515b, 515c: imaging device
[0175] 16: illuminating device
[0176] 17: control device
[0177] 17b: ROM
[0178] 17c: RAM
[0179] 17d: hard disk
[0180] 17e: bus
[0181] 21, 22: member to be welded
[0182] 21a: first surface
[0183] 22a: first surface
[0184] 31: molten pool
[0185] 32: keyhole
[0186] 33: weld
[0187] 40: generating device
[0188] 171: acquisition section
[0189] 172: image processing section
[0190] 173: control section
[0191] 174: storage section
[0192] 200: learning model
[0193] 210: generator
[0194] 211: input layer
[0195] 212a: first intermediate layer
[0196] 212b: second intermediate layer
[0197] 212c: third intermediate layer
[0198] 213a: fourth intermediate layer
[0199] 213b: fifth intermediate layer
[0200] 213c: sixth intermediate layer
[0201] 214: output layer
[0202] 220: recognizer
[0203] A1, A2, and A3: areas
[0204] F11 to F1b: filters
[0205] F21 to F2c: filters
[0206] F31 to F3d: filters
[0207] F41 to F4e: filters
[0208] F51 to F5f: filters
[0209] F61 to F6g: filters
[0210] F71 to F7h: filters
[0211] F81 to F83: filters
[0212] IA1, IA2, and IA3: multiple input images for control
[0213] IB: feature extraction image for control
[0214] IC1, IC2, IC3: multiple input images for learning
[0215] ID2: feature extraction image for learning
[0216] ID1, ID3: feature extraction images for pre-processing
[0217] IE: feature extraction image for learning
[0218] IM1 to IM3: pre-processed images
[0219] K11 to K1e: first enlarged view
[0220] K21 to K2f: second enlarged view
[0221] K31 to K3d: third enlarged view
[0222] K41 to K4g: fourth enlarged view
[0223] K51 to K5c: fifth enlarged view
[0224] L: laser
[0225] M1: first mask
[0226] M2: second mask
[0227] M3: image after overall blurring P11 to P1b: first feature map P21 to P2c: second feature map P31 to P3d: third feature map P41 to P4e: fourth feature map P51 to P5e: fifth feature map P61 to P6g: sixth feature map P71 to P7h: seventh feature map P81 to P83: eighth feature map R1, R2 to R12, R5a, R5b, R5c: line TD: teacher data
[0228] f1: element
[0229] im1: element
[0230] im2: element
[0231] im3: element
[0232] n1 to n8: convolution kernel size
[0233] x: lateral direction
[0234] y: longitudinal direction
[0235] Δx: change amount
[0236] Δy: change amount
Claims
1. A method of generating a learning model, wherein, comprises: a process of acquiring teacher data including a plurality of learning input images and a learning feature extraction image obtained by extracting a feature from one of the plurality of learning input images; and a process of causing a learning model to learn using the teacher data, the learning model outputting an extraction image of the feature estimated from a plurality of input images, the learning model including an input layer that performs convolution, the positions of the features in the respective plurality of learning input images being different from each other, the amount of change in the positions of the features in the plurality of learning input images being smaller than the size of a convolution kernel of a filter of the input layer.
2. The learning model generation method according to claim 1, wherein the learning model includes an output layer that performs convolution, the amount of change being smaller than the size of a convolution kernel of a filter of the output layer.
3. The learning model generation method according to claim 1 or 2, wherein the learning model includes an intermediate layer that performs convolution, the amount of change being smaller than the size of a convolution kernel of a filter of the intermediate layer.
4. The learning model generation method according to any one of claims 1 to 3, wherein the learning model includes another intermediate layer that performs deconvolution, the amount of change being smaller than the size of a convolution kernel of a filter of the other intermediate layer.
5. The learning model generation method according to any one of claims 1 to 4, wherein U-NET is used in the learning model.
6. The learning model generation method according to any one of claims 1 to 5, wherein the process further includes a process of creating a plurality of pre-processed images obtained by blurring the features of the plurality of learning input images, before the learning process, in the learning process, the plurality of pre-processed images are input to the input layer.
7. The learning model generation method according to claim 6, wherein in the process of creating the plurality of pre-processed images, the degree of blurring the features in one of the plurality of learning input images is different from the degree of blurring the features in the other learning input images.
8. The learning model generation method according to any one of claims 1 to 7, wherein the plurality of learning input images are images in which at least one of a time when a subject is photographed, a polarization direction of light, a photographing position, a photographing angle, a wavelength of light, and an exposure time is different from each other.
9. The learning model generation method according to claim 8, wherein the plurality of learning input images are images in which at least one of a time when a subject is photographed, a polarization direction of light, a photographing position, a photographing angle, a wavelength of light, and an exposure time is different from each other.
10. The learning model generation method according to any one of claims 1 to 7, wherein the plurality of learning input images are images constituting a dynamic image of a subject.
11. The learning model generation method according to any one of claims 1 to 10, wherein The plurality of learning input images are images obtained by photographing a welding site at the time of welding; The feature is at least a part of a contour of a molten pool, at least a part of a contour of a keyhole, or at least a part of a contour of a welded component.
12. An image processing method comprising: a step of obtaining a plurality of input images; and a step of outputting, using a learned model, an extraction image of a feature estimated from the plurality of input images, the learned model includes an input layer that performs convolution, and learning is completed using teacher data including a plurality of learning input images and a learning feature extraction image obtained by extracting the feature from one of the plurality of learning input images, the positions of the feature in the plurality of learning input images being different from each other, and the amount of change in the positions of the feature in the plurality of learning input images being smaller than the size of a convolution kernel of a filter of the input layer.
13. The image processing method according to claim 12, wherein the amount of change in the positions of the feature in the plurality of input images is smaller than the size of a convolution kernel of a filter of the input layer.
14. An image processing system comprising: an image processing unit that outputs, using a learned model, an extraction image of a feature estimated from a plurality of input images, the learned model includes an input layer that performs convolution, and learning is completed using teacher data including a plurality of learning input images and a learning feature extraction image obtained by extracting the feature from one of the plurality of learning input images, the positions of the feature in the plurality of learning input images are different from each other, the amount of change in the positions of the feature in the plurality of learning input images is smaller than the size of a convolution kernel of a filter of the input layer.
15. A welding system comprising: a welding unit that welds a welded component; one or more photographing devices that photograph a welding site of the welded component; an image processing unit that outputs, using a learned model, an extraction image of a feature of welding estimated from a plurality of images photographed by the photographing devices; and a control unit that controls the welding unit based on the extraction image of the feature output by the image processing unit, the learned model includes an input layer that performs convolution, and learning is completed using teacher data including a plurality of learning input images and a learning feature extraction image obtained by extracting the feature from one of the plurality of learning input images, the positions of the feature in the plurality of learning input images are different from each other, the amount of change in the positions of the feature in the plurality of learning input images is smaller than the size of a convolution kernel of a filter of the input layer.
Citation Information
Patent Citations
Laser processing system
JP2019141902A
Molten pool contour detection method based on deep neural network
CN110363781A
Welding system, and method for welding workpiece in which same is used
WO2020129618A1