Metal product surface paint spraying color quality inspection method based on deep learning
By constructing a color recognition and consistency evaluation model based on deep learning, the problem of difficult and time-consuming and labor-consuming for metal paint color machines is solved, and high accuracy and high efficiency color recognition and evaluation are achieved.
Patent Information
- Application Number
- CN202510232224.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The color of metal paint is difficult to identify and manually identify the time-consuming and labor-consuming problem of machine and manual identification. Especially due to environmental factors and lighting changes, the color difference characteristics are not obvious, making it difficult to accurately compare with standard color cards.
A deep learning-based method is used to build a color recognition model (YOLOv5 and Resnet-18) and a color consistency evaluation model (U-net and twin neural networks) to realize pixel-level segmentation prediction and color consistency evaluation of metal workpiece images.
It improves the accuracy and efficiency of metal workpiece paint color recognition, can better identify the color of workpiece paint under the background and match it with standard color cards, reducing the cost and time of manual inspection.
Smart Images

Figure CN120147735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of workpiece image processing, and particularly to a method for quality inspection of the painting color on the surface of metal products based on deep learning. Background Technique
[0002] With the continuous development of China's industrial technology, the demand for the workpiece manufacturing industry has also been increasing year by year. Some metals in the process need to be subjected to surface painting treatment. As a part of the delivery standard, the color of metal workpieces must eliminate the workpieces with colors different from the standard color card. However, due to the differences in surface painting processes, the existence of threads and holes, and the large interference of the paint coating on the metal surface by light traps, shadows, and ambient light, the color difference features are not obvious, making it difficult to compare with the standard color card, resulting in certain difficulties in color quality inspection.
[0003] Nowadays, in the quality inspection of the painting color of workpieces in industry, most still use manual methods to judge the quality of the colors of production parts and standard parts. However, the disadvantages of this method are that it is impossible to unify the detection environmental conditions, and even the same inspector may make mistakes in judging the workpieces due to state reasons, which undoubtedly reduces the precision and detection efficiency of industrial production.
[0004] Due to the nature of the metal itself, its color is affected by various environmental factors, and has a great impact on the image shooting and subsequent color judgment. In order to obtain consistent measurement results, the CIE standard colorimetric system is currently mostly used for judgment, including the CIE1931 XYZ standard colorimetric system, the CIE1976 LUV and CIE1976 LAB color spaces, to meet the needs of color judgment in different scenarios.
[0005] Due to machine and human errors during the production process, there will be a certain deviation between the metal painting color and the standard color required for production. Moreover, the color differences between similar metals are extremely small, and the metal color changes greatly with light and environment. Due to reasons such as the sensitivity of inspectors to metal color errors, manual inspection is in a situation of great difficulty, high cost, long time consumption, and low accuracy. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method for quality inspection of the painting color on the surface of metal products based on deep learning, so as to realize pixel-level segmentation prediction of workpieces.
[0007] To achieve the above purpose, the present invention adopts the following technical solutions: A method for quality inspection of the painting color on the surface of metal products based on deep learning includes the following steps:
[0008] Step 1: Determine the painting color of the metal workpiece to be detected and the color of the standard color card, collect the data set, process and label the data set, and use a script to convert the format of the generated annotation file;
[0009] Step 2: Build a color recognition model, that is, use YOLOv5 and Resnet-18 for color recognition to achieve the recognition of the color types of metal workpieces after image input;
[0010] Step 3: Build a color consistency evaluation model, that is, first use the U-net network for target segmentation, and then perform twin neural network training for qualification inspection;
[0011] Step 4: Adjust and improve the model parameters to obtain a network model, and compare the performance of the two; finally, take pictures of the workpieces under actual conditions to verify the applicability of the model and analyze the advantages of the model in different scenarios.
[0012] In a preferred embodiment, in step 1, a darkroom is built during the process of collecting images of metal workpieces. The metal workpiece is placed in the darkroom, illuminated with lights of fixed position and fixed light intensity, and the equipment for collecting images is placed at a fixed height and position; the pictures are labeled. Since the shape of the metal workpiece is an irregular polygon, a json file corresponding to the picture is generated using labelme under the segmentation network, and then the json file tags are converted into a completely black mask for segmentation using a script.
[0013] In a preferred embodiment, in step 2, first send the image of the metal workpiece to be detected into YOLOv5 for color target segmentation; in the labeled dataset, after the image extracts the texture and metal color features on the surface of the workpiece through the network, the recognized feature parameter information is further enhanced, and finally used to achieve the segmentation of the target workpiece. After that, the segmented image is sent into the Resnet-18 color classification network, and after passing through the convolutional layer, the max-pooling layer, and the fully connected layer, the model finally outputs the color category of the workpiece.
[0014] In a preferred embodiment, in step 3, the U-Net network includes a decoder, an encoder, and a bottleneck layer; the input image enters the encoder part, and after path contraction, it enters the decoder part through the extended path;
[0015] The process is as follows:
[0016] (1) After the image is input, through a number of convolutional kernels and downsampling, the encoder starts to extract the target features of the workpiece and learn; the convolutional structure used in the first to fifth layers is uniformly three 3×3 convolutional kernels, as shown in formula (1);
[0017]
[0018] where n out represents the output feature map, n inDenote the input feature map, p represents the padding number, k represents the kernel size of the pooling layer, and s represents the stride;
[0019] According to padding = 0, striding = 1, we get
[0020] n out = n in - 2(2)
[0021] The first 4 convolutional layers enter the next layer through max pooling. The kernel size of each pooling layer is k = 2, padding = 0, striding = 2. So there is
[0022]
[0023] The output feature map of the 5th convolutional layer is sent to the decoder of the extended path;
[0024] (2) The decoder part consists of convolution, upsampling and skip connection structures; The feature map output by the previous encoder is sent into the decoder, restored to the original resolution, and cascaded with the convolutional layer of the encoding for feature mapping. The skip connection is used to fuse the position information of the shallow layer and the semantic information of the deep layer;
[0025] (3) When the entire network outputs an image finally, the response values of each channel at each pixel position are normalized. According to Equation (4), K represents the number of categories, a k (x) represents the response value of pixel x at feature channel k, p k (x) represents the probability of the corresponding category;
[0026]
[0027] Among them, k' represents the category sequence.
[0028] In a preferred embodiment, in step 3, the formula of the weighted cross-entropy weighted loss function is as shown in (5):
[0029] E = ∑ x∈Ω w(x)log(p l ( x )(x)) (5)
[0030] Among them, w(x) is the pixel weight. When the pixel is closer to the cell boundary, w(x) is larger, otherwise smaller, so as to enhance the learning of edge pixels. The specific calculation method is:
[0031]
[0032] Among them, w c (x) represents the weight map for balancing the category frequency, w 0Denote the initial weight set as d 1 d(x) represents the distance to the nearest cell boundary 2 d(x) represents the distance to the second nearest cell boundary
[0033] In a preferred embodiment, in the Siamese neural network, the input x 1 and x 2 are fed into the constructed neural network f to obtain the feature vectors h 1 and h 2 . The distance between the two feature vectors h 1 and h 2 is used to evaluate the similarity degree of the two inputs; the two feature vectors h 1 and h 2 are subtracted and then the absolute value is taken to obtain the vector z = |h 1 -h 2 |, which represents the difference between these two vectors. Through several fully connected layers, finally, the sigmoid activation function is used to map the value between 0 and 1
[0034] This finally output sim(x 1 , x 2 ) is used to measure the similarity between two pictures. If the two pictures are similar, the output should be close to 1; otherwise, it should be close to 0. 1 represents the same class, 0 represents different classes. Combining the label and the just output sim(x 1 , x 2 ) can select the loss function to calculate the loss, and then perform gradient descent and backpropagation to update the parameters from the fully connected layer parameters to the convolutional layer parameters
[0035] In a preferred embodiment, in the Siamese neural network, the loss function adopted is
[0036] (1) The contrastive loss function, defined as shown in Equation (7)
[0037]
[0038] where N represents the number of samples; X 1 , X 2 are the input data pairs; Y = 1 indicates that the result predicted by the model is that the two inputs belong to similar classes of data, otherwise Y = 0; m represents the distance threshold between the two inputs when they are not similar, that is, the distance between two dissimilar samples is in [0, m]. When it exceeds the distance threshold, that is, the dissimilarity between the two samples is 0
[0039] E w is defined as the Euclidean distance between the outputs of the Siamese neural network, that is
[0040] E w = |X 1-X 2 | 2 (8)
[0041] The contrast loss function is mainly used for dimensionality reduction, that is, in the feature extraction of images;
[0042] (2) Triplet loss function, whose input is a triplet including an anchor, a positive sample, and a negative sample; the definition is shown in Equation (9);
[0043] L = max(d(a, p) - d(a, n) + margin, 0) (9)
[0044] Where a is the anchor; p is the positive sample, which is of the same category as a; n is the sample, which is of a different category from a; margin is a constant greater than 0; the output L of the whole formula should follow the criterion of shortening the distance between a and p and lengthening the distance between a and n.
[0045] Compared with the prior art, the present invention has the following beneficial effects: Aiming at the problem that it is difficult for machines to identify the colors of metal paints and manual identification is time-consuming and laborious, the present invention studies a quality inspection method for the colors of metal surface paints based on deep learning, and proposes a color recognition model based on YOLOv5 and Resnet-18 to detect the color types of input metal workpiece images; and a color consistency evaluation model based on U-net and Siamese neural network to detect whether the input metal workpiece images are qualified. Both models have high recognition accuracy and can better identify the paint colors of workpieces in the background and match them with the standard color cards. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is an image of the paint color on the metal surface collected in the preferred embodiment of the present invention;
[0047] Figure 2 It is an image of the standard color card collected in the preferred embodiment of the present invention;
[0048] Figure 3 It is an image of the actual metal workpiece in the preferred embodiment of the present invention;
[0049] Figure 4 It is a schematic diagram of the dark box in the preferred embodiment of the present invention;
[0050] Figure 5 It is a schematic diagram of the lighting in the preferred embodiment of the present invention;
[0051] Figure 6 It is a schematic diagram of the image obtained by taking supplementary light measurement according to different workpieces in the preferred embodiment of the present invention;
[0052] Figure 7 It is the color card acquisition process in the preferred embodiment of the present invention;
[0053] Figure 8 Conversion script for the preferred embodiment of the present invention;
[0054] Figure 9 Original image and segmented mask image of the preferred embodiment of the present invention;
[0055] Figure 10 Schematic diagram of the overall network structure of the preferred embodiment of the present invention;
[0056] Figure 11 Schematic diagram of the YOLOv5 network structure of the preferred embodiment of the present invention;
[0057] Figure 12 Schematic diagram of the residual function of the preferred embodiment of the present invention;
[0058] Figure 13 Schematic diagram of the residual function structure of the preferred embodiment of the present invention;
[0059] Figure 14 Schematic diagram of the Resnet-18 network structure of the preferred embodiment of the present invention;
[0060] Figure 15 Schematic diagram of the overall network structure of the color consistency evaluation model of the preferred embodiment of the present invention;
[0061] Figure 16 Schematic diagram of the U-net network structure of the preferred embodiment of the present invention;
[0062] Figure 17 Schematic diagram of the siamese neural network of the preferred embodiment of the present invention;
[0063] Figure 18 Schematic diagram of the relationship between the contrast loss function value and the Euclidean distance of the samples of the preferred embodiment of the present invention;
[0064] Figure 19 Schematic diagram of the principle of the triplet loss function of the preferred embodiment of the present invention;
[0065] Figure 20 Schematic diagram of the training and prediction accuracy curves and loss curves of the model of the preferred embodiment of the present invention;
[0066] Figure 21 Schematic diagram of the model prediction confusion matrix of the preferred embodiment of the present invention;
[0067] Figure 22 Schematic diagram of the target segmentation effect of the preferred embodiment of the present invention;
[0068] Figure 23 Schematic diagram of the model demonstration interface of the preferred embodiment of the present invention;
[0069] Figure 24 Schematic diagram of the training and prediction accuracy curves and loss curves of the U-net model according to the preferred embodiment of the present invention;
[0070] Figure 25 Schematic diagram of the training and prediction accuracy curves and loss curves of the Siamese neural network model according to the preferred embodiment of the present invention;
[0071] Figure 26 Schematic diagram of the input demonstration interface according to the preferred embodiment of the present invention;
[0072] Figure 27 Schematic diagram of the input image and recognition result according to the preferred embodiment of the present invention;
[0073] Figure 28 Schematic diagram of the recognition of the real-shot image results of the same color category according to the preferred embodiment of the present invention. Detailed implementation manners
[0074] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0075] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0076] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0077] A method for quality inspection of the paint color on the surface of metal products based on deep learning, referring to Figure 1-28 , includes the following steps: Step 1: Determine the paint color of the metal workpiece to be detected and the color of the standard color card, collect the data set, process and label the data set, and use a script to convert the format of the generated annotation file;
[0078] Step 2: Build a color recognition model, that is, use YOLOv5 and Resnet-18 for color recognition to realize the recognition of the color types of the metal workpiece after image input;
[0079] Step 3: Build a color consistency evaluation model, that is, first use the U-net network for target segmentation, and then perform Siamese neural network training for qualification inspection;
[0080] Step 4: Adjust the parameters of the improved model to obtain a network model, and compare the performance of the two. Finally, take pictures of the workpiece under actual conditions to verify the applicability of the model and analyze the advantages of the model in different scenarios.
[0081] Specifically, dataset construction:
[0082] Introduction to the workpieces used in the experiment:
[0083] A total of 8 types of painted colors on the metal surface to be detected were used for image acquisition. There are 2 - 3 different specifications of workpieces under each color classification. There are 16 colors in total for the standard color cards used to compare with the photographed workpieces.
[0084] Among them, the actual workpiece diagrams of various colors are as Figure 3 shown.
[0085] (a) The paint surface of the GR class features a black smooth glaze, and the edge of the texture is golden - red. The workpiece includes three styles, two are cuboids, and one is a small frustum with a reticulated concave - convex surface.
[0086] (b) The paint surface of the MB class features black matte particles. The workpiece includes three styles, all of which are cuboid - like.
[0087] (c) The paint surface of the US4 class features a brass - colored smooth glaze, and the surface has filamentous longitudinal textures. The workpiece includes three styles, two are cuboids, and one is a cylinder with a rhombic concave - convex surface.
[0088] (d) The paint surface of the SS class features a silver - colored smooth glaze, which easily reflects the surrounding light, and the surface has filamentous longitudinal textures. The workpiece includes two styles, both of which are cuboids.
[0089] (e) The paint surface of the US7 class features a golden mirror surface, and the surface has filamentous longitudinal textures. The workpiece includes three styles, one is a sunken platform - type with fish - scale - like concave - convex surfaces at the edge, and two are irregular combined convex - types.
[0090] (f) The paint surface of the US15A class features a lead - black smooth glaze, and the surface has filamentous longitudinal textures. The workpiece includes two styles, both of which are irregular convex - types.
[0091] (g) The paint surface of the USIOB class is a black matte paint surface. The workpiece includes three styles, two are cuboid - like, and one is a small frustum with a reticulated concave - convex surface.
[0092] (h) The paint surface of the US19 class features a black smooth glaze. The workpiece includes two styles, both of which are cuboids.
[0093] Data acquisition
[0094] Due to the reflective properties of metals, the painted colors on their surfaces are extremely susceptible to environmental background and lighting factors. Therefore, the following operations were carried out during image acquisition to extract more features during subsequent network training:
[0095] (1) Simulating the production line under real conditions, the same background color was used for image acquisition of the workpieces. At the same time, to avoid the reflection of the surrounding background color onto the workpieces and causing color confusion, we chose off-white as the background for image acquisition to reduce color interference factors.
[0096] (2) Since the workpieces are sensitive to light changes, in order to unify the color acquisition of the workpieces, a darkroom was constructed during the image acquisition process in this paper. The workpieces were placed in the darkroom, illuminated with lights of fixed position and fixed light intensity, and the image acquisition device was placed at a fixed height and position to simulate the situation of the detection device on the production line. The schematic diagram of the darkroom settings and the actual shooting situation are as Figure 4 shown.
[0097] In the schematic diagram, the outermost cuboid is the set dark box, which does not transmit light externally to ensure that the color of the workpieces is not affected by ambient light and the ambient background. There are openings on the top of the dark box, and the acquisition device and the upper-side light are set respectively. There is an opening on the right side to set the right-side light. The metal workpiece to be acquired is placed in the center of the dark box.
[0098] (3) Since the workpieces are relatively thick and have holes, in order to avoid the influence of a single light source on the exposure effect of the workpiece surface and affect the color, for the same metal workpiece, three different lighting effects, namely upper-side fill light, right-side fill light, and upper-side and right-side fill light, will be used for shooting. The placement angle and position are changed multiple times under each lighting condition, and 20 pictures are acquired. A total of 60 pictures can be acquired for each workpiece. The lighting situation is shown as Figure 5 shown.
[0099] During the shooting process, it was found that for smooth metal workpieces with a certain thickness, although the upper-side fill light reduced the shadow area, it also increased the specular noise on the workpiece surface. The right-side fill light was prone to producing large-area shadows. Therefore, the effect of using both upper-side and right-side fill light at the same time was the best. For smooth metal workpieces with a certain thickness and a matte surface, good color pictures could be acquired with upper-side fill light alone. The right-side fill light was prone to producing large-area shadows, which was not convenient for subsequent image segmentation and recognition.
[0100] For thin smooth-surface metal workpieces, large-area shadows are not likely to appear with right-side fill light, while large-area specular noise will appear with upper-side fill light and upper-side and right-side fill light at the same time; conversely, for thin matte workpieces, upper-side fill light alone is sufficient. According to the specific workpiece situation, the upper-side fill light, right-side fill light, and upper-side and right-side fill light at the same time are as Figure 6 shown.
[0101] (4) When collecting standard color cards, the color cards should occupy as much of the picture as possible, and at the same time, the original etched fonts on the color cards should be avoided to affect the network's judgment of color features. The fonts of the color cards should be cropped, and only the color and texture parts of the color cards should be retained. After cropping, each color card collection image has a total of 5 pictures, and 16 color cards have a total of 80 pictures. The color card collection process is as follows: Figure 7 shown.
[0102] Specifically, the data processing and labeling acquisition equipment outputs an image format of 5792×4344px. The metal workpiece image is reduced to 640×480 and then the image is labeled. Since the shape of the workpiece is an irregular polygon, labelme is used under the segmentation network to generate a json file corresponding to the image, and then a script is used to convert the json file label into a segmented black mask. The conversion script and the final effect are shown below. Figure 8 and Figure 9 shown.
[0103] The final collected original workpiece dataset contains 8 metal color categories, each category contains images of three lighting conditions, a total of 1,200 metal workpiece images, and the dataset size after scaling is 45.9MB. 80% of the images are used as training sets, with a total of 960 images, and 20% of the images are used for testing and verification, with a total of 240 images in the test set. The color card dataset contains 16 metal color categories, and after cropping, there are a total of 80 color card images, and the dataset size is 50MB.
[0104] Build a color recognition model based on YOLOv5 and Resnet-18:
[0105] The color recognition model consists of a YOLOv5 target segmentation network and a Resnet-18 color classification network. Figure 10 As shown in the figure, the metal workpiece image to be detected is first sent to YOLOv5 for color target segmentation; in the labeled dataset, the image is extracted through the network to obtain the texture and metal color features of the workpiece surface, and then the identified feature parameter information is further enhanced, and finally used to achieve the segmentation of the target workpiece. After that, the segmented image is sent to the Resnet-18 color classification network, and after the convolution layer, the maximum pooling layer, and the fully connected layer, the model finally outputs the workpiece color category.
[0106] The structure of YOLOv5 is divided into three parts: backbone network, neck network, and detection head. Figure 11 As shown. Among them:
[0107] (1) The backbone network is usually a convolutional neural network with strong adaptability and small lightweight, which extracts as many image features as possible while maintaining the efficiency of the network. Commonly used networks include CSPDarknet-53, etc. In the first stage of the model, the CSP module is not added, and only a convolution with a large kernel is used for the first downsampling, and then the CSP module is used.
[0108] (2) The neck network is the network layer responsible for further processing and enhancing the feature map. Usually, PANet (Path Aggregation Network) and FPN (Feature Pyramid Network) are used in it. The CSP module is added to the feature pyramid, and at the same time, a depth factor is added to adjust the depth. Feature fusion is enhanced through top-down and bottom-up paths.
[0109] (3) The head network, as the last part of YOLOv5, adopts an anchor-free approach to achieve target classification and predict the precise location of the detected target.
[0110] Specifically, the Resnet-18 color recognition network is designed as shown in Equation (1) to solve the problem of network degradation.
[0111] H(x) = F(x) + x (1)
[0112] After simple transformation, the residual function is as shown in Equation (2), and the specific structure is as Figure 12 shown.
[0113] F(x) = H(x) - x (2)
[0114] As long as F(x) = 0, it constitutes an identity mapping H(x) = x. At this time, it is easier to fit the residual. F(x) is the network mapping before summation, and H(x) is the network mapping from the input to after summation.
[0115] The residual learning structure consists of a forward neural network and a direct connection (shortcut), as Figure 13 shown. Among them, the latter constitutes an identity mapping. However, compared with deep networks, it has fewer parameters and computational complexity, can be optimized through faster convergence, and solves the problem of network accuracy degradation in backpropagation training. The principle of the identity mapping is as shown in Equation (3) and Equation (4).
[0116] y = F(x, Wi) + x (3)
[0117] F = W 2 σ(W, x) (4)
[0118] The shortcut path can directly output the input x. However, in more practical applications, it is necessary to perform 1×1 convolution for dimensionality increase or downsampling to make the shape of the path output consistent with that of the F(x) path output. The two structures are as Figure 13 shown.
[0119] The Resnet-18 network structure is as Figure 14 shown. After the first two convolutional layers, a 7×7 convolutional layer with 64 output channels and a stride of 2 is followed by a 3×3 max pooling layer with a stride of 2. A batch normalization layer is added after the convolutional layer, and 4 modules composed of residual blocks are used, with each module using several residual blocks with the same number of output channels. When the image is input into the first convolutional layer, the consistency between the image channels and the convolutional layer channels is ensured. In each subsequent module, the number of channels of the previous module is doubled in the first residual block, and the height and width are halved.
[0120] Its specific network parameters are shown in Table 1. There are five convolutional stages in the network, each stage contains 3 convolutional layers and a fully connected layer, and each layer extracts features from the input image through different numbers of convolutional kernels and filters. The final fully connected layer is used for tasks such as classifying or regressing the extracted features.
[0121] Table 1 Resnet Network Parameters
[0122]
[0123]
[0124] Specifically, the construction of the color consistency evaluation model based on U-net and Siamese network is as follows:
[0125] The color consistency evaluation model consists of a U-net target segmentation network and a Siamese neural network. As Figure 15 shown, first, the image of the metal workpiece to be detected is sent into the U-net network. After convolution and downsampling by the encoder, it is sent into the decoder through the skip connection layer, and after upsampling and convolution, the segmented image is output. The image is sent into one of the subnets of the Siamese network, and the standard color card is sent into the other subnet of the Siamese network. By training and comparing the loss function, the similarity of the colors of the two pictures is evaluated. Finally, the model outputs the result of whether the color of the workpiece image and the standard color card belong to the same category and whether the color of the workpiece is qualified.
[0126] The main structure of the U-Net network includes three parts: a decoder, an encoder, and a bottleneck layer. The structure is as Figure 16 shown. The input image enters the encoder part shown in the red box. After path contraction, it enters the decoder part shown in the purple box on the right through the expansion path. The middle gray arrow line is called the skip connection.
[0127] The process is as follows:
[0128] (1) After the image is input, through the convolution kernels and downsampling of several of the four program blocks, the encoder starts to extract the target features of the workpiece and learn. The convolution structure used in the first to fifth layers is uniformly three 3×3 convolution kernels, as shown in formula (5).
[0129]
[0130] According to padding = 0 and striding = 1, we get
[0131] n out = n in - 2 (6)
[0132] The first four convolutional layers enter the next layer through max pooling. The kernel size of each pooling layer is k = 2, padding = 0, and striding = 2. Therefore, there is
[0133]
[0134] The fifth convolutional layer outputs the feature map and sends it to the decoder in the extended path.
[0135] (2) The decoder part consists of convolution, upsampling, and skip connection structures. The feature map output by the previous encoder is sent into the decoder, restored to the original resolution, and cascaded with the convolutional layer of the encoding in terms of feature mapping. The skip connection is used to fuse the position information of the shallow layer and the semantic information of the deep layer.
[0136] (3) When the entire network finally outputs an image, the response values of each channel at each pixel position are normalized. According to formula (8), K represents the number of categories, a k (x) represents the response value of pixel x at feature channel k, and p k (x) represents the probability of the corresponding category.
[0137]
[0138] The weighted cross-entropy weighted loss function formula is as shown in (9):
[0139] E = ∑ x∈Ω w(x) log(p l ( x )(x)) (9)
[0140] Among them, w(x) is the pixel weight. When the pixel is closer to the cell boundary, w(x) is larger, otherwise smaller, so as to enhance the learning of edge pixels. The specific calculation method is as follows:
[0141]
[0142] The Siamese network structure is as follows Figure 17 shown. The input x 1 and x 2 are fed into the constructed neural network f to obtain the feature vectors h 1 and h 2 . The distance between the two feature vectors can be used to evaluate the similarity of the two inputs. Subtract the two vectors and then take the absolute value to get the vector z = |h 1 - h 2 |, which represents the difference between the two vectors. Through several fully connected layers, finally, the sigmoid activation function is used to map the value between 0 and 1. The blue dashed line represents the backpropagation process.
[0143] This final output sim(x 1 , x 2 ) is used to measure the similarity between two images. If the two images are similar, the output should be close to 1; otherwise, it should be close to 0. 1 represents the same class, and 0 represents different classes. Combining the labels and the just output sim(x 1 , x 2 ) can select the loss function to calculate the loss, and then perform gradient descent, and backpropagation is used to update the parameters from the fully connected layer parameters to the convolutional layer parameters.
[0144] The loss function usually adopts:
[0145] (1) Contrastive Loss, defined as shown in Equation (11).
[0146]
[0147] Where N represents the number of samples; X 1 , X 2 are the input data pairs; Y = 1 indicates that the result predicted by the model is that the two inputs belong to similar classes of data, otherwise Y = 0; m represents the distance threshold between the two inputs when they are not similar, that is, the distance between two dissimilar samples is in [0, m], and when the distance threshold is exceeded, that is, the dissimilarity between the two samples is 0.
[0148] E w is defined as the Euclidean distance between the outputs of the Siamese neural network, that is:
[0149] E w = |X 1 - X 2 | 2 (12)
[0150] The contrast loss function is mainly used for dimensionality reduction, that is, feature extraction of images. Normally, assuming that the two original inputs belong to different categories, after a series of network processing such as convolution and downsampling, even if some high-dimensional information is lost, they are still regarded as two dissimilar categories in the mapped feature space, so as not to confuse the category information at the output.
[0151] like Figure 18 As shown, it is the relationship between the loss function value and the Euclidean distance of the sample features, where the red dotted line represents the loss value of similar samples and the blue solid line represents the loss value of dissimilar samples.
[0152] (2) Triplet Loss: Its input is a triplet consisting of an anchor point, a positive sample, and a negative sample. The similarity between samples is calculated by optimizing the distance between the anchor point and the positive sample to make it smaller than the distance between the anchor point and the negative sample. The definition is shown in formula (13).
[0153] L=max(d(a,p)-d(a,n)+margin,0) (13)
[0154] Where a is the anchor point; p is a positive sample, which is of the same category as a; n is a sample, which is of a different category from a; margin is a constant greater than 0. The output L of the integer should follow the principle of shortening the distance between a and p and increasing the distance between a and n. The actual network diagram is as follows Figure 19 As shown in the figure, the standard color card is used as the anchor point, the artifacts of the same color are used as positive samples, and the artifacts of different colors are used as negative samples. The network is trained to optimize the distance between the artifacts of the same color and maximize the distance between artifacts of different colors.
[0155] The samples can be divided into three categories:
[0156] 1. Simple triples refer to triples that satisfy L = 0 without training. At this time, the network does not need training to meet the requirements of the loss function, that is:
[0157] d(a,p)+margin<d(a,n) (14)
[0158] 2. Difficult triples: L>margin. The distance between the negative sample and the anchor point is smaller than the distance between the positive sample and the anchor point. This will cause a large fluctuation in the loss value, that is:
[0159] d(a,p)<d(a,n) (15)
[0160] 3. For general triples, L < margin is satisfied, and the distance between the negative sample and the anchor point is greater than the distance between the positive sample and the anchor point, that is:
[0161] d(a, p) < d(a, n) < d(a, p) + margin (16).
[0162] Experimental Results and Analysis
[0163] Programming is carried out using the Python language. Since the GPU is used in the model training process of the experiment, the model is configured to be trained on a cloud server in this paper, and the network model file is generated and then deployed locally for demonstration. The GPU used by the cloud server is RTX4090 (16GB), the tensorflow version is 2.1, the cuda version is 10.1, and ubuntu is 18.04; the local CPU is 11th Gen Intel(R) Core(TM) i7-1165G7@2.80GHz 2.80GHz, and the GPU is NVIDIA GeForce MX450.
[0164] During training, the batch_size is set to 32, the num_workers is set to 2, the initial learning rate is 1e-5, the number of training epochs is set to 50, and the Adam optimizer is set to update the neural network parameters according to the gradient information to minimize the loss function. In the learning rate scheduler, the learning rate is decayed by 3% per epoch. During the training process, the model is saved every 5 times. When it is judged that the current score is greater than the historical highest score, the model is saved and the highest score and the best model path are updated.
[0165] During the 50 epochs of model training and prediction, the accuracy curve and loss curve plotted are as Figure 20 shown. As can be seen from the figure, until the 15th epoch of training, the accuracy curve shows a continuous and steady upward trend, and the loss curve shows a continuous and steady downward trend. After this training epoch, the training accuracy value of the model fluctuates stably above 90%, and the prediction accuracy value fluctuates around 95%, indicating that the model can correctly predict the category of the input number; the training loss value fluctuates stably below 0.2, and the prediction loss value fluctuates stably below 0.1, indicating that the model can correctly capture the rules in the training data. The prediction accuracy curve is above the training accuracy curve, and the prediction loss curve is below the training loss curve, indicating that the model has not experienced overfitting.
[0166] After the color recognition model is trained, as shown in Table 2, the prediction effect of USIOB is better, while that of MB is worse. The reason is the nature of the workpiece surface itself. The former is a matte surface and is less likely to be confused with other color categories. The latter is a smooth glazed surface. Since there are many glazed surface processes in the test workpieces, it is easy to cause confusion. Due to the imbalance of color categories in the training model, the average data cannot be used as the overall model effect. In the final model prediction results, the accuracy rate reaches 97.68%, the recall rate is 97.6%, the precision is 0.9776, and the F1 score is 0.9768, indicating that the model's ability to accurately retrieve and comprehensively retrieve is at a good level, meeting the requirements of industrial actual production detection.
[0167] Table 2 Prediction Results of Various Colors
[0168]
[0169]
[0170] After making predictions on the test set, print the confusion matrix as Figure 21 shown. As can be seen from the figure, most of the predicted sample labels match the actual sample labels correctly, but there are a small number of cases where the MB class is misclassified as the USIOB class. This is because the two classes are generally similar in color and the patterns are both non-reflective, resulting in incorrect model recognition. The segmentation effect of the model on the image is as Figure 21 shown, and it can basically detect the contour of the metal workpiece and perform the next step of color type recognition.
[0171] Use the Gradio library to create an interactive interface. On the left side of the interface, upload the original image to be recognized. You can upload the image by switching to select local file reading, camera shooting, or clipboard copying. The "Clear" button deletes the uploaded image, and the "Color Recognition" button performs color type recognition on the input image.
[0172] The specific interface is as Figure 23 shown. On the right side of the interface, output the color recognition results. Output the top five confidence levels of the color types recognized from the image to the recognition results, and output the corresponding color category with the highest confidence level as the color type of the workpiece. In order to simulate the situation in industrial production where qualified products are misrecognized as unqualified products due to accidental lighting or incorrect camera focusing, and to improve the anti-interference ability of the model, the following rule is set, that is, when the highest confidence level ranking is still less than 50%, it is output as an unqualified product.
[0173] During the recognition process, in order to simulate the situation in a real production line where there are color categories identical to those of other qualified workpieces but the color error exceeds the range of qualified products, it is necessary to conduct a robustness test on the model. Select pictures from a dataset and perform preprocessing on the images by adjusting the contrast by 10% - 40% respectively. Simulate the situation of unqualified workpieces and send them into the model for detection again. Ideally, the correct category should be output but the confidence level should decrease.
[0174] After the contrast decreases, the confidence levels and misrecognition types of workpieces of each color are shown in Table 3. At a contrast of 10%, the confidence levels of all colors are within the range of qualified products. Among them, the confidence levels of USIOB, MB, SS, and US4 remain at a relatively high level, and the confidence level of US19 decreases more significantly; at a contrast of 20%, the confidence levels of most colors are within the range of qualified products, US4 and USIOB maintain relatively high confidence levels, and the US19 category has been identified as unqualified; at contrasts of 30% and 40%, all colors are in the category of unqualified products, indicating that the image no longer meets the detection standard at this time.
[0175] Table 3 Confidence levels of the correct category of the model at different contrasts
[0176]
[0177]
[0178] It is found that in the case of a contrast of 10% - 20%, the reason for the relatively high confidence level of US4 is that as the only gold mirror workpiece in the dataset, neither its color nor its texture feature vector is likely to be confused with other workpieces; the confidence level of USIOB remains relatively high because its matte surface is not prone to reflection, and black is less affected by the decrease in contrast, so it has strong anti-interference ability. The reason for the relatively fast decrease in the confidence level of US15A is that the types of workpieces participating in the training are relatively similar, and it is easy for the model to identify them as a type of workpiece, and its color pattern is similar to that of other workpieces, so the confidence level is not high.
[0179] After the contrast decreases by more than 30%, there is a situation of confusion among similar color categories, and the detection effect is poor. The color detection is basically regarded as unqualified. One reason is that some of the metal workpieces being detected are prone to high reflection themselves, resulting in the proportion of reflected light pixels in the color of the workpiece itself being too large. For some workpieces such as USIOB with a matte surface, when the contrast decreases and the precise details of the image are lost, it is easy to misidentify the same color as the matte paint color process.
[0180] Second, although the YOLOv5 model and Resnet-18 are relatively fast, to achieve good recognition results, a large amount of data is required for training. However, the scale of the collected dataset and the manually labeled dataset is small, and false detections of similar colors are likely to occur. The dataset obtained in actual industrial production is much larger than the collected dataset. Therefore, when training, try to input image data in various situations so that the model can accurately fit more cases of unqualified workpieces during training.
[0181] Training and Results of the Color Consistency Evaluation Model Based on U-net and Siamese Network
[0182] Define the loss function of U-net as the binary cross-entropy loss function, batch_size as 32, num_workers as 8, the initial learning rate as 1e-4, and the number of training epochs as 50. The training and prediction accuracy curves and loss curves of the model are as Figure 24 shown. In the first 5 epochs of training, the accuracy curve continued to rise steadily, and the loss curve continued to decline steadily. After 5 epochs of training, the training accuracy value and prediction accuracy value of the model stabilized around 95%; after 30 epochs of training, the training loss value and prediction loss value stabilized below 0.1.
[0183] Define the loss function of the Siamese neural network as the contrastive loss function, set margin = 1, batch_size = 4, and num_workers = 1. Define the initial learning rate as 1e-4, and the number of training epochs as 50. Since the indicators in the initial stage of the model are meaningless, the accuracy of the validation set can be evaluated after the 30th epoch of training to shorten the time-consuming of the training process.
[0184] During the training process, print the accuracy and loss values. The final predicted value curve and loss curve of the model are as Figure 25 shown. Until before the 40th epoch of training, that is, when the model is not yet stable, the accuracy curve generally shows a continuous and steady upward trend, and the loss curve generally shows a continuous and steady downward trend; after that, the training accuracy value and prediction accuracy value of the model stabilize around 98%; after 30 epochs of training, the training loss value and prediction loss value stabilize below 0.025.
[0185] After training the color consistency evaluation model, the accuracy rate in the predicted results of the model reaches 98.56%, the recall rate is 98.6%, the precision is 0.9856, and the F1 score is 0.9858. This shows that the model's ability to accurately retrieve and retrieve all is at a good level, and it can accurately output whether the color of the workpiece image is qualified, meeting the requirements of actual industrial production detection.
[0186] Use the Gradio library to create an interactive interface. The demonstration interface is as Figure 26As shown, the upper left image 1 section uploads the user image, and the lower left image 2 section uploads the standard color card image. The right side outputs the workpiece segmentation result, workpiece color recognition type result, standard color card recognition result, and the consistency judgment result of the two in the user uploaded image.
[0187] After identifying the color card and the input image respectively, the conclusion that the output workpiece conforms to the standard color card is not simply taken as the same identification category label. Instead, when the distance between the input user image and the standard color card image in the twin network is less than 0.5, the final output is concluded to be of the same category as the standard color card. Otherwise, the output is not of the same category. Therefore, it is possible that the categories are the same but the output comparison results are different, which is consistent with the situation in industrial production where the detected colors belong to the same category but the product color is unqualified.
[0188] Taking the US4 color card as an example, the output of the input image is as follows: Figure 27 As shown in the table. It can be seen that in the US4 category, a workpiece image is selected to be compared with the US4 standard color card. The recognition result of the workpiece image before being processed and put into the network is US4, and the calculated distance from the standard color card is less than 0.5, which meets the requirements of qualified products, so a similar result is output; the workpiece image is reduced in brightness and the recognition result is still US4 when put into the network, but the distance from the standard color card is greater than 0.5, which belongs to the situation where the color of the workpiece meets the category to which it belongs, but the color quality cannot meet the industrial standard, and the final output is inconsistent with the color card; in order to simulate the situation where other workpieces are mixed into the normal production line, an image is selected in the SS category and put into the network, and the color type is identified as SS, and the output is not similar to the standard color card. In the three cases, the segmentation of metal workpieces has achieved good results, and the contour of the workpiece is basically segmented.
[0189] In order to meet the actual industrial production needs, a non-gray background and non-fixed light intensity are introduced for re-testing, such as Figure 28 As shown in the figure, the final color recognition accuracy remains above 90%, but there is a slight shape shift in the artifact segmentation. This is because the artifact shadow causes the model to recognize the artifact covered by the shadow as a shadow and ignore it.
[0190] The high recognition rate is due to the fact that U-net is suitable for training small data sets and precise pixel-level segmentation, which just meets the data set collection and annotation requirements of this article. In addition, the generalization degree of the twin neural network is good, and it still has a high recognition rate for real-life pictures. However, since it is necessary to compare one by one with the standard color card set and calculate the distance between images, the detection and segmentation time for a single image is relatively long. It is suitable for scenes with fewer standard color cards and requiring high-precision detection of color pass rate.
[0191] Aiming at the problem that it is difficult for metal painting color machines to identify and manual identification is time-consuming and laborious, this invention has studied a quality inspection method for metal surface painting colors based on deep learning, and proposed a color recognition model based on YOLOv5 and Resnet-18 to detect the color types of input metal workpiece images; and a color consistency evaluation model based on U-net and Siamese neural network to detect whether the input metal workpiece images are qualified. Both models have a high recognition accuracy and can better identify the painting colors of workpieces in the background and match them with the standard color card.
Claims
1. A method for quality inspection of paint color on metal product surfaces based on deep learning, characterized in that: The following steps are involved: Step 1: Determine the paint color of the metal workpiece to be tested and the color of the standard color card, collect the data set, process and annotate the data, and use the script to convert the format of the generated annotation file; Step 2: Build a color recognition model, that is, use YOLOv5 and Resnet-18 for color recognition to realize the color type recognition of metal workpieces after image input; Step 3: Build a color consistency judgment model, that is, first use the U-net network to segment the target, and then perform twin neural network training for qualification inspection; Step 4: Adjust and improve the model parameters, obtain the network model, and compare the performance of the two. Finally, take pictures of workpieces in actual situations to verify the applicability of the model and analyze the advantages of the model in different scenarios.
2. The method for quality inspection of metal product surface paint color based on deep learning according to claim 1 is characterized in that: In step 1, a darkroom is constructed during the process of collecting images of metal workpieces, the metal workpieces are placed in the darkroom, illuminated by lights at a fixed position and fixed light intensity, and the image collection equipment is placed at a fixed height and position; Annotate the image. Since the shape of the metal workpiece is an irregular polygon, use labelme to generate a json file corresponding to the image under the segmentation network, and then use the script to convert the json file label into a segmented black mask.
3. The method for inspecting the color of paint on the surface of metal products based on deep learning according to claim 1 is characterized in that: In step 2, the image of the metal workpiece to be inspected is first sent to YOLOv5 for color target segmentation; in the labeled dataset, the image is extracted through the network to obtain the surface texture and metal color features of the workpiece, and then the identified feature parameter information is further enhanced, and finally used to achieve the segmentation of the target workpiece. After that, the segmented image is sent to the Resnet-18 color classification network, and after the convolution layer, the maximum pooling layer, and the fully connected layer, the model finally outputs the workpiece color category.
4. The method for quality inspection of metal product surface paint color based on deep learning according to claim 1 is characterized in that: In step 3, the U-Net network includes a decoder, an encoder, and a bottleneck layer; the input image enters the encoder part, and after the path is contracted, it enters the decoder part through the expansion path; The process is as follows: (1) After the image is input, it passes through several convolution kernels and downsampling, and the encoder begins to extract the target features of the workpiece and learn. The convolution structure used in the 1st to 5th layers is unified into three 3×3 convolution kernels, as shown in formula (1); Among them, n out represents the output feature map, n in represents the input feature map, p represents the padding number, k represents the kernel size of the pooling layer, and s represents the step size; According to padding = 0, striding = 1, we get n out =n in -2 (2) The first four convolutional layers enter the next layer through maximum pooling. The kernel size of each pooling layer is k=2, padding=0, striding=2, so there is The 5th convolutional layer outputs a feature map, which is fed into the decoder of the extended path; (2) The decoder part consists of convolution, upsampling and skip structures. The feature map output by the previous encoder is sent to the decoder, restored to the original resolution, and concatenated with the feature map of the encoder convolution layer. The skip connection is used to fuse the shallow position information with the deep semantic information. (3) When the entire network outputs the image, the response values of each channel at each pixel position are normalized. According to formula (4), K represents the number of categories, a k (x) represents the response value of pixel x at feature channel k, p k (xv represents the probability of the corresponding category; Among them, k' represents the category sequence.
5. The method for inspecting the color of paint on the surface of metal products based on deep learning according to claim 4 is characterized in that: In step 3, the weighted cross entropy loss function formula is shown in (5): E=∑ x∈Ω w(x)log(p l(x) (x)) (5) Among them, w(x) is the pixel weight. When the pixel is close to the cell boundary, w(x) is larger, otherwise it is smaller, so as to enhance the learning of edge pixels. The specific calculation method is: Among them, w c (x) represents the weight map of balanced class frequencies, w0 represents the initial weight set, d1(x) represents the distance to the nearest cell border, and d2(x) represents the distance to the second nearest cell border.
6. The method for quality inspection of metal product surface paint color based on deep learning according to claim 1 is characterized in that: In the twin neural network, the input x1 and x2 are sent to the constructed neural network f to obtain the feature vectors h1 and h2. The distance between the two feature vectors h1 and h2 is used to judge the similarity of the two inputs. The two feature vectors h1 and h2 are subtracted and then the absolute value is calculated to obtain the vector z = |h1-h2|, which represents the difference between the two vectors. After passing through several fully connected layers, the sigmoid activation function is used to map the value between 0 and 1. The final output sim(x1,x2) is used to measure the similarity between two images. If the two images are similar, the output should be close to 1; otherwise, it should be close to 0.1 for the same class, and 0 for different classes. Combining the label and the output sim(x1,x2) just now, we can select the loss function to calculate the loss, and then perform gradient descent and back propagation to update the parameters from the fully connected layer parameters to the convolutional layer parameters.
7. The method for inspecting the color of paint on the surface of metal products based on deep learning according to claim 6 is characterized in that: In the twin neural network, the loss function is: (1) Contrastive loss function, defined as shown in formula (7); Where N is the number of samples; X1, X2 are input data pairs; Y = 1 means that the model predicts that the two inputs belong to similar categories, otherwise Y = 0; m represents the distance threshold between the two inputs when they are dissimilar, that is, the distance between the two dissimilar samples is [0, m]. When the distance threshold is exceeded, the dissimilarity between the two samples is 0; E w It is defined as the Euclidean distance between the outputs of the twin neural network, that is: AND w =|X1-X2|2 (8) The contrast loss function is mainly used for dimensionality reduction, that is, feature extraction of images; (2) Triplet loss function, whose input is a triplet including anchor point, positive sample and negative sample; the definition is shown in formula (9); L=max(d(a,p)-d(a,n)+margin,0) (9) Where a is the anchor point; p is a positive sample, which is of the same category as a; n is a sample, which is of a different category from a; margin is a constant greater than 0; the integer output L should follow the principle of shortening the distance between a and p and increasing the distance between a and n.
Citation Information
Cited By
Vehicle painting make-up color matching method based on computer vision
CN120823414A