Self-supervised multi-category industrial product image defect detection method and device based on twin flow model
Through a self-supervised method based on the twin flow model, realistic defect samples are generated and the loss function is optimized, which solves the problem of high computing resources and storage costs of the flow model in multi-category detection and achieves efficient multi-category defect detection.
Patent Information
- Application Number
- CN202510840609.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing flow models have problems with high computational and storage costs and slow inference speed in multi-category industrial product image defect detection, and the self-supervised defect detection method is insufficiently applied, which affects the detection effect.
A self-supervised multi-category industrial product image defect detection method based on the twin flow model is adopted. By generating realistic artificial defect samples, combining the feature fusion and self-supervised training of the twin flow model, the loss function is optimized to improve the detection accuracy and speed.
While ensuring the accuracy of multi-category detection, it reduces computing and storage resources, improves the inference speed, and realizes lightweight multi-category defect detection.
Smart Images

Figure CN120765553A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image defect detection, and in particular relates to a self-supervised multi-category industrial product image defect detection method and device based on a twin flow model. Background Art
[0002] Recently, research on industrial anomaly detection has made significant progress, combining the design and application of various model architectures. With the development of inference models and architectures, multi-category anomaly detection has become a popular branch in the field of visual anomaly detection. The main challenge in this field lies in the diverse characteristics of various industrial product styles. Most methods focus on increasing network parameters to improve perception capabilities, with only a few research methods focusing on the actual computational and storage costs. Furthermore, due to the limited perception of target distributions by the sample feature mapping model, few lightweight flow model methods have been applied to multi-category anomaly detection and localization settings. Due to the limited number of parameters in the flow model, the model's defect detection capabilities are significantly affected. Balancing the number of parameters and performance is also a challenge that models must address and deal with in practical deployments.
[0003] The Normalizing Flow model essentially maps the distribution of the inspection target directly to a desired distribution through a flow model. A standard normal distribution is often chosen as the mapping target. Because this mapping process is reversible and can be combined with a known mapping target, the distribution of the image or its original features can be derived through reverse tracing, allowing analysis of whether the sample under test contains defects.
[0004] Because flow models inherently model and estimate the distribution of normal images, self-supervised defect detection methods are rarely used in flow model training to ensure the most accurate estimation of the normal distribution. Experimental and deployment studies have demonstrated that the loss of synthetic defects through self-supervised defect detection using realistic defects improves the flow model's normal distribution estimation training results, providing an effective path for improving multi-category defect detection results using flow models. Summary of the Invention
[0005] The purpose of the present invention is to provide a self-supervised multi-category industrial product image defect detection method and device based on a twin flow model to solve the above-mentioned technical problems.
[0006] To solve the above technical problems, the specific technical solutions of the self-supervised multi-category industrial product image defect detection method and device based on the twin flow model of the present invention are as follows:
[0007] A self-supervised multi-category industrial product image defect detection method based on a twin flow model includes the following steps:
[0008] Step 1: Randomly sample from a multi-category dataset of normal samples, randomly cut and flip another part of the normal samples, and combine them with self-supervised Poisson editing to generate realistic artificial defect samples;
[0009] Step 2: The artificial abnormal image and the normal image are passed through a pre-trained feature extractor, and the deep-shallow fusion features and deep features are respectively input into the twin stream model. The deep features and shallow features are fused through upsampling, convolution, concatenation, and linear mapping alignment. The result is used as the input of stream model 1, and the result of the deep features is used as the input of stream model 2.
[0010] Step 3: Upsample the output normal feature distribution estimates of flow model 1 and flow model 2 respectively, and fuse the feature estimation maps of the two outputs by taking the mean;
[0011] Step 4: Estimate the negative logarithm distribution of normal samples. Use the fused feature estimation distribution map to compare the estimation of the self-supervised process with the true anomaly segmentation map. Calculate the loss of defect localization and use it as the adversarial term for the negative logarithm estimation of normal samples. If the mixed loss meets the requirements at the current step, continue training until the required computational accuracy is met.
[0012] Furthermore, the process of generating realistic artificial defect samples in step 1 includes the following steps:
[0013] Step 1.1: First generate a mask set containing rectangular shapes Among them, at the boundary between the normal area and the abnormal area, the width w of each rectangle is i ) and height h i ) follows the Gamma distribution, that is, w i ~Γ(α w ,β w ) and h i ~Γ(α h ,β h ), where i = 1,…,k;
[0014] Step 1.2: With random rotation and flipping Perform element-by-element dot multiplication to get the output
[0015] Step 1.3: Place X and Perform seamless combination based on Poisson editing, and finally obtain
[0016] Furthermore, the steps of Poisson editing are as follows:
[0017] S1: Randomly reduce the original image X locally, and according to According to the pixel value, the color of the corresponding area is changed by enhancing, attenuating, and replacing the RGB color channel gradient;
[0018] S2: Set the upper, lower, left, and right boundary pixels of the new image block to X and The boundary pixels and half of , and solve the gradient fields of the two separately;
[0019] S3: Smoothing the image: Combining X and The gradient relationship, under the premise of the pixel value of X, is The gradient is constrained to solve the pixel value and the same method is used to solve the pixel value. Solve for the appropriate pixel value and average the gradient results obtained from the two as the pixel value of the abnormal area.
[0020] Furthermore, the step 2 includes the following specific steps:
[0021] Step 2.1: Apply a pre-trained open-source feature extractor f ea and multi-layer perceptron feedforward network MLFF to obtain the fused deep non-defect and defect features, namely H deep 、 Among them H deep 、 Directly used as input to flow model 2;
[0022] Step 2.2: Using shallow features H shallow 、 and deep features H deep 、 Combining projection and interpolation, we get the input of flow model 1:
[0023]
[0024] where LP(·), C(·), Intp(·), and Conv(·) denote linear projection, concatenation, interpolation, and convolution layers, respectively.
[0025] Furthermore, the step 3 includes the following specific steps:
[0026] Step 3.1: A multi-scale twin flow MSSF is proposed to estimate the probability and obtain a reversible multi-step twin flow f sia ={f 1 :H shallow →Z shallow ;f 2 :H deep →Z deep}, each flow f consists of K flow blocks {f1,…,f K}, the unknown normal image distribution p H Bidirectional mapping to an accessible latent space distribution p which is usually set to a Gaussian distribution Z , the reversible process is as follows:
[0027]
[0028] Step 3.2: Under differential expression, the formula for the probability estimation process is as follows:
[0029]
[0030] Here, |·| represents the calculation of the determinant. After replacing it with the flow f in the mapping process, the expression is written as follows:
[0031]
[0032] Each step f of the flow model f i They can be expressed using the following expressions:
[0033]
[0034] If it is step 1, enter H in For H = {H deep ,H shallow},as well as As for the last step of the flow, that is, the Kth step, the corresponding original sample log-likelihood atlas H out yes It contains both defect-free samples and artificial anomaly samples from mixed sources. In addition, due to the step-by-step estimation of the Jacobian matrix H and Z can be converted to each other in each step, where Subnet represents a 3×3 convolution block Conv 3×3 and activation function ReLU, H in,1 Corresponding to H in The initial half of the features, and H in,2 Corresponding to the remaining part, H′ in,1,1 and H′ in,1,2 denotes similar definitions, with C(·), s(·), and t(·) representing the concatenation, scaling, and exponential functions, respectively.
[0035] Furthermore, the step 4 includes the following specific steps:
[0036] Step 4.1: The basis of the flow-based method is to optimize the original defect-free image feature H = {H shallow ,H deepThe distribution of} H (h), to be as close as possible to the non-defective zero mapping, so the ideal result is to make p H (h) close to zero, twin flow structure f sia ={f 1 ,f 2 The loss of} is calculated by negative log-likelihood -log(p H (h)) as follows:
[0037]
[0038] Step 4.2: Mapping Results Jacobian matrix By limiting the upper and lower bounds of tanh(s(H′ in,1,1 )) is estimated, where tanh is the hyperbolic tangent function,
[0039]
[0040] Step 4.3: Design artificial anomalies in the flow model as a new type of loss that hinders the training of the flow model, and the average L1 loss L of the abnormal positioning pixels mask They are listed below:
[0041]
[0042] where Y norm and are the mappings of zero anomalies for the input normal samples and synthetic defect mappings, respectively;
[0043] Step 4.4: Final training loss L tr The adversarial approach of negative log-likelihood and segmentation loss is used as follows:
[0044] L tr =L neg -λ·epoch·L mask
[0045] where λ is an influence L mask Parameters of constraint ability, a larger λ indicates stronger adversarial behavior, and epoch indicates the current training round number.
[0046] The self-supervised multi-category industrial product image defect detection method and device based on the twin flow model of the present invention have the following advantages:
[0047] The present invention's self-supervised multi-category industrial product image defect detection method and device based on the twin flow model makes up for the deficiencies of existing flow model methods and devices in the use of self-supervised information. Through the design of the model architecture, the loss function is adapted to multi-category detection tasks and detection targets. While ensuring the accuracy of multi-category detection, this technology makes up for the shortcomings of the existing multi-step flow model framework in multi-category defect detection application scenarios, such as insufficient reasoning speed and too many model parameters. The use of a lightweight twin flow model structure saves computing and storage resources, speeds up reasoning, and has obvious advantages over existing multi-category defect methods and devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Flowchart of training and testing of traditional normalized flow model.
[0049] Figure 2 Flowchart for feature extraction of twin flow model.
[0050] Figure 3 Flowchart for self-supervised realistic sample mixing.
[0051] Figure 4 Flowchart of a simple basic image enhancement method.
[0052] Figure 5 This is a schematic diagram of the physical structure of the device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of the self-supervised multi-category industrial product image defect detection method and device based on the twin flow model of the present invention in conjunction with the accompanying drawings.
[0054] The present invention provides a self-supervised multi-category industrial product image defect detection method based on a twin flow model, comprising the following steps:
[0055] Step 1: Randomly sample from a multi-category dataset of normal samples, randomly cut and flip another part of the normal samples, and combine them with self-supervised Poisson editing to generate realistic artificial defect samples.
[0056] The process of generating realistic artificial defect samples in step 1 includes the following steps:
[0057] Step 1.1: First generate a mask set containing rectangular shapes Among them, at the boundary between the normal area and the abnormal area, the width of each rectangle (w i ) and height (h i ) follows the Gamma distribution, that is, w i ~Γ(αw ,β w ) and h i ~Γ(α h ,β h ), where i = 1,…,k.
[0058] Step 1.2: With random rotation and flipping Perform element-by-element dot multiplication to get the output
[0059] Step 1.3: Place X and Perform seamless combination based on Poisson editing, and finally obtain The steps of Poisson editing are as follows:
[0060] (1) Randomly reduce the original image X locally, and according to According to the pixel value situation, the color of the corresponding area is changed by enhancing, attenuating, and replacing the RGB color channel gradient.
[0061] (2) Set the upper, lower, left, and right boundary pixels of the new image block to X and The boundary pixels and half of , and solve the gradient fields of the two separately.
[0062] (3) Smooth the image. Combine X and The gradient relationship, under the premise of the pixel value of X, is The gradient of is constrained to solve the pixel value. The same method is used to solve Find the appropriate pixel value. Take the average of the gradient results obtained from the two and use it as the pixel value of the abnormal area.
[0063] Step 2: The artificial abnormal image and normal image are passed through the pre-trained feature extractor, the deep features are input into the twin flow model 1, and the deep and shallow fused features are input into the twin flow model 2.
[0064] Step 2.1: Apply a pre-trained open-source feature extractor f ea and multi-layer perceptron feedforward network (MLFF) to obtain the fused deep non-defective and defective features, namely H deep 、 Among them H deep 、 Directly used as input to flow model 2.
[0065] Step 2.2: Using shallow features H shallow 、 and deep features H deep、 Combine projection and interpolation to obtain the input of flow model 1.
[0066]
[0067] Among them, LP(·), C(·), Intp(·) and Conv(·) represent linear projection, concatenation, interpolation and convolution layers respectively. The intuitive expression of the processing and fusion of this part of features is as follows Figure 3 shown.
[0068] Step 3: Upsample the output normal feature distribution estimates of flow model 1 and flow model 2 respectively, and fuse the feature estimation maps of the two outputs by taking the mean.
[0069] Step 3.1: A multi-scale twin flow (MSSF) is proposed to estimate the probability and obtain a reversible multi-step twin flow f sia ={f 1 :H shallow →Z shallow ;f 2 :H deep →Z deep Each flow f consists of K flow blocks {f1,…,f K}, the unknown normal image distribution p H Bidirectional mapping to an accessible latent space distribution p which is usually set to a Gaussian distribution Z The reversible process is as follows:
[0070]
[0071] Step 3.2: Under differential expression, the formula for the probability estimation process is as follows
[0072]
[0073] Here, |·| represents the computation of the determinant. After replacing it with the flow f during the mapping process, the expression can be written as follows.
[0074]
[0075] Each step f of the flow model f i They can be expressed using the following expressions:
[0076]
[0077] If it is step 1, enter H in For H = {H deep ,H shallow},as well as As for the last step of the flow, that is, the Kth step, the corresponding original sample log-likelihood atlas H out yes It contains both defect-free samples and artificial abnormal samples from mixed sources. In addition, due to the step-by-step estimation of the Jacobian matrix H and Z can be converted to each other in each step. Subnet represents a 3×3 convolution block Conv 3×3 and activation function ReLU. in,1 Corresponding to H in The initial half of the features, and H in,2 Corresponding to the remaining part. in,1,1 and H′ in,1,2 C(·), s(·), and t(·) denote the concatenation, scaling, and exponential functions, respectively.
[0078] Step 4: Estimate the negative log-likelihood distribution of normal samples. Using the fused feature estimate distribution, compare the self-supervised estimate with the ground-truth anomaly segmentation map. Calculate the defect localization loss and use it as an adversarial term for the negative log-likelihood estimate of normal samples. If the hybrid loss meets the requirements at the current step, continue training until the required accuracy is achieved.
[0079] Step 4.1: The basis of the flow-based method is to optimize the original defect-free image feature H = {H shallow ,H deep The distribution of} H (h), to get as close as possible to the non-defective zero map. Therefore, the ideal result is to make p H (h) is close to zero. Twin flow structure f sia ={f 1 ,f 2 The loss of} can be obtained by negative log likelihood -log(p H (h)) as follows:
[0080]
[0081] Step 4.2: Mapping Results Jacobian matrix By limiting the upper and lower bounds of tanh(s(H′ in,1,1 )) is estimated, where tanh is the hyperbolic tangent function.
[0082]
[0083] Step 4.3: Artificial anomalies seem to be redundant in the flow model and hinder the correct inference of distribution evaluation. In fact, we design this part as a new type of loss that hinders the training of the flow model, the anomaly localization pixel average L1 loss L mask They are listed below:
[0084]
[0085] where Y norm and are the mapping of zero anomalies for input normal samples and synthetic defect mapping, respectively.
[0086] Step 4.4: Final training loss L tr The adversarial approach of negative log-likelihood and segmentation loss is used as follows:
[0087] L tr =L neg -λ·epoch·L mask
[0088] where λ is an influence L mask Parameter for constraint capability. A larger λ indicates stronger adversarial behavior. epoch indicates the current training round number.
[0089] Example 1:
[0090] The present invention can be applied to any occasion in the product image defect detection stage, and can be applied to multi-category defect detection occasions including wooden floors and leather products. Figure 1 、 2 , 3, 4, and 5 are used as examples.
[0091] Most existing flow models are targeted at detecting defects in a single category of products, and only a few studies have explored optimizing basic properties of the model, such as the number of parameters and inference speed. Few existing flow methods incorporate self-supervised defect segmentation, and there are also few methods and devices that can effectively combine corresponding artificial defects to improve the estimation results of normal distribution. Figure 3 As shown in Figure 2, generating realistic artificial defect samples includes the following steps:
[0092] Step 1.1: First, generate a set of masks containing rectangular shapes for the image dataset of wooden floors and leather products. Among them, at the boundary between the normal area and the abnormal area, the width of each rectangle (w i ) and height (h i ) follows the Gamma distribution, that is, w i ~Γ(α w ,β w ) and h i ~Γ(αh ,β h ), where i = 1,…,k.
[0093] Step 1.2: With random rotation and flipping Perform element-by-element dot multiplication to get the output
[0094] Step 1.3: Place X and Perform seamless combination based on Poisson editing, and finally obtain The steps of Poisson editing are as follows:
[0095] (1) Randomly reduce the original image X locally, and according to According to the pixel value situation, the color of the corresponding area is changed by enhancing, attenuating, and replacing the RGB color channel gradient.
[0096] (2) Set the upper, lower, left, and right boundary pixels of the new image block to X and The boundary pixels and half of , and solve the gradient fields of the two separately.
[0097] (3) Smooth the image. Combine X and The gradient relationship, under the premise of the pixel value of X, is The gradient of is constrained to solve the pixel value. The same method is used to solve Find the appropriate pixel value. Take the average of the gradient results obtained from the two and use it as the pixel value of the abnormal area.
[0098] For example Figure 4 Compared with the general image enhancement methods in other flow models shown above, we combine self-supervised defects with normal data to fully exploit the characteristics of the data itself, thereby improving the effect and ability of defect detection and localization.
[0099] Example 2:
[0100] Compared with simple image preprocessing and unsupervised flow models (such as Figure 1 ), the flow model method we invented can make full use of the characteristics of the target image information and make full use of the multi-level feature content, improving the accuracy and effect of distribution estimation. Based on Example 1, training the twin flow model to complete the distribution estimation of the normal image and the original normal image distribution estimation specifically includes the following steps:
[0101] Step 2.1: Apply a pre-trained open-source feature extractor f eaand multi-layer perceptron feedforward network (MLFF) to obtain the fused deep non-defective and defective features, namely H deep 、 Among them H deep 、 Directly used as input to flow model 2.
[0102] Step 2.2: Using shallow features H shallow 、 and deep features H deep 、 Combine projection and interpolation to obtain the input of flow model 1.
[0103]
[0104] Among them, LP(·), C(·), Intp(·) and Conv(·) represent linear projection, concatenation, interpolation and convolution layers respectively. The intuitive expression of the processing and fusion of this part of features is as follows Figure 3 shown.
[0105] Step 2.3: Combine the multi-scale twin flow (MSSF) to estimate the probability and obtain the reversible multi-step twin flow f sia ={f 1 :H shallow →Z shallow ;f 2 :H deep →Z deep Each flow f consists of K flow blocks {f1,…,f K}, bidirectionally mapping the unknown normal image distribution pH to an accessible latent space distribution p which is usually set to a Gaussian distribution Z .
[0106] Step 2.4: Under differential expression, the formula for the probability estimation process is as follows
[0107]
[0108] Here, |·| represents the computation of the determinant. After replacing it with the flow f during the mapping process, the expression can be written as follows.
[0109]
[0110] Each step f of the flow model f i They can be expressed using the following expressions:
[0111]
[0112] If it is step 1, enter H in For H = {Hdeep ,H shallow},as well as As for the last step of the flow, that is, the Kth step, the corresponding original sample log-likelihood atlas H out yes It contains both defect-free samples and artificial abnormal samples from mixed sources. In addition, due to the step-by-step estimation of the Jacobian matrix H and Z can be converted to each other in each step. Subnet represents a 3×3 convolution block Conv 3×3 and activation function ReLU. in,1 Corresponding to H in The initial half of the features, and H in,2 Corresponding to the remaining part. in,1,1 and H′ in,1,2 C(·), s(·), and t(·) denote the concatenation, scaling, and exponential functions, respectively.
[0113] Step 2.5: The basis of the flow-based method is to optimize the original defect-free image feature H = {H shallow ,H deep The distribution of} H (h), to get as close as possible to the non-defective zero map. Therefore, the ideal result is to make p H (h) is close to zero. Twin flow structure f sia ={f 1 ,f 2 The loss of} can be obtained by negative log likelihood -log(p H (h)) as follows:
[0114]
[0115] Step 2.6: Mapping the results Jacobian matrix By limiting the upper and lower bounds of tanh(s(H′ in,1,1 )) is estimated, where tanh is the hyperbolic tangent function.
[0116]
[0117] Step 2.7: Artificial anomalies seem to be redundant in the flow model and hinder the correct inference of distribution evaluation. In fact, we design this part as a new type of loss that hinders the training of the flow model, the anomaly localization pixel average L1 loss L mask They are listed below:
[0118]
[0119] where Ynorm and are the mapping of zero anomalies for input normal samples and synthetic defect mapping, respectively.
[0120] Step 2.8: Final training loss L tr The adversarial approach of negative log-likelihood and segmentation loss is used as follows:
[0121] L tr =L neg -λ·epoch·L mask
[0122] where λ is an influence L mask Parameter for constraint capability. A larger λ indicates stronger adversarial behavior. epoch indicates the current training round number.
[0123] Example 3:
[0124] Based on Example 1 and Example 2, the present invention can also be used in the estimation of the distribution of multi-category targets and defect detection conditions in the actual defect detection pipeline. Figure 1 and attached Figure 2 As shown, the present invention effectively combines the self-supervisory method with the twin flow model, improving the model's feature synthesis and distribution estimation capabilities of the original detection target. The specific steps of defect detection in the production process are as follows:
[0125] Step 1: Receive product images from multiple production lines on the production site using an external image acquisition device, and obtain images that meet the model size and format requirements by scaling.
[0126] Step 2: The pre-processed image is passed to the trained twin stream model with fixed parameters to complete the prediction and segmentation of the defective part of the captured image;
[0127] Step 3: Obtain the segmentation map and scores of defects and detection locations, and complete the defect detection and statistical analysis of multiple categories of products in this batch;
[0128] Step 4: Wait for the subsequent batch of products to be input and return to step 1.
[0129] The actual detection process has verified that, under the same hardware configuration, the gap in detection indicators between the twin flow model proposed in the present invention and the cutting-edge multi-category defect detection method is controlled within 3.5%, the number of trainable parameters of the model is reduced to a quarter of that of the cutting-edge model, and the inference speed of the model is accelerated to more than three times that of the cutting-edge detection framework.
[0130] The beneficial results of the present invention are mainly manifested in: combining the flow model's ability to efficiently evaluate whether the image to be tested is abnormal or not, adopting a detection framework that takes into account both detection accuracy and detection speed, while controlling the number of parameters, through self-supervision methods, making full use of artificial simulated samples to significantly reduce the number of trainable parameters and improve the speed of the inference process.
[0131] Corresponding to the aforementioned method for self-supervised multi-category industrial product image defect detection based on a twin flow model, an embodiment of a device for self-supervised multi-category industrial product image defect detection based on a twin flow model is provided, as shown in the attached Figure 5 As shown. The system comprises a memory and one or more processors, and the memory contains all executable code, preprocessed images, and representative features. When the processor executes the executable code, it implements the training or inference of the self-supervised multi-category industrial product image defect detection method based on the twin flow model.
[0132] The embodiment of the self-supervised multi-category industrial product image defect detection method based on the twin flow model provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, if Figure 5 As shown, the present invention provides a hardware structure diagram of any device with data processing capability for a self-supervised multi-category industrial product image defect detection device based on a twin flow model, except Figure 5 In addition to the storage unit, visual unit, computing unit, network interface, multi-pipeline image acquisition device and fixed memory shown, any device with data processing capabilities in which the device in the embodiment is located may also include other hardware based on the actual function of the device with data processing capabilities, which will not be described in detail.
[0133] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0134] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement the present invention without inventive work.
[0135] An embodiment of the present invention also provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, it implements a self-supervised multi-category industrial product image defect detection method based on a twin flow model in the above embodiment. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the above embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, smart memory card (Smart Media Card), SD card, flash card (Flash Card), graphics card (Graphics Card), etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.
[0136] It will be understood that the present invention is described by way of some embodiments, and it will be appreciated by those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are intended to be protected by the present invention.
Claims
1. A self-supervised multi-category industrial product image defect detection method based on a twin flow model, characterized by: The steps include: Step 1: Randomly sample from a multi-category dataset of normal samples, randomly cut and flip another part of the normal samples, and combine them with self-supervised Poisson editing to generate realistic artificial defect samples; Step 2: The artificial abnormal image and the normal image are passed through a pre-trained feature extractor, and the deep-shallow fusion features and deep features are respectively input into the twin stream model. The deep features and shallow features are fused through upsampling, convolution, concatenation, and linear mapping alignment. The result is used as the input of stream model 1, and the result of the deep features is used as the input of stream model 2. Step 3: Upsample the output normal feature distribution estimates of flow model 1 and flow model 2 respectively, and fuse the feature estimation maps of the two outputs by taking the mean; Step 4: Estimate the negative logarithm distribution of normal samples. Use the fused feature estimation distribution map to compare the estimation of the self-supervised process with the true anomaly segmentation map. Calculate the loss of defect localization and use it as the adversarial term for the negative logarithm estimation of normal samples. If the mixed loss meets the requirements at the current step, continue training until the required computational accuracy is met.
2. The self-supervised multi-category industrial product image defect detection method based on the twin flow model according to claim 1, It is characterized in that The process of generating realistic artificial defect samples in step 1 includes the following steps: Step 1.1: First generate a mask set containing rectangular shapes Among them, at the boundary between the normal area and the abnormal area, the width w of each rectangle is i ) and height h i ) follows the Gamma distribution, that is, w i ~Γ(α w ,β w ) and h i ~Γ(α h ,β h ), where i = 1,…,k; Step 1.2: With random rotation and flipping Perform element-by-element dot multiplication to get the output Step 1.3: Place X and Perform seamless combination based on Poisson editing, and finally obtain 3. The self-supervised multi-category industrial product image defect detection method based on the twin flow model according to claim 2 is characterized in that: The steps of Poisson editing are as follows: S1: Randomly reduce the original image X locally, and according to According to the pixel value, the color of the corresponding area is changed by enhancing, attenuating, and replacing the RGB color channel gradient; S2: Set the upper, lower, left, and right boundary pixels of the new image block to X and The boundary pixels and half of , and solve the gradient fields of the two separately; S3: Smoothing the image: Combining X and The gradient relationship, under the premise of the pixel value of X, is The gradient is constrained to solve the pixel value and the same method is used to solve the pixel value. Solve for the appropriate pixel value and average the gradient results obtained from the two as the pixel value of the abnormal area.
4. The self-supervised multi-category industrial product image defect detection method based on the twin flow model according to claim 1 is characterized in that: The step 2 includes the following specific steps: Step 2.1: Apply a pre-trained open-source feature extractor f ea and multi-layer perceptron feedforward network MLFF to obtain the fused deep non-defect and defect features, namely H deep 、 Among them H deep 、 Directly used as input to flow model 2; Step 2.2: Using shallow features H shallow 、 and deep features H deep 、 Combining projection and interpolation, we get the input of flow model 1: where LP(·), C(·), Intp(·), and Conv(·) denote linear projection, concatenation, interpolation, and convolution layers, respectively.
5. The self-supervised multi-category industrial product image defect detection method based on the twin flow model according to claim 1 is characterized in that: The step 3 includes the following specific steps: Step 3.1: A multi-scale twin flow MSSF is proposed to estimate the probability and obtain a reversible multi-step twin flow f sia ={f 1 :H shallow →z shallow ;f 2 :H deep →Z deep }, each flow f consists of K flow blocks {f1,…,f K }, the unknown normal image distribution p H Bidirectional mapping to an accessible latent space distribution p which is usually set to a Gaussian distribution Z , the reversible process is as follows: Step 3.2: Under differential expression, the formula for the probability estimation process is as follows: Here, |·| represents the calculation of the determinant. After replacing it with the flow f in the mapping process, the expression is written as follows: Each step f of the flow model f i They can be expressed by the following expressions: If it is step 1, enter H in For H = {H deep ,H shallow },as well as As for the last step of the flow, that is, the Kth step, the corresponding original sample log-likelihood atlas H out yes It contains both defect-free samples and artificial anomaly samples from mixed sources. In addition, due to the step-by-step estimation of the Jacobian matrix H and Z can be converted to each other in each step, where Subnet represents a 3×3 convolution block Conv 3×3 and activation function ReLU, H in,1 Corresponding to H in The initial half of the features, and H in,2 Corresponding to the remaining part, H′ in,1,1 and H′ in,1,2 denotes similar definitions, with C(·), s(·), and t(·) representing the concatenation, scaling, and exponential functions, respectively.
6. The self-supervised multi-category industrial product image defect detection method based on the twin flow model according to claim 1 is characterized in that: The step 4 includes the following specific steps: Step 4.1: The basis of the flow-based method is to optimize the original defect-free image feature H = {H shallow ,H deep The distribution of} H (h), to be as close as possible to the non-defective zero mapping, so the ideal result is to make p H (h) close to zero, twin flow structure f sia ={f 1 ,f 2 The loss of} is calculated by negative log-likelihood -log(p H (h)) as follows: Step 4.2: Mapping result Z~N(0,1), Jacobian matrix By limiting the upper and lower bounds of tanh(s(H′ in,1,1 )) is estimated, where tanh is the hyperbolic tangent function, Step 4.3: Design artificial anomalies in the flow model as a new type of loss that hinders the training of the flow model, and the average L1 loss L of the abnormal positioning pixels mask They are listed below: where Y norm and are the mappings of zero anomalies for the input normal samples and synthetic defect mappings, respectively; Step 4.4: Final training loss L tr The adversarial approach of negative log-likelihood and segmentation loss is used as follows: L tr =L neg -λ·epoch·L mask where λ is an influence L mask Parameters of constraint ability, a larger λ indicates stronger adversarial behavior, and epoch indicates the current training round number.
7. A device for self-supervised multi-category industrial product image defect detection based on a twin flow model, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the self-supervised multi-category industrial product image defect detection method based on the twin flow model described in any one of claims 1 to 6.