Industrial product surface defect detection method based on multi-scale feature reconstruction of variational autoencoder
Patent Information
- Application Number
- CN202611257144.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
这是分类器方法的根本性局限,即决策边界的封闭性导致其对未知缺陷的泛化能力为零
[0023](1)彻底消除对缺陷(NG)样本的依赖。本发明在训练阶段完全不需要任何缺陷样本,仅使用正常(OK)产品图像即可完成模型训练。这从根本上解决了工业检测领域NG样本获取难、获取贵、分布不均的核心痛点。对于新产品线或新型号,只需采集正常产品的图像即可快速部署检测模型,部署周期从传统方法的数周甚至数月缩短到数天。
Smart Images

Figure CN122820698A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of automatic detection technology for surface defects in industrial products, and relates to a method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder. Background Technology
[0002] With the development of industrial automation and intelligent manufacturing, machine vision-based surface defect detection of industrial products has become a key link in quality control and is widely used in the quality inspection of various industrial products such as electronic components, metal parts, textiles, and glass panels.
[0003] Currently, a representative technical approach in the field of industrial defect detection is as follows:
[0004] First, a multi-granularity scale (MGS) feature extraction module is used to extract visual feature vectors at multiple granularity levels (coarse-grained, medium-grained, and fine-grained) from the product surface image. After pooling, these vectors are concatenated into a high-dimensional multi-scale feature vector. Then, this feature vector is input into a multi-grained cascade forest (gcForest) classifier, which outputs a binary classification result of "OK" or "NG". The core technology of this scheme lies in the fact that the feature extraction stage utilizes MGS to capture image feature information at different scales, while the classification stage relies on the cascaded structure of the deep forest to perform layer-by-layer feature transformation and the final binary classification decision.
[0005] This technical approach has a certain degree of rationality in terms of feature extraction, as features of different scales can capture large-area defects (such as scratches and stains) and small defects (such as pinholes and pits) respectively. However, its classification and determination process is based on a supervised learning paradigm, which requires the prior collection of a large number of labeled OK and NG samples to train the classifier.
[0006] The aforementioned technical approach of "MGS feature extraction - gcForest classifier - OK / NG binary classification" faces the following core problems in actual industrial production line deployment:
[0007] (1) Incomplete defect coverage, new defects cannot be detected. The classifier needs to prepare sufficient NG (no good) samples of various types for training in advance. The essence of its learning is the "decision boundary between OK samples and known NG samples". When a completely new defect type appears on the production line that has never appeared in the training set, the model will not be able to identify these new defects because the decision boundary of the classifier only covers the known defect type space, and will directly miss them as OK (qualified). This is the fundamental limitation of the classifier method, that is, the closed nature of the decision boundary leads to its zero generalization ability to unknown defects.
[0008] (2) Obtaining NG sample data is difficult and the deployment cycle is long. In actual industrial production, the number of NG samples is scarce and the distribution of types is severely uneven. Some rare defects (such as latent cracks under specific process conditions) may require several months or even longer to collect enough samples, resulting in extremely high collection costs and long cycles. This data acquisition bottleneck severely restricts the rapid deployment of classifier-based detection methods on new models and new production lines.
[0009] (3) Limited output information. Existing solutions can only output OK / NG binary judgment results, and cannot provide fine-grained information such as the severity of the defect (e.g., minor, moderate, severe), the specific spatial location of the anomaly, and the anomalous components of each feature scale. This coarse-grained output method of "black and white" is difficult to meet the needs of modern industrial quality inspection for defect traceability and graded processing (e.g., downgrading minor defects for use and directly scrapping severe defects).
[0010] (4) The model's decision boundary is fixed and cannot adapt to production line drift. After the classifier is trained, its decision boundary remains fixed. However, the process conditions (such as temperature and humidity), lighting environment, and raw material batches in industrial production lines may all experience slow drift within the normal range. This drift will cause the classifier's misclassification rate to gradually increase, and the model itself cannot adapt and adjust. It is necessary to collect data again and retrain, and each adjustment will return to the dilemma of "needing NG samples".
[0011] (5) Poor model interpretability. As an ensemble learning model, deep forest involves a multi-layered cascade structure and a voting mechanism of a large number of decision trees, which is a "black box" decision-making process for users. When misjudgment occurs, it is difficult to trace the specific basis of the model's judgment (which feature or scale led to the judgment result), which cannot meet the compliance requirements of industrial quality inspection for traceability and auditability.
[0012] The root cause of the above five problems lies in the fact that the essence of classifier learning is "drawing a dividing line (decision boundary) between OK and NG". Once the NG type space expands (new defect types appear), the original decision boundary is no longer effective. Therefore, there is an urgent need for a new detection paradigm that does not rely on prior knowledge of NG samples and can make anomaly judgments based solely on "what normal products look like". Summary of the Invention
[0013] To address the aforementioned technical problems, the purpose of this invention is to provide a method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder.
[0014] The present invention provides a method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder, comprising:
[0015] Step 1: Acquire the surface image of the industrial product to be inspected, and perform image scaling, illumination correction, noise removal, and data type conversion preprocessing in sequence;
[0016] Step 2: Using a multi-granularity, multi-scale feature extraction network, extract feature vectors of three scales—coarse-grained, medium-grained, and fine-grained—from the preprocessed image, and then concatenate them after pooling to form the original multi-scale feature vector.
[0017] Step 3: Input the original multi-scale feature vector into the encoder part of the pre-trained variational autoencoder (VAE), compress it step by step through a multi-layer fully connected network, output the mean and log-variance of the latent space, and obtain the latent vector through reparameterization techniques.
[0018] Step 4: Input the latent vector into the decoder part of the pre-trained variational autoencoder (VAE), and restore it step by step through a multi-layer fully connected network to output the reconstructed multi-scale feature vector.
[0019] Step 5: Decompose the reconstructed multi-scale feature vector into three scales, and independently calculate the reconstruction error between the original features and the reconstructed features at each scale; adaptively weight and fuse the reconstruction errors at each scale using a set of learnable weight parameters to obtain the comprehensive reconstruction error;
[0020] Step 6: Map the overall reconstruction error to an anomaly score between 0 and 1 using the Sigmoid activation function, determine normal and defective based on the preset judgment threshold, and output the judgment result, anomaly score, reconstruction error at each scale, and anomaly level.
[0021] Step 7: Collect M images that are judged to be normal to form an image batch, and feed them into the pre-trained variational autoencoder (VAE). After obtaining the reconstructed multi-scale feature vector, calculate the anomaly score. If the anomaly scores of N consecutive batches of normal images all fall outside the confidence interval, then update the variational autoencoder (VAE).
[0022] The present invention provides a method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder, which has the following beneficial effects:
[0023] (1) Completely eliminates dependence on defective (NG) samples. This invention requires no defective samples at all during the training phase, and can complete model training using only normal (OK) product images. This fundamentally solves the core pain points of difficult, expensive, and unevenly distributed NG samples in the industrial inspection field. For new product lines or new models, only images of normal products need to be collected to quickly deploy the inspection model, shortening the deployment cycle from several weeks or even months in traditional methods to several days.
[0024] (2) It has a natural ability to detect unknown defects. Since the variational autoencoder (VAE) has never learned what defects look like, its detection mechanism is based on the principle that deviations from the normal are abnormal. Therefore, it is also effective for novel unknown defects that have never appeared during training. In comparative experiments, for known defect types, the detection rates of both the present invention and the gcForest classification scheme are above 95%; for the three completely new unknown defects, the detection rate of the gcForest scheme drops sharply to below 40%, while the detection rate of the present invention remains above 90%. This advantage is particularly prominent in the early production stages when the process is unstable and the defect types are varied.
[0025] (3) Provides rich, multi-dimensional detection outputs. Compared with the single OK / NG binary output of traditional methods, this invention outputs anomaly scores (0~1, which can finely distinguish the degree of anomaly), independent reconstruction errors at each scale, and anomaly degree levels. This rich information provides a data foundation for the digital transformation of industrial quality inspection (such as defect trend analysis, production line health monitoring, and intelligent grading processing).
[0026] (4) The model has adaptive evolution capability. When the production line process conditions drift normally, only a small number of new OK samples and low learning rate fine-tuning are needed to complete the model update, without the need to collect NG samples again. This feature enables the model to maintain high detection accuracy throughout the entire production line life cycle, avoiding the periodic maintenance dilemma of traditional methods where the model accuracy slowly declines and requires re-collecting data for training.
[0027] (5) Strong model interpretability. Compared with the black-box ensemble decision-making mechanism of deep forests, the VAE reconstruction detection principle of this invention is intuitive and transparent: large reconstruction error → anomaly. The independent output of error components at each scale further enhances the interpretability of the results—quality inspectors can clearly know which scale's anomaly led to the NG judgment. This interpretability meets the compliance requirements of industrial quality inspection for traceability and auditability, and also facilitates quality inspectors' understanding and trust in the model's judgment results.
[0028] (6) High computational efficiency, suitable for real-time deployment on production lines. During online detection, the VAE model only needs to perform one forward propagation (encoding + decoding), without the need for iterative optimization or complex post-processing. The inference time for a single image is usually in the millisecond range. The symmetrical encoding and decoding network structure is simple to design, with a moderate total number of parameters, and can be easily deployed on edge computing devices in industrial production lines to meet the real-time detection needs of high-speed production lines. Attached Figure Description
[0029] Figure 1 This is a flowchart of the industrial product surface defect detection method based on multi-scale feature reconstruction using a variational autoencoder, according to the present invention. Detailed Implementation
[0030] This invention provides a method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a Variational Autoencoder (VAE). This method fundamentally changes the technological paradigm of defect detection—shifting from "classification and discrimination" to "reconstruction detection": during the training phase, only OK (normal) samples are used to train the model, enabling it to compress multi-scale features into a low-dimensional latent space and accurately reconstruct (restore) them; during the inference phase, the model's reconstruction error of the input features (i.e., the difference between the input and the reconstructed version) is directly used as a measure of the degree of anomaly, achieving defect detection without requiring out-of-the-box (NG) samples for training.
[0031] like Figure 1 As shown, the industrial product surface defect detection method based on multi-scale feature reconstruction using a variational autoencoder according to the present invention includes:
[0032] Step 1: Acquire a surface image of the industrial product to be inspected, and perform image scaling, illumination correction, noise removal, and data type conversion preprocessing in sequence to ensure consistent input quality for subsequent processing. The preprocessing specifically includes:
[0033] Image scaling: Bilinear or bicubic interpolation algorithms are used to uniformly scale surface images of industrial products of different resolutions and sizes to a fixed size. This fixed size ensures consistent input dimensions for subsequent multi-granularity, multi-scale feature extraction networks.
[0034] Illumination Correction Processing: Due to potential fluctuations in lighting conditions at industrial sites (such as lamp aging and changes in ambient light), image illumination normalization is necessary. Specifically, the mean grayscale value of all pixels in the image is calculated. Each pixel value is then subtracted from this mean and divided by the standard deviation, resulting in an image with zero mean and unit variance. This operation effectively eliminates the impact of global brightness shifts on subsequent feature extraction.
[0035] Noise Removal: Industrial cameras may introduce sensor noise during high-speed acquisition. Gaussian or median filtering is used to perform mild noise reduction on the image, suppressing random noise while preserving defect details.
[0036] Data type conversion processing: Convert the pixel values of the image into floating-point tensors, and normalize the numerical range to the interval [0,1] or [-1, 1].
[0037] Step 2: Using a Multi-Granularity Scales (MGS) feature extraction network, coarse-grained, medium-grained, and fine-grained feature vectors are extracted from the preprocessed image. After pooling, these vectors are concatenated to form the original multi-scale feature vector. Specifically:
[0038] Step 2.1: Select a convolutional neural network, ResNet or VGG, pre-trained on the ImageNet large-scale image dataset as the backbone network for feature extraction. This preserves defect details while suppressing random noise.
[0039] Step 2.2: Define three feature extraction scales: coarse-grained, medium-grained, and fine-grained, and perform feature extraction operations at these three scales.
[0040] Coarse-grained feature extraction: Feature maps are extracted from the shallow or mid-shallow layers of the network and compressed into an a-dimensional coarse-grained feature vector through global average pooling. Shallow features retain more spatial structure information, have a smaller receptive field but higher feature map resolution, and are suitable for capturing the overall outline of the product, large-scale textures and structural changes.
[0041] Medium-granularity feature extraction: Feature maps are extracted from the intermediate layers of the network and then subjected to global average pooling to obtain b-dimensional medium-granularity feature vectors. The intermediate layer features have a medium receptive field, which can capture local structural patterns and medium-scale texture changes.
[0042] Fine-grained feature extraction: Feature maps are extracted from the deep layers of the network and then subjected to global average pooling to obtain c-dimensional fine-grained feature vectors. Deep features have the largest receptive field and the richest semantic information, and are highly sensitive to minute details and local anomalies in the image.
[0043] Step 2.3: Concatenate the feature vectors from the three scales together to form the original multi-scale feature vector with a total dimension of d, where d = a + b + c. The original multi-scale feature vector contains visual information at three levels, from coarse to fine, and serves as the input for the subsequent VAE.
[0044] Step 3: Input the original multi-scale feature vector into the encoder part of the pre-trained variational autoencoder (VAE), compress it step by step through a multi-layer fully connected network, output the mean and log-variance of the latent space, and obtain the latent vector through reparameterization techniques. This completes the condensed representation of high-dimensional features.
[0045] The encoder portion of the variational autoencoder (VAE) includes: an input layer, a first hidden layer, a second hidden layer, a latent spatial parameter layer, and a reparameterized sampling layer.
[0046] Input layer: Receives the d-dimensional multi-scale feature vector output from step 2.
[0047] The first hidden layer maps the input d-dimensional multi-scale feature vector to a c-dimensional space through a fully connected layer, and then processes it sequentially through a batch normalization layer and a LeakyReLU activation function.
[0048] The second hidden layer further compresses the c-dimensional multi-scale feature vector to b-dimensionality, and also processes it through a batch normalization layer and the LeakyReLU activation function.
[0049] The purpose of the batch normalization layer is to normalize the output of this layer for each mini-batch of data, making its mean 0 and variance 1. This can accelerate model convergence, stabilize the training process, and allow for the use of a larger learning rate.
[0050] The purpose of the LeakyReLU activation function is to prevent neurons from "dying" during training (i.e., output always being zero, gradient always being zero, and parameters no longer being updated) from having a small slope (e.g., 0.01) on the negative half-axis, unlike the standard ReLU which sets the negative half-axis to zero. This is crucial for the stable training of deep networks.
[0051] Latent Space Parameter Layer: The b-dimensional multi-scale feature vectors are mapped to two a-dimensional vectors through two independent fully connected layers, namely the mean vector μ of the b-dimensional multi-scale feature vectors and the log-variance vector logσ² of the b-dimensional multi-scale feature vectors, which together define the Gaussian distribution in the latent space; μ describes the central location of the feature, and σ² describes the reasonable range of variation of the feature.
[0052] Reparameterized sampling layer: Direct random sampling from the distributions defined by μ and σ is non-differentiable (backpropagation is impossible), therefore a reparameterization technique is employed. Specifically, a random noise vector ε is sampled from the standard normal distribution N(0, 1), and the a-dimensional latent vector is calculated using the following formula:
[0053]
[0054] Where z is the potential vector.
[0055] This operation shifts randomness from the sampling process to the noise variable ε, making the entire computation process differentiable, thus supporting end-to-end gradient backpropagation training.
[0056] Step 4: Input the latent vector into the decoder part of the pre-trained variational autoencoder (VAE), and reconstruct it step by step through a multi-layer fully connected network to output the reconstructed multi-scale feature vector. This completes the decompression and reconstruction of the features.
[0057] The decoder section includes an input layer, a third hidden layer, a fourth hidden layer, and an output layer.
[0058] Input layer: Receives the a-dimensional latent vector z from the encoder part of the output.
[0059] The third hidden layer maps the a-dimensional latent vector to a b-dimensional space through fully connected layers, batch normalization, and the LeakyReLU activation function.
[0060] The fourth hidden layer maps the b-dimensional latent vectors to a c-dimensional space through fully connected layers, batch normalization, and the LeakyReLU activation function.
[0061] Output layer: Maps the c-dimensional latent vector back to d-dimensional through a fully connected layer, and outputs the reconstructed multi-scale feature vector.
[0062] The decoder does not directly reconstruct the original image, but instead reconstructs the MGS multi-scale feature vector. The advantage of doing so is that the feature space is more compact and semantic than the pixel space, and the VAE can focus on learning "what the features of a normal product should be like", without having to learn pixel-level image generation (the latter is more difficult and easily introduces irrelevant image details).
[0063] Step 5: Decompose the reconstructed multi-scale feature vector into three scales, and independently calculate the reconstruction error between the original features and the reconstructed features at each scale; adaptively weight and fuse the reconstruction errors at each scale using a set of learnable weight parameters to obtain the comprehensive reconstruction error.
[0064] This step is the core innovation of this invention. Traditional reconstruction error calculation methods calculate a uniform difference value for the entire spliced feature vector. However, this method ignores the heterogeneity of features at different scales—the numerical ranges and importance of coarse-grained features (256 dimensions) and fine-grained features (1280 dimensions) may be completely different, and calculating them together will mask subtle anomalous signals at the fine-grained scale. Specifically:
[0065] Step 5.1: Scale-based splitting: The reconstructed multi-scale feature vector output from the decoder is split into coarse-grained feature components x in the first a dimensions according to three scales. r1 The intermediate b-dimensional medium-granularity feature component x r2 Finally, the c-dimensional fine-grained feature components x r3 .
[0066] Step 5.2: Independently calculate the reconstruction error at each scale: Calculate the a-dimensional coarse-grained feature vector x1 and the a-dimensional coarse-grained feature component x. r1 The squared Euclidean distance between them scale1 As a coarse-grained error component; calculate the b-dimensional medium-grained eigenvector x2 and the b-dimensional medium-grained eigencomponent x. r2 The squared Euclidean distance between them scale2 As a medium-grained error component; calculate the c-dimensional coarse-grained eigenvector x3 and the c-dimensional coarse-grained eigencomponent x. r3 The squared Euclidean distance between themscale3 As a component of fine granularity error.
[0067] coarse-grained error component E scale1 This primarily reflects the deviation in the reconstruction of the product's overall structure and large-scale texture. This error component increases significantly when the product has large-area scratches, stains, or structural deformation.
[0068] Medium particle size error component E scale2 This reflects the reconstruction deviation of medium-scale patterns and local structures. This error component increases when local structural anomalies occur in the product (such as regional color difference or local deformation).
[0069] Fine particle size error component E scale3 This reflects the reconstruction deviation of minute details and local textures. When a product has minor defects such as pinholes, pits, or microcracks, this error component will increase significantly because fine-grained features are most sensitive to these minute anomalies.
[0070] Step 5.3: Adaptively fuse the error components of the three scales using learnable weights to obtain the comprehensive reconstruction error:
[0071]
[0072] in, These are learnable weight parameters at a coarse-grained scale. These are learnable weight parameters at a medium granularity scale. β represents the learnable weight parameters at a fine-grained scale, and β is the global bias parameter. , , The four parameters β are used as trainable parameters of the variational autoencoder (VAE), and are automatically learned during training through the backpropagation algorithm.
[0073] The core advantage of learnable weights is that the model automatically identifies which scale to focus on for which type of defect during training. For example, if the main defects on the production line are large-area scratches, then learnable weights at a coarse-grained scale would be more suitable. It will naturally learn to be larger; if the main defect is a tiny pinhole, the learnable weight parameters at a fine-grained scale will be better. They will learn more naturally. The entire process requires no manual pre-setting and is completely data-driven.
[0074] Furthermore, non-negativity constraints can be added to the weight parameters (e.g., through softmax normalization or absolute value operations) to ensure the non-negativity of contributions at each scale, thus improving the physical interpretability of the fusion results. After training convergence, observing the relative magnitudes of the weights can help quality control personnel understand the scale distribution characteristics of the main defect types on the production line.
[0075] In specific implementation, the pre-trained variational autoencoder (VAE) in step 3 or step 4 is trained using a training dataset composed of normal surface image samples of industrial products. The specific training process is as follows:
[0076] (1) Data preparation: Collect sufficient normal surface image samples of industrial products. All normal surface image samples are preprocessed according to step 1, and multi-scale feature vectors are extracted according to step 2 to construct training set and validation set.
[0077] (2) Forward propagation: The multi-scale feature vector X of normal surface image samples in the training set is input into the variational autoencoder (VAE), and passes through the encoder part and the decoder part in sequence to obtain the reconstructed feature vector X. recon .
[0078] (3) Calculate the total loss L total :
[0079]
[0080] in, To reconstruct the loss, For KL divergence loss, These are the weighting coefficients for the KL divergence.
[0081] Reconstruction loss drives the model to reconstruct input features as accurately as possible, serving as the primary driving force for training. KL divergence loss aims to ensure the continuity and smoothness of the latent space—similar inputs correspond to similar latent codes, avoiding gaps or discontinuous regions in the latent space. This is crucial for anomaly detection: only when the latent space is well-structured can anomalous samples produce significant reconstruction bias.
[0082]
[0083] Where S is the total number of scales. This represents the original eigenvector at scale s. Let be the feature components at the s-th scale of the reconstruction.
[0084]
[0085] Where 'a' represents the dimension of the latent variable. This is the mean vector of the predictions from the encoder portion. This represents the standard deviation component of the encoder's predictions.
[0086] (4) Backpropagation and parameter update: The backpropagation algorithm is used to simultaneously update all weights and bias parameters of the encoder part of the VAE, all weights and bias parameters of the decoder part, and multi-scale fusion weight parameters. , , And the global bias parameter β.
[0087] (5) Iterative training: Repeat the forward propagation to parameter update step until the model converges or reaches the preset maximum number of training rounds. When the loss on the validation set no longer decreases for several consecutive rounds, terminate the training in advance to avoid overfitting.
[0088] (6) Threshold calibration: After training, the samples in the training set are fed into the variational autoencoder (VAE), and the outlier scores are calculated and the mean of the outlier scores is calculated. and standard deviation And calculate the judgment threshold according to the following formula:
[0089]
[0090] Where T is the decision threshold and k is an adjustable coefficient.
[0091] k=2 corresponds to a 95% confidence interval, suitable for high recall scenarios. It has a lower threshold, more sensitive anomaly detection, extremely low false negative rate but slightly higher false positive rate, making it suitable for safety-critical scenarios (such as automotive parts inspection).
[0092] k=3 corresponds to a 99.7% confidence interval, suitable for standard, routine scenarios. It balances recall and precision, making it suitable for most common industrial testing scenarios.
[0093] k=4 corresponds to a 99.99% confidence interval, suitable for high-precision scenarios. The threshold is relatively high, and only very significant anomalies are judged as defective (NG) images. The false alarm rate is extremely low, but it may miss minor defects. It is suitable for scenarios with extremely low tolerance for false alarms.
[0094] (7) Mean based on outlier scores and standard deviation Set the following confidence intervals for outlier scores: .
[0095] Step 6: Map the overall reconstruction error to an anomaly score between 0 and 1 using the Sigmoid activation function. Determine whether the error is normal or defective based on a preset threshold, and output the determination result, anomaly score, reconstruction error at each scale, and anomaly severity level. Specifically:
[0096] Step 6.1: Outlier score mapping: The integrated reconstruction error is mapped to a value between 0 and 1 using the Sigmoid activation function, which is used as the final outlier score S.
[0097] The Sigmoid function's characteristic is that it compresses any real number input into the (0, 1) interval, with larger positive values closer to 1 and smaller negative values closer to 0. Therefore, a larger reconstruction error → an anomaly score closer to 1 → more likely a defective product; a smaller reconstruction error → an anomaly score closer to 0 → more likely a normal product. Continuous score design provides much richer information than binary judgment; for example, both 0.35 and 0.92 are judged as NG, but their degrees of anomaly are clearly quite different.
[0098] Step 6.2: Compare the anomaly score with the judgment threshold. At that time, the surface images of the industrial products were normal. At that time, the surface image of the industrial product to be inspected contained defects.
[0099] Step 6.3: Determine the level of abnormality:
[0100] when At that time, it was determined that the surface image of the industrial product to be inspected had a slight anomaly.
[0101] when At that time, it was determined that the surface image of the industrial product to be inspected had a moderate anomaly.
[0102] when At that time, it was determined that the surface image of the industrial product to be inspected had serious anomalies.
[0103] Step 6.4: Output the final judgment result, and also output the anomaly score, reconstruction error at each scale, and anomaly level.
[0104] Step 7: Collect M images (far fewer than the pre-training sample size) that are judged as normal to form an image batch. Feed these images into the pre-trained variational autoencoder (VAE) to obtain the reconstructed multi-scale feature vectors. Then calculate the anomaly score. If the anomaly scores of N consecutive batches of normal images all fall outside the confidence interval, then update the VAE. Specifically:
[0105] Step 7.1: During online detection, M images judged as normal are collected to form an image batch, which is then fed into a pre-trained variational autoencoder (VAE) to obtain reconstructed multi-scale feature vectors before calculating anomaly scores. If the anomaly scores of N consecutive batches of normal images all fall within the confidence interval... In addition, the variational autoencoder (VAE) is updated as follows;
[0106] Step 7.2: Using the parameters of the currently pre-trained variational autoencoder (VAE) as initial values, fine-tune the training using M newly acquired images that are judged to be normal:
[0107] The learning rate and number of training epochs for fine-tuning training are lower than those for pre-training. During fine-tuning training, the parameters of the encoder part are frozen, and only the weights and bias parameters of the decoder part, as well as the multi-scale fusion weight parameters and global bias parameters are updated.
[0108] Step 7.3: After fine-tuning, recalculate the abnormal score distribution using the new normal image and update the judgment threshold.
[0109] The above description is only a preferred embodiment of the present invention and is not intended to limit the ideas of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder, characterized in that, include: Step 1: Acquire the surface image of the industrial product to be inspected, and perform image scaling, illumination correction, noise removal, and data type conversion preprocessing in sequence; Step 2: Using a multi-granularity, multi-scale feature extraction network, extract feature vectors of three scales—coarse-grained, medium-grained, and fine-grained—from the preprocessed image, and then concatenate them after pooling to form the original multi-scale feature vector. Step 3: Input the original multi-scale feature vector into the encoder part of the pre-trained variational autoencoder (VAE), compress it step by step through a multi-layer fully connected network, output the mean and log-variance of the latent space, and obtain the latent vector through reparameterization techniques. Step 4: Input the latent vector into the decoder part of the pre-trained variational autoencoder (VAE), and restore it step by step through a multi-layer fully connected network to output the reconstructed multi-scale feature vector. Step 5: Decompose the reconstructed multi-scale feature vector into three scales, and independently calculate the reconstruction error between the original features and the reconstructed features at each scale; adaptively weight and fuse the reconstruction errors at each scale using a set of learnable weight parameters to obtain the comprehensive reconstruction error; Step 6: Map the overall reconstruction error to an anomaly score between 0 and 1 using the Sigmoid activation function, determine normal and defective based on the preset judgment threshold, and output the judgment result, anomaly score, reconstruction error at each scale, and anomaly level. Step 7: Collect M images that are judged to be normal to form an image batch, and feed them into the pre-trained variational autoencoder (VAE). After obtaining the reconstructed multi-scale feature vector, calculate the anomaly score. If the anomaly scores of N consecutive batches of normal images all fall outside the confidence interval, then update the variational autoencoder (VAE).
2. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 1, characterized in that, The preprocessing in step 1 specifically includes: Image scaling processing: Using bilinear interpolation or bicubic interpolation algorithms, surface images of industrial products with different resolutions and sizes are uniformly scaled to a fixed size; Illumination correction processing: Calculate the mean gray value of the pixels in the entire image, subtract the mean from each pixel value and then divide by the standard deviation to make the processed image have zero mean and unit variance; Noise removal: The image is lightly filtered and denoised using Gaussian filtering or median filtering. Data type conversion processing: Convert the pixel values of the image into floating-point tensors, and normalize the numerical range to the interval [0, 1] or [-1, 1].
3. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 1, characterized in that, Step 2 specifically involves: Step 2.1: Select a convolutional neural network, ResNet or VGG, pre-trained on the ImageNet large-scale image dataset as the feature extraction backbone network; Step 2.2: Define three feature extraction scales: coarse-grained, medium-grained, and fine-grained, and perform feature extraction operations at these three scales; Coarse-grained scale feature extraction: Feature maps are extracted from the shallow or mid-shallow layers of the network and compressed into an a-dimensional coarse-grained feature vector through global average pooling; Medium-granularity feature extraction: Feature maps are extracted from the intermediate layers of the network and obtained as b-dimensional medium-granularity feature vectors after global average pooling; Fine-grained scale feature extraction: Feature maps are extracted from the deep layers of the network and obtained as c-dimensional fine-grained feature vectors after global average pooling; Step 2.3: Concatenate the feature vectors of the above three scales end to end to form the original multi-scale feature vector with a total dimension of d, where d = a + b + c.
4. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 3, characterized in that, The encoder part of the variational autoencoder (VAE) in step 3 includes: an input layer, a first hidden layer, a second hidden layer, a latent spatial parameter layer, and a reparameterized sampling layer. Input layer: Receives the d-dimensional multi-scale feature vector output from step 2; First hidden layer: The input d-dimensional multi-scale feature vector is mapped to c-dimensional space through a fully connected layer, and then processed by batch normalization layer and LeakyReLU activation function in sequence; The second hidden layer further compresses the c-dimensional multi-scale feature vector to b-dimensionality, and also processes it through a batch normalization layer and the LeakyReLU activation function. Latent Space Parameter Layer: The b-dimensional multi-scale feature vectors are mapped to two a-dimensional vectors through two independent fully connected layers, namely the mean vector μ of the b-dimensional multi-scale feature vectors and the log-variance vector logσ² of the b-dimensional multi-scale feature vectors, which together define the Gaussian distribution in the latent space; μ describes the central location of the feature, and σ² describes the reasonable range of variation of the feature. Reparameterized sampling layer: Random noise vector ε is sampled from the standard normal distribution N(0, 1), and the a-dimensional latent vector is calculated using the following formula: Where z is the potential vector.
5. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 4, characterized in that, The decoder part in step 4 includes: an input layer, a third hidden layer, a fourth hidden layer, and an output layer. Input layer: Receives the a-dimensional latent vector z from the encoder portion; The third hidden layer: maps the a-dimensional latent vector to a b-dimensional space through fully connected layers, batch normalization, and the LeakyReLU activation function; The fourth hidden layer maps the b-dimensional latent vectors to a c-dimensional space using a fully connected layer, batch normalization, and the LeakyReLU activation function. Output layer: Maps the c-dimensional latent vector back to d-dimensional through a fully connected layer, and outputs the reconstructed multi-scale feature vector.
6. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 3, characterized in that, Step 5 specifically involves: Step 5.1: Scale-based splitting: The reconstructed multi-scale feature vector output from the decoder is split into coarse-grained feature components x in the first a dimensions according to three scales. r1 The intermediate b-dimensional medium-granularity feature component x r2 Finally, the c-dimensional fine-grained feature components x r3 ; Step 5.2: Independently calculate the reconstruction error at each scale: Calculate the a-dimensional coarse-grained feature vector x1 and the a-dimensional coarse-grained feature component x. r1 The squared Euclidean distance between them scale1 , as a coarse-grained error component; Calculate the b-dimensional medium-granularity eigenvector x2 and the b-dimensional medium-granularity eigencomponent x. r2 The squared Euclidean distance between them scale2 , as a medium-granularity error component; Calculate the c-dimensional coarse-grained eigenvector x3 and the c-dimensional coarse-grained eigencomponent x. r3 The squared Euclidean distance between them scale3 , as a fine-grained error component; Step 5.3: Adaptively fuse the error components of the three scales using learnable weights to obtain the comprehensive reconstruction error: in, For coarse-grained scale weighting parameters, For medium-granularity scale weighting parameters, β represents the weighting parameters at a fine-grained scale, and β represents the global bias parameter. , , The four parameters β are used as trainable parameters of the variational autoencoder (VAE), and are automatically learned during training through the backpropagation algorithm.
7. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 6, characterized in that, The variational autoencoder (VAE) pre-trained in step 3 or 4 is trained using a training dataset composed of normal surface image samples of industrial products. The specific training process is as follows: (1) Data preparation: Collect sufficient normal surface image samples of industrial products. All normal surface image samples are preprocessed according to step 1, and multi-scale feature vectors are extracted according to step 2 to construct training set and validation set; (2) Forward propagation: The multi-scale feature vector X of normal surface image samples in the training set is input into the variational autoencoder (VAE), and passes through the encoder part and the decoder part in sequence to obtain the reconstructed feature vector X. recon ; (3) Calculate the total loss L total : in, To reconstruct the loss, For KL divergence loss, These are the weighting coefficients for the KL divergence; Where S is the total number of scales. This represents the original eigenvector at scale s. For the reconstructed feature components at the s-th scale; Where 'a' represents the dimension of the latent variable. This is the mean vector of the predictions from the encoder portion. The standard deviation component of the encoder's prediction; (4) Backpropagation and parameter update: The backpropagation algorithm is used to simultaneously update all weights and bias parameters of the encoder part of the VAE, all weights and bias parameters of the decoder part, and multi-scale fusion weight parameters. , , and global bias parameter β; (5) Iterative training: Repeat the forward propagation to parameter update step until the model converges or reaches the preset maximum number of training rounds. When the loss on the validation set no longer decreases for several consecutive rounds, terminate the training in advance to avoid overfitting. (6) Threshold calibration: After training, the samples in the training set are fed into the variational autoencoder (VAE) to obtain the reconstructed multi-scale feature vectors. Then, the anomaly scores are calculated and the mean of the anomaly scores is calculated. and standard deviation And calculate the judgment threshold according to the following formula: Where T is the judgment threshold and k is an adjustable coefficient: k=2 corresponds to a 95% confidence interval suitable for high recall scenarios, k=3 corresponds to a 99.7% confidence interval suitable for standard scenarios, and k=4 corresponds to a 99.99% confidence interval suitable for high precision scenarios. (7) Mean based on outlier scores and standard deviation Set the following confidence intervals for outlier scores: .
8. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 7, characterized in that, Step 6 specifically involves: Step 6.1: Outlier score mapping: The integrated reconstruction error is mapped to a value between 0 and 1 using the Sigmoid activation function, which is used as the final outlier score S; Step 6.2: Compare the anomaly score with the judgment threshold. At that time, the surface images of the industrial products were normal. At that time, the surface image of the industrial product to be inspected contained defects; Step 6.3: Determine the level of abnormality: when At that time, it was determined that the surface image of the industrial product to be inspected had a slight anomaly; when At that time, it was determined that the surface image of the industrial product to be inspected showed moderate anomalies; when At that time, it was determined that the surface image of the industrial product to be inspected had serious anomalies; Step 6.4: Output the final judgment result, and also output the anomaly score, reconstruction error at each scale, and anomaly level.
9. The method for detecting surface defects in industrial products based on multi-scale feature reconstruction using a variational autoencoder according to claim 7, characterized in that, Step 7 specifically involves: Step 7.1: During online detection, M images judged as normal are collected to form an image batch, which is then fed into a pre-trained variational autoencoder (VAE) to obtain reconstructed multi-scale feature vectors before calculating anomaly scores. If the anomaly scores of N consecutive batches of normal images all fall within the confidence interval... In addition, the variational autoencoder (VAE) is updated as follows; Step 7.2: Using the parameters of the currently pre-trained variational autoencoder (VAE) as initial values, fine-tune the training using M newly acquired images that are judged to be normal: The learning rate and number of training epochs for fine-tuning training are lower than those for pre-training. During fine-tuning training, the parameters of the encoder part are frozen, and only the weights and bias parameters of the decoder part, as well as the multi-scale fusion weight parameters and global bias parameters are updated. Step 7.3: After fine-tuning, recalculate the abnormal score distribution using the new normal image and update the judgment threshold.