Industrial picture anomaly detection method and device based on deep learning
Through data augmentation and Y-shaped neural networks, deep learning has solved the problem of data imbalance and insufficient generalization capabilities in industrial image anomaly detection, and achieved efficient and accurate anomaly detection and image reconstruction, which is suitable for complex production environments.
Patent Information
- Application Number
- CN202510440567.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing deep learning technology has problems such as imbalance in training data, insufficient generalization capabilities of model and high real-time requirements in image anomaly detection in industrial production, making it difficult to adapt to complex and changeable production environments.
The number of abnormal images is expanded through the data augmentation method, an abnormality detection model based on Deepsvdd is constructed, and a neural network with a Y-shaped structure is adopted, and the loss functions of the main task and the auxiliary task are combined for joint optimization, and some model parameters are shared to realize feature learning of abnormal detection and image reconstruction.
It improves the feature extraction ability and adaptability of the model, enhances the accuracy and real-time nature of abnormal detection, and can effectively identify abnormalities in complex production environments, and is suitable for tasks such as image classification and object detection.
Smart Images

Figure CN120298380A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image anomaly detection, and in particular, to an industrial image anomaly detection method and device based on deep learning. Background Art
[0002] With the advent of the Industrial 4.0 era, China's industrial production is gradually transforming towards intelligence and automation. During the industrial production process, a large number of equipment, processes, and products need to be monitored in real time to ensure production quality and efficiency. Image recognition technology is increasingly widely used in industrial production. Among them, image anomaly detection, as an important branch of image recognition technology, is of great significance for improving production quality and reducing production costs.
[0003] During the industrial production process, product quality is the lifeline of an enterprise. By monitoring the images in the production process in real time, problems such as defects and shape inconsistencies on the product surface can be detected in a timely manner, and then measures can be taken for rectification to ensure product quality.
[0004] Traditional anomaly detection methods mainly rely on manual operation, which is not only inefficient but also easily affected by subjective factors, resulting in missed detections and false detections. The use of image anomaly detection technology based on deep learning can improve the detection speed and accuracy and reduce labor costs.
[0005] During the industrial production process, problems such as equipment failures and operation errors may lead to safety accidents. By monitoring the operating status and relevant parameters of the equipment in real time, abnormal situations can be detected in a timely manner to prevent accidents.
[0006] The image anomaly detection technology based on deep learning can provide rich data support for enterprises, helping enterprises analyze problems in the production process, thereby optimizing the production process and improving production efficiency.
[0007] Traditional image processing methods mainly rely on manually designed features and classifiers, and their detection effects are often not good for complex scenarios and changing production environments. In addition, traditional methods have high computational complexity when dealing with a large number of samples and are difficult to meet the requirements of real-time detection.
[0008] Although deep learning technology has achieved remarkable results in the field of image recognition, there are still the following limitations in the image anomaly detection in industrial production:
[0009] (1) Data imbalance: The abnormal images in industrial production are relatively few, resulting in an unbalanced training data set and affecting the detection effect;
[0010] (2) Insufficient model generalization ability: The models trained for specific scenarios are difficult to adapt to the anomaly detection in other scenarios;
[0011] (3) High real-time requirements: The industrial production environment is complex and changeable, requiring the anomaly detection algorithm to have high real-time performance. However, the existing deep learning models still need to be improved in terms of computing resources and speed. Summary of the Invention
[0012] The purpose of this application is to overcome the problems of unbalanced training data, insufficient model generalization ability, and poor feature extraction ability and scalability in the image anomaly detection of deep learning technology in industrial production, and to provide a deep learning-based industrial image anomaly detection method and device.
[0013] In the first aspect, a deep learning-based industrial image anomaly detection method is provided, including:
[0014] Obtain normal images and abnormal images containing objects to be detected, and expand the number of abnormal images through data augmentation;
[0015] Preprocess and label the normal images, abnormal images, and augmented abnormal images to construct a dataset;
[0016] Based on Deepsvdd, construct an anomaly detection model. The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images that share some model parameters. The main task part is used for feature extraction and mapping feature vectors to a feature representation space for image anomaly detection, and the auxiliary task part is used to learn the features of normal images by reconstructing normal images;
[0017] Use the dataset to train the anomaly detection model. Among them, combine the loss functions of the main task part and the auxiliary task part to jointly optimize all parameters, so that the anomaly detection model can simultaneously learn the features of anomaly detection and image reconstruction;
[0018] Use the trained anomaly detection model to perform anomaly detection on industrial images.
[0019] In some possible implementation manners, expanding the number of abnormal images through data augmentation includes:
[0020] Randomly crop a first region on a normal image, and paste the cropped first region to a random position of other normal images to generate abnormal images;
[0021] Randomly inject noise into normal images to simulate interference factors in production.
[0022] In some possible implementation manners, the network representation φ(xi;W) of normal samples falls inside the hypersphere:
[0023] ||φ(x i;W) - c|| 2 ≤R
[0024] The network representation φ(xi;W) of the abnormal sample falls outside the hypersphere;
[0025] ||φ(x i ;W) - c|| 2 >R
[0026] Where the input space X⊆Rd and the output space F⊆Rp, φ(x;W)∈F is the feature representation of the network φ for the input x∈X under the parameter W, φ(·;W) represents the neural network, x i represents the input, W represents the joint learning network parameter, ||φ(x i ;W) - c|| 2 represents the distance from the sample to the center point, c represents the center of the hypersphere, R represents the radius of the hypersphere. Assuming that the radius R>0 of the hypersphere and the center c∈F are pre - given, the soft - margin Deepsvdd objective is defined as:
[0027]
[0028] Where v and λ represent hyperparameters.
[0029] In some possible implementation manners,
[0030] In some possible implementation manners, the One - ClassDeepsvdd algorithm shrinks the hypersphere by minimizing the average distance from all data representations to the center. For the given test point x∈X, the anomaly score s is expressed as:
[0031] .
[0032] In some possible implementation manners, the anomaly detection model includes the main function of Deepsvdd and the self - supervised auxiliary function constructed by the decoder corresponding to the encoder. The main function defines a standard K - layer neural network. Among them, the parameter of the k - th layer is θ k , the stacked parameter vector θ=(θ1,…,θ K ) defines the entire model for the anomaly detection task. The loss function of the main function is lm(x,y;θ), where (x,y) is the test sample;
[0033] The loss function of the self - supervised auxiliary function is ls(x), and the auxiliary task shares part of the model parameters θ e =(θ1,…,θ κ ), until a certain κ∈{1,…,K}. These κ layers are called the shared feature extractor. The specific - task parameters θs=(θ′ κ+1, …, θ′ K ) is called the self-supervised task branch, and θm = (θ κ +1, …, θ K ) is called the main task branch.
[0034] In some possible implementation manners, the main task part sequentially includes a first input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fully-connected layer. Among them, a pooling layer and a first batch normalization layer are sequentially arranged after the first convolutional layer, the second convolutional layer, and the third convolutional layer.
[0035] In some possible implementation manners, the auxiliary task part sequentially includes a second input layer, a reshaping layer, a first transposed convolutional layer, a second transposed convolutional layer, a third transposed convolutional layer, a fourth transposed convolutional layer, and an output layer. Among them, a second batch normalization layer is arranged after the first transposed convolutional layer, the second transposed convolutional layer, the third transposed convolutional layer, and the fourth transposed convolutional layer. An upsampling layer is arranged after the first transposed convolutional layer, the second transposed convolutional layer, and the third transposed convolutional layer. An activation function layer is arranged after the first transposed convolutional layer, the second transposed convolutional layer, the third transposed convolutional layer, the fourth transposed convolutional layer, and the upsampling layer.
[0036] In a second aspect, an industrial picture anomaly detection device based on deep learning is provided, including:
[0037] A data acquisition and enhancement module, configured to acquire normal images and abnormal images containing objects to be detected, and expand the number of abnormal images through data enhancement;
[0038] A dataset construction module, configured to preprocess and label the normal images, abnormal images, and the expanded abnormal images, and then construct a dataset;
[0039] A model construction module, configured to construct an anomaly detection model based on Deepsvdd. The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images that share some model parameters. The main task part is used for feature extraction and mapping feature vectors to a feature representation space for picture anomaly detection. The auxiliary task part is used for learning the features of normal images by reconstructing normal images;
[0040] A training module, configured to train the anomaly detection model by using the dataset. Among them, the loss functions of the main task part and the auxiliary task part are combined to jointly optimize all parameters, so that the model can simultaneously learn the features of anomaly detection and image reconstruction;
[0041] A detection module, configured to perform anomaly detection on industrial pictures by using the trained anomaly detection model.
[0042] In a third aspect, a computer-readable storage medium is provided. The computer-readable medium stores program code for a device to execute, and the program code includes steps for executing the method in any of the implementation manners in the first aspect described above.
[0043] In a fourth aspect, an electronic device is provided. The electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the method in any of the implementation manners in the first aspect described above is implemented.
[0044] The present application has the following beneficial effects:
[0045] 1. By creating an auxiliary task part for image reconstruction in the present application, which shares some parameters with the main task part for image anomaly detection, an anomaly detection model with a Y-shaped structure is formed. By jointly training to improve the model performance, the loss functions of the main task part and the auxiliary task part are combined to jointly optimize all parameters, enabling the model to simultaneously learn the features of anomaly detection and image reconstruction. The Y-shaped structure and the joint training method enable the model to make full use of the information of the main task and the auxiliary task, learn more comprehensive and accurate feature representations, thereby improving the performance of the model in the anomaly detection task. By sharing some features, the model can learn more general feature representations, which are not only applicable to anomaly detection but also to other tasks such as image reconstruction, enhancing the feature extraction ability of the model. Among them, the flexibility of the Y-shaped structure enables the model to be easily extended to other related tasks, such as image classification, object detection, etc., improving the scalability and application scope of the model;
[0046] 2. In the test stage, a self-supervised learning problem is created based on a single test sample, and the parameters of the shared feature extractor are fine-tuned by minimizing the loss of the auxiliary task, enabling the anomaly detection model to be dynamically adjusted according to the characteristics of the test sample and better adapt to the distribution of the test sample. Especially when there are differences between the test sample and the training sample distributions, it can significantly improve the adaptability and detection performance of the model;
[0047] 3. The present application produces diverse anomaly samples through data augmentation, enabling the anomaly detection model to learn more types of anomaly features. Thus, when facing complex situations in actual production, it has stronger robustness and adaptability. Secondly, the richer anomaly samples enable the model to more accurately identify and distinguish normal and abnormal images, thereby improving the accuracy of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application.
[0049] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0050] Figure 1 is the flowchart of the industrial image anomaly detection method based on deep learning in Embodiment 1 of the present application;
[0051] Figure 2 is the schematic diagram of neural network data processing in the industrial image anomaly detection method based on deep learning in Embodiment 1 of the present application;
[0052] Figure 3 is the schematic diagram of the anomaly detection model in the industrial image anomaly detection method based on deep learning in Embodiment 1 of the present application;
[0053] Figure 4 is the schematic diagram of the main task part in the industrial image anomaly detection method based on deep learning in Embodiment 1 of the present application;
[0054] Figure 5 is the schematic diagram of the auxiliary task part in the industrial image anomaly detection method based on deep learning in Embodiment 1 of the present application;
[0055] Figure 6 is the flowchart of the cable aging degree monitoring method in Embodiment 1 of the present application;
[0056] Figure 7 is the structural block diagram of the industrial image anomaly detection device based on deep learning in Embodiment 2 of the present application;
[0057] Figure 8 is the internal structural schematic diagram of the electronic device in Embodiment 4 of the present application.
[0058] Reference numerals:
[0059] 100, data acquisition and enhancement module; 200, dataset construction module; 300, model construction module; 400, training module; 500, detection module. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0061] Embodiment 1
[0062] As Figure 1 shown, a method for detecting industrial image anomalies based on deep learning involved in Embodiment 1 of this application includes:
[0063] S100. Obtain normal images and abnormal images containing objects to be detected, and expand the number of abnormal images through data augmentation;
[0064] Since the number of normal image samples is far more than that of abnormal image samples in actual industrial production, in order to expand the samples of abnormal images, a random area is cropped from a normal image and pasted to a random position of another normal image to generate abnormal samples.
[0065] In addition, in order to simulate the interference in the actual production process, random noise is injected into the normal image samples
[0066] S200. Preprocess and label the normal images, abnormal images, and augmented abnormal images, and then construct a data set;
[0067] Specifically, the preprocessing includes normalization processing, that is, normalizing the normal images, abnormal images, and augmented abnormal images to make the sample data distribution consistent.
[0068] The preprocessed images are labeled to distinguish between normal and abnormal state images. Among them, the labeling can be completed by professional technicians using a labeling tool. After the labeling is completed, spot checks or full checks are carried out to ensure the accuracy of the labeling. The labeled data is made into a data set and divided into a training set, a validation set, and a test set for model training, validation, and testing.
[0069] S300. Build an anomaly detection model based on Deepsvdd. The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images that share some model parameters. The main task part is used for feature extraction and mapping the feature vectors to the feature representation space for image anomaly detection, and the auxiliary task part is used for learning the features of normal images by reconstructing normal images;
[0070] As Figure 2As shown in the figure, the neural network is trained based on "Deep Support Vector Data Description" (i.e., DeepSVDD), while minimizing the volume of the hypersphere that encloses the network representation of the data. In the left figure, \(X\subseteq\mathbb{R}^d\) is the input space, which contains all the samples during the training process. In the right figure, \(F\subseteq\mathbb{R}^p\) is the output space, and the neural network maps the input data to this output space. \(\varphi(\cdot; W)\) represents a neural network that maps the data in the input space to the output space. \(W\) in \(\varphi(\cdot; W)\) represents the weight parameters of the network, \(c\) is the center of the hypersphere, which is located in the output space \(F\), and \(R\) is the radius of the hypersphere. The goal of DeepSVDD is to train the neural network \(\varphi(\cdot; W)\) such that the neural network maps most of the data points into a hypersphere with the smallest volume. This hypersphere is determined by the center \(c\) and the radius \(R\). The network representation \(\varphi(x i ; W)\) of the normal samples falls inside the hypersphere, that is:
[0071] \(\|\varphi(x i ; W)-c\|\) 2 \(\leq R\)
[0072] The network representation \(\varphi(x i ; W)\) of the abnormal samples falls outside the hypersphere, that is:
[0073] \(\|\varphi(x i ; W)-c\|\) 2 \(>R\)
[0074] Among them, the input space \(X\subseteq\mathbb{R}^d\) and the output space \(F\subseteq\mathbb{R}^p\), \(\varphi(x; W)\in F\) is the feature representation of the input \(x\in X\) by the network \(\varphi\) under the parameter \(W\), \(\varphi(\cdot; W)\) represents the neural network, \(x i represents the input, \(W\) represents the joint learning network parameters, \(\|\varphi(x i ; W)-c\|\) 2 represents the distance from the sample to the center point, \(c\) represents the center of the hypersphere, and \(R\) represents the radius of the hypersphere.
[0075] For some input space \(X\subseteq\mathbb{R}^d\) and output space \(F\subseteq\mathbb{R}^p\), let \(\varphi(\cdot; W): X\rightarrow F\) be a neural network layer with \(L\) hidden layers and the weight set \(W = \{W_1,\ldots,W L \}\). That is to say, \(\varphi(x; W)\in F\) is the feature representation of the input \(x\in X\) by the network \(\varphi\) under the parameter \(W\). The goal of DeepSVDD is to jointly learn the network parameters \(W\) while minimizing the volume of the smallest volume hypersphere that contains the data in the output space \(F\). Assume that the radius \(R>0\) of the hypersphere and the center \(c\in F\) are given in advance. Define the soft margin DeepSVDD objective function as:
[0076]
[0077] The first term of the objective function is to minimize the volume of the hypersphere, and the second term is a penalty term for the points located outside the sphere after passing through the neural network. That is, if the distance from a point to the center point is greater than the radius R, the hyperparameter v will control the trade-off between the volume of the sphere and the violation of the boundary, allowing it to appear outside the boundary.
[0078] In the industrial production process, the training objects are often positive samples, and the objective function can be simplified to:
[0079]
[0080] where v and λ represent hyperparameters. One-Class DeepSVDD simply uses a quadratic loss function to penalize the distance of each network representation φ(x i ;W) to the center c ∈ F. The second term is still the network weight decay regularization term with the hyperparameter λ > 0. One-Class DeepSVDD can also be regarded as finding a hypersphere with the smallest volume centered at c. However, different from Soft-Boundary DeepSVDD, Soft-Boundary DeepSVDD shrinks the hypersphere by directly penalizing the radius and the data representations mapped outside the hypersphere, while One-Class DeepSVDD shrinks the hypersphere by minimizing the average distance of all data representations to the center. Once again, to make the data (on average) as close as possible to the center c, the neural network must extract the common variation factors of the data. Penalizing the average distance for all data points instead of allowing some data points to fall outside the hypersphere is consistent with the assumption that the training data mainly comes from one class.
[0081] For a given test point x ∈ X, the anomaly score s of the two DeepSVDD variants can be naturally defined by the distance of this point to the center of the hypersphere, that is:
[0082]
[0083] where W∗ are the network parameters of the trained anomaly detection model. For Soft-Boundary DeepSVDD, this score can be adjusted by subtracting the final radius R∗ of the trained model, so that the outliers (points outside the hypersphere) have positive scores, while the normal values have negative scores. Note that the network parameters W∗ (and R∗) completely describe a DeepSVDD model, and no data needs to be stored for prediction. Therefore, DeepSVDD has a very low memory complexity. This also allows for quick testing by simply evaluating the network φ at the test point x ∈ X using the learned parameters W∗, which is usually just a concatenation of simple functions.
[0084] The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images, which share some model parameters. The main task part is used for feature extraction and mapping feature vectors to a feature representation space for image anomaly detection. The auxiliary task part is used to learn the features of normal images by reconstructing normal images, as Figure 3 shown. That is, the anomaly detection model includes the main function of Deepsvdd and a self-supervised auxiliary function constructed by a decoder corresponding to an encoder. Here, Input_new represents the original image restored by the decoder. The loss2 here measures the difference between the restored image and the original image. The loss_total represents the total loss function, which is the weighted sum of the main task loss loss1 and the self-supervised auxiliary task loss loss2. By combining these two loss functions, the model can optimize anomaly detection and feature learning simultaneously during training. Loss1 is the main loss for anomaly detection, based on the distance from the data point to the center of the hypersphere. The role of the self-supervised auxiliary function is to update the main task part by updating the auxiliary task part during testing through sharing some parameters.
[0085] First, define a standard K-layer neural network, where the parameters of the k-th layer are θ k . The stacked parameter vector θ = (θ1, …, θ K ) defines the entire model for the anomaly detection task, and its loss function is lm(x, y; θ), where (x, y) is the test sample.
[0086] Assume that the training data (x1, y1), …, (x n , y n ) are independently and identically distributed samples drawn from the distribution P. The standard empirical risk minimization solves the following optimization problem:
[0087]
[0088] A self-supervised auxiliary task part is needed, and its loss function is ls(x). Here, the decoder corresponding to the encoder is selected as the auxiliary function, and the auxiliary task part shares some model parameters θ e = (θ1, …, θ κ ) until a certain κ ∈ {1, …, K}. These κ layers are called the shared feature extractor. The auxiliary task uses its own specific task parameters θ s = (θ′ κ + 1, …, θ′ K ). The unshared parameters θ s are called the self-supervised task branch, and θ m = (θ κ+1,…,θ K ) is called the main task branch. Intuitively, the joint architecture is a Y-shaped structure with a shared bottom and two branches, as Figure 3 shown.
[0089] Add the losses of the two tasks together and take the gradient with respect to the set of all parameters. Thus, the joint training problem is:
[0090]
[0091] Describe training during testing on a single test sample x. Briefly, training during testing fine-tunes the shared feature extractor θ by minimizing the auxiliary task loss on x e . This can be formulated as:
[0092]
[0093] As Figure 4 shown, it is the main task part. The main task part successively includes a first input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fully connected layer. Among them, a pooling layer and a first batch normalization layer are successively set after the first convolutional layer, the second convolutional layer, and the third convolutional layer. Specifically:
[0094] First input layer (Input1): Receives color image data of 32×32×3, where 3 represents the three color channels of RGB.
[0095] First convolutional layer (Conv1): Uses 32 5×5 convolutional kernels, with a stride of 1 and padding of 2 before and after, performs a convolutional operation on the input image without adding a bias term, and the output feature map size is 32×32×32. The convolutional operation formula is:
[0096]
[0097] Among them, represents the value of the k-th channel at the position (i,j) of the output feature map, represents the pixel value of the input image at the position (i+m,j+n), represents the weight of the k-th convolutional kernel at the position (m,n).
[0098] Second convolutional layer (Conv2): Performs a convolutional operation on the output of the first convolutional layer, uses 64 5×5 convolutional kernels, with a stride of 1 and padding of 2 before and after, without adding a bias term, and the output feature map size is 32×32×64.
[0099] Third Convolutional Layer (Conv3): Performs a convolution operation on the output of the second convolutional layer. It uses 128 5×5 convolutional kernels with a stride of 1, padding of 2 on both sides, and no bias term. The output feature map has a size of 32×32×128.
[0100] Pooling Layer: A 2×2 max pooling operation with a stride of 2 is used after each convolutional layer to reduce the spatial dimension of the feature map. The pooling operation formula is:
[0101]
[0102] where represents the value of the k-th channel of the pooled feature map at position (i,j).
[0103] First Batch Normalization Layer (BN1+Relu1, the first batch normalization layer + activation function layer): Batch normalization operation is used after each convolutional layer to accelerate the training process and improve the generalization ability of the model. The batch normalization operation formula is:
[0104]
[0105] where represents the value of the k-th channel of the input feature map at position (i,j), and represent the mean and variance of the k-th channel respectively, is a small constant to prevent division by zero, are learnable scale and shift parameters.
[0106] Fully Connected Layer (FC): Flattens the output of the last pooling layer into a one-dimensional vector, and then maps it to a 128-dimensional feature representation space through a fully connected layer. The weight matrix of the fully connected layer is W, the bias term is 0, and the output feature vector is: z = Wx, where x is the flattened input vector and z is the output feature vector.
[0107] In the main task part, the input 32×32×3 color image first passes through the first convolutional layer to extract low-level features, and then the spatial dimension of the feature map is reduced by the max pooling layer to reduce the computational amount. Subsequently, the image data sequentially passes through the second convolutional layer, the third convolutional layer, and the pooling layer to gradually extract higher-level features. After each convolutional layer, the first batch normalization layer normalizes the feature map to accelerate the training process and improve the generalization ability of the model. Finally, the flattened feature vector is mapped to a 128-dimensional feature representation space through the fully connected layer for anomaly detection of the image.
[0108] As Figure 5As shown, the auxiliary task part successively includes a second input layer, a reshaping layer, a first transposed convolutional layer, a second transposed convolutional layer, a third transposed convolutional layer, a fourth transposed convolutional layer, and an output layer. Among them, a second batch normalization layer is provided after each of the first transposed convolutional layer, the second transposed convolutional layer, the third transposed convolutional layer, and the fourth transposed convolutional layer. An upsampling layer is provided after each of the first transposed convolutional layer, the second transposed convolutional layer, and the third transposed convolutional layer. An activation function layer is provided after each of the first transposed convolutional layer, the second transposed convolutional layer, the third transposed convolutional layer, the fourth transposed convolutional layer, and the upsampling layer. Specifically:
[0109] Second input layer (Input2): Receives the low-dimensional feature representation output by the encoder, with a dimension of 128.
[0110] Reshaping layer: Reshapes the 128-dimensional feature vector into a three-dimensional tensor of 4×4×8 to prepare for subsequent transposed convolutional operations.
[0111] Transposed convolutional layer:
[0112] First transposed convolutional layer (ConvTranSpose2d1): Uses 128 5×5 transposed convolutional kernels, with a padding of 2, to perform a transposed convolutional operation on the reshaped tensor without adding a bias term. The formula for the transposed convolutional operation is:
[0113]
[0114] Where, represents the value of the k-th channel at the position (i,j) of the output feature map, represents the value of the input tensor at the position (i - m, j - n), represents the weight of the k-th transposed convolutional kernel at the position (m,n).
[0115] Second transposed convolutional layer (ConvTranSpose2d2): Performs a transposed convolutional operation on the output of the first transposed convolutional layer, uses 64 5×5 transposed convolutional kernels, with a padding of 2 and a stride of 2, without adding a bias term.
[0116] Third transposed convolutional layer (ConvTranSpose2d3): Performs a transposed convolutional operation on the output of the second transposed convolutional layer, uses 32 5×5 transposed convolutional kernels, with a padding of 2 and a stride of 2, without adding a bias term.
[0117] Fourth transposed convolutional layer (ConvTranSpose2d4): Performs a transposed convolutional operation on the output of the third transposed convolutional layer, uses 3 5×5 transposed convolutional kernels, with a padding of 2 and a stride of 2, without adding a bias term, and the output feature map size is 32×32×3, which is the same as the input image size.
[0118] Second Batch Normalization Layer (BN2): Batch normalization is used after each transposed convolutional layer to accelerate the training process and improve the generalization ability of the model. The formula for batch normalization is:
[0119]
[0120] where represents the value of the k-th channel at the position (i,j) of the input feature map, and represent the mean and variance of the k-th channel respectively, is a small constant to prevent division by zero, are learnable scaling and offset parameters.
[0121] Upsampling Layer: Upsampling is used after the first, second, and third transposed convolutional layers to double the spatial dimension of the feature map. In this embodiment, the bilinear interpolation method is used. The bilinear interpolation formula is:
[0122]
[0123] +
[0124] where represents the pixel value after upsampling, represent the values of the four surrounding pixels respectively.
[0125] Activation Function Layer (Relu2): The LeakyReLU activation function is used after each transposed convolutional layer and upsampling layer to introduce non-linearity. The formula for the LeakyReLU activation function is:
[0126]
[0127] where α is a small positive number, and in this embodiment, α = 0.01.
[0128] Output Layer: The Sigmoid activation function is used in the last layer to map each pixel value of the output feature map to the interval [0,1], obtaining a reconstructed image with the same size as the input image. The formula for the Sigmoid activation function is:
[0129]
[0130] In the auxiliary task part, the 128-dimensional low-dimensional feature vector output by the encoder is first converted into a three-dimensional tensor of 4×4×8 through a reshaping layer. Then, this tensor sequentially passes through the first, second, and third transposed convolutional layers and batch normalization layers, gradually increasing the spatial dimension and the number of channels of the feature map. After each transposed convolutional layer, an upsampling layer is used to double the spatial dimension of the feature map to restore it to a size close to that of the input image. Next, the number of channels of the feature map is reduced to 3 through the fourth transposed convolutional layer, which is the same as the number of RGB channels of the input image. Finally, each pixel value of the output feature map is mapped to the interval [0,1] through the Sigmoid activation function to obtain the reconstructed CIFAR10 image.
[0131] The reconstructed image in the auxiliary task part is used to evaluate the quality of the input image on the one hand. By comparing the differences between the input image and the reconstructed image, it can be judged whether there are abnormalities in the input image. If the difference between the reconstructed image and the input image is large, it indicates that the input image may contain abnormalities. On the other hand, the reconstructed image helps the anomaly detection model learn the features of normal samples. During the training process, the anomaly detection model enhances the understanding of the features of normal samples by reconstructing normal samples, thereby improving the recognition ability of abnormal samples.
[0132] The main task part (anomaly detection) and the auxiliary task part (reconstructing images) share some model parameters, so that the learning of the auxiliary task can be used to enhance the performance of the main task part. The main task part focuses on anomaly detection, while the auxiliary task part learns the features of normal samples by reconstructing normal samples. The main task part and the auxiliary task part complement each other and jointly improve the overall performance of the model.
[0133] S400. Use the dataset to train the anomaly detection model. Among them, the loss functions of the main task part and the auxiliary task part are combined to jointly optimize all parameters, so that the anomaly detection model can simultaneously learn the features of anomaly detection and image reconstruction;
[0134] Specifically, use the DeepSVDD model for training, and the goal is to minimize the volume of the hypersphere that encloses the network representation of the data. Combine data augmentation methods to improve the model's recognition ability for abnormal samples. Use the training-at-test-time method to fine-tune the model parameters at test time through self-supervised learning to improve the model's generalization ability. After training, use the validation set to evaluate the performance of the model, and calculate evaluation metrics such as accuracy, recall, and F1 score. Adjust the model parameters according to the validation results to optimize the model performance.
[0135] S500. Use the trained anomaly detection model to detect anomalies in industrial pictures.
[0136] Specifically, deploy the trained model to the industrial production environment for real-time monitoring. Ensure that the model can process input images in real time and detect abnormal situations in a timely manner.
[0137] As Figure 6 shown, taking the monitoring of cable aging degree as an example:
[0138] Step 1: Take pictures of the cable through a camera
[0139] Install a high-definition camera in the cable laying area and take pictures of the cable regularly.
[0140] The shooting frequency of the camera can be adjusted according to the cable's usage environment and aging speed, and gradually adjusted to once a week or once a month according to the situation.
[0141] Step 2: Collect historical pictures
[0142] Collect the pictures of the cable taken in the past few years, including pictures in normal state and known aging states. These historical pictures can be used as the data source for model training.
[0143] Ensure the diversity and representativeness of the historical pictures, covering the cable states under different environmental conditions (such as different temperatures, humidities, and lighting conditions).
[0144] Step 3: Perform data augmentation
[0145] Use the data enhancement method proposed in this application to process the collected normal cable pictures and generate more diverse abnormal samples. The specific methods include:
[0146] Random cropping and pasting: Randomly crop an area on a normal cable picture and paste it to a random position on another normal cable picture to create an abnormal data set.
[0147] Random noise injection: Randomly inject noise into the normal cable pictures to simulate the interference in actual production.
[0148] Step 4: Prepare annotations to form a training data set
[0149] Annotate the collected cable pictures and the pictures obtained by augmentation to distinguish between normal and aging state pictures. The annotation can be completed by professional technicians, and all of them should be checked again after completion to ensure the accuracy of the annotation.
[0150] Divide the annotated pictures into a training set, a validation set, and a test set for model training, validation, and testing.
[0151] Step 5: Model training
[0152] The DeepSVDD model is trained with the goal of minimizing the volume of the hypersphere that encloses the network representation of the data.
[0153] Combined with data enhancement methods, the model's ability to identify abnormal samples is improved.
[0154] Use the test-time training method to fine-tune model parameters during testing through self-supervised learning to improve the generalization ability of the model.
[0155] Step 6: Capture video footage from the real-time video captured by the camera at preset time intervals and feed it into the model for anomaly detection
[0156] In actual applications, the camera captures the video of the cable in real time and captures the video footage as the input image of the model at preset time intervals (such as every hour or every 12 hours).
[0157] The captured video footage is sent to the trained model for anomaly detection, which can automatically output whether the cable is not aged or has aged, thereby enabling real-time monitoring of the cable aging condition.
[0158] If an abnormality is detected, the system will automatically sound an alarm and send the abnormal image to maintenance personnel so that timely measures can be taken.
[0159] In this embodiment, by randomly cutting and pasting to create an abnormal data set, and injecting random noise to simulate interference in actual production, a variety of abnormal samples can be generated, and the model can learn more types of abnormal features, so that it has stronger robustness and adaptability when facing complex situations in actual production. Secondly, richer abnormal samples enable the model to more accurately identify and distinguish normal and abnormal images, thereby improving the accuracy of abnormality detection.
[0160] In the test phase, a self-supervised learning problem is created based on a single test sample, and the parameters of the shared feature extractor are fine-tuned by minimizing the auxiliary task loss. The test-time training method enables the model to dynamically adjust according to the characteristics of the test sample and better adapt to the distribution of the test sample, especially when there is a difference between the distribution of the test sample and the training sample, which can significantly improve the adaptability and detection performance of the model. Secondly, by fine-tuning the model during testing, the model can better capture the characteristics of the test sample, thereby improving the detection ability of unknown anomaly types and further enhancing the generalization ability of the model.
[0161] In this embodiment, a Y-shaped structure is adopted, with a shared feature extractor at the bottom. The two branches are respectively used for the main task part (anomaly detection) and the auxiliary task part (reconstructing images). By jointly training, the model performance is improved. The loss functions of the main task and the auxiliary task are combined to jointly optimize all parameters, enabling the model to simultaneously learn the features of anomaly detection and image reconstruction. The Y-shaped structure and the joint training method enable the model to make full use of the information of the main task and the auxiliary task, learning more comprehensive and accurate feature representations, thereby enhancing the model's performance in the anomaly detection task. Additionally, by sharing the feature extractor, the model can learn more general feature representations, which are applicable not only to anomaly detection but also to other tasks such as image reconstruction, enhancing the model's feature extraction ability. Secondly, the flexibility of the Y-shaped structure enables the model to be easily extended to other related tasks, such as image classification and object detection, improving the model's scalability and application scope.
[0162] Embodiment 2
[0163] As Figure 7 shown, an industrial picture anomaly detection device based on deep learning according to Embodiment 2 of the present application includes:
[0164] A data acquisition and enhancement module 100, configured to acquire normal images and abnormal images containing objects to be detected, and expand the number of abnormal images through data enhancement;
[0165] A dataset construction module 200, configured to preprocess and label the normal images, abnormal images, and the expanded abnormal images to construct a dataset;
[0166] A model construction module 300, configured to construct an anomaly detection model based on Deepsvdd. The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images that share part of the model parameters. The main task part is used for feature extraction and mapping the feature vectors to a feature representation space for picture anomaly detection, and the auxiliary task part is used for learning the features of normal images by reconstructing normal images;
[0167] A training module 400, configured to train the anomaly detection model using the dataset. Among them, the loss functions of the main task part and the auxiliary task part are combined to jointly optimize all parameters, so that the model can simultaneously learn the features of anomaly detection and image reconstruction;
[0168] A detection module 500, configured to perform anomaly detection on industrial pictures using the trained anomaly detection model.
[0169] It should be noted that for other specific implementations of the industrial image anomaly detection device based on deep learning in this embodiment, reference can be made to the specific implementations of the above-mentioned industrial image anomaly detection method based on deep learning. To avoid redundancy, it will not be elaborated here.
[0170] Embodiment 3
[0171] A computer-readable storage medium involved in Embodiment 3 of the present application, the computer-readable medium stores program codes for a device to execute, and the program codes include steps for executing the method in any one of the implementation manners in Embodiment 1 of the present application;
[0172] Among them, the computer-readable storage medium can be a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM); the computer-readable storage medium can store program codes. When the program stored in the computer-readable storage medium is executed by a processor, the processor is used to execute the steps of the method in any one of the implementation manners in Embodiment 1 of the present application.
[0173] Embodiment 4
[0174] As Figure 8 shown, an electronic device involved in Embodiment 4 of the present application, the electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the method in any one of the implementation manners in Embodiment 1 of the present application;
[0175] Among them, the processor can adopt a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits, and is used to execute relevant programs to implement the method in any one of the implementation manners in Embodiment 1 of the present application.
[0176] The processor can also be an integrated circuit electronic device with signal processing capabilities. In the implementation process, each step of the method in any one of the implementation manners in Embodiment 1 of the present application can be completed by the integrated logic circuit in the hardware of the processor or the instruction in the form of software.
[0177] The above-mentioned processor may also be a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the functions required to be executed by the units included in the data processing device of the embodiments of the present application, or executes the method in any one of the implementation manners in Embodiment 1 of the present application.
[0178] The above is only a preferred specific implementation manner of the present application; however, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application, according to the technical solution of the present application and its improved concept, makes an equivalent replacement or change, and should be covered by the protection scope of the present application.
Claims
1. An industrial image anomaly detection method based on deep learning, characterized in that, Including: Obtain normal images and abnormal images containing the object to be detected, and augment the number of abnormal images through data augmentation; Preprocess and label the normal images, abnormal images and augmented abnormal images to construct a dataset; Build an anomaly detection model based on Deepsvdd. The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images that share some model parameters. The main task part is used for feature extraction and mapping feature vectors to a feature representation space for image anomaly detection, and the auxiliary task part is used to learn the features of normal images by reconstructing normal images; Use the dataset to train the anomaly detection model. Among them, combine the loss functions of the main task part and the auxiliary task part to jointly optimize all parameters, so that the anomaly detection model can learn the features of anomaly detection and image reconstruction at the same time; Use the trained anomaly detection model to detect anomalies in industrial images.
2. The industrial image anomaly detection method based on deep learning according to claim 1, characterized in that, Augment the number of abnormal images through data augmentation, including: Randomly crop a first region on a normal image, and paste the cropped first region to a random position of other normal images to generate abnormal images; Randomly inject noise into normal images to simulate interference factors in production.
3. The industrial image anomaly detection method based on deep learning according to claim 1, characterized in that The network representation φ(xi;W) of normal samples falls inside the hypersphere: ||φ(x i ;W)−c|| 2 ≤R The network representation φ(xi;W) of abnormal samples falls outside the hypersphere; ||φ(x i ;W)−c|| 2 >R Among them, the input space \(X\subseteq\mathbb{R}^d\) and the output space \(F\subseteq\mathbb{R}^p\), \(\varphi(x;W)\in F\) is the feature representation of the input \(x\in X\) by the network \(\varphi\) under the parameter \(W\), \(\varphi(\cdot;W)\) represents the neural network, \(x\) i represents the input, \(W\) represents the jointly learned network parameters, \(\|\varphi(x\) i ;W)-c\| 2 represents the distance from the sample to the center point, \(c\) represents the center of the hypersphere, \(R\) represents the radius of the hypersphere. Assuming that the radius \(R > 0\) of the hypersphere and the center \(c\in F\) are given in advance, the soft-boundary DeepSVDD objective is defined as: ; Among them, v and λ represent hyperparameters.
4. The industrial image anomaly detection method based on deep learning according to claim 3, wherein, The One-Class Deepsvdd algorithm shrinks the hypersphere by minimizing the average distance from all data representations to the center. For a given test point x∈X, the anomaly score s is expressed as: 。 5. The industrial image anomaly detection method based on deep learning according to claim 4, wherein, The anomaly detection model includes the main function of Deepsvdd and the self-supervised auxiliary function constructed by the decoder corresponding to the encoder. The main function defines a standard K-layer neural network, where the parameters of the k-th layer are θ k , and the stacked parameter vector θ = (θ1, …, θ K ) defines the entire model for the anomaly detection task. The loss function of the main function is lm(x, y; θ), where (x, y) is the test sample; The loss function \(l_s(x)\) of the self-supervised auxiliary function, and the auxiliary task shares some model parameters \(\theta\). e \(= (\theta_1, \ldots, \theta\) κ ), until a certain \(\kappa\in\{1,\ldots,K\}\), these \(\kappa\) layers are called the shared feature extractor, and the specific task parameters \(\theta_s = (\theta' κ _{+1}, \ldots, \theta' K ) used by the self-supervised auxiliary function are called the self-supervised task branch, and \(\theta_m = (\theta κ _{+1}, \ldots, \theta K ) is called the main task branch.
6. The industrial image anomaly detection method based on deep learning according to claim 5, characterized in that, The main task part sequentially includes a first input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer and a fully connected layer. Among them, a pooling layer and a first batch normalization layer are sequentially arranged after the first convolutional layer, the second convolutional layer and the third convolutional layer.
7. The industrial image anomaly detection method based on deep learning according to claim 5, wherein The auxiliary task part sequentially includes a second input layer, a reshaping layer, a first transposed convolutional layer, a second transposed convolutional layer, a third transposed convolutional layer, a fourth transposed convolutional layer and an output layer. Among them, a second batch normalization layer is arranged after the first transposed convolutional layer, the second transposed convolutional layer, the third transposed convolutional layer and the fourth transposed convolutional layer. An upsampling layer is arranged after the first transposed convolutional layer, the second transposed convolutional layer and the third transposed convolutional layer. An activation function layer is arranged after the first transposed convolutional layer, the second transposed convolutional layer, the third transposed convolutional layer, the fourth transposed convolutional layer and the upsampling layer.
8. An industrial image anomaly detection device based on deep learning, characterized in that, Including: A data acquisition and augmentation module, which is used to obtain normal images and abnormal images containing the object to be detected, and augment the number of abnormal images through data augmentation; A dataset construction module, which is used to preprocess and label the normal images, abnormal images and augmented abnormal images to construct a dataset; A model construction module, which is used to construct an anomaly detection model based on Deepsvdd. The anomaly detection model includes a main task part for anomaly detection and an auxiliary task part for reconstructing images that share part of the model parameters. The main task part is used for feature extraction and mapping the feature vector to the feature representation space for picture anomaly detection, and the auxiliary task part is used for learning the features of normal images by reconstructing normal images; A training module, which is used to train the anomaly detection model by using the data set. Among them, the loss functions of the main task part and the auxiliary task part are combined to jointly optimize all parameters, so that the model can simultaneously learn the features of anomaly detection and image reconstruction; A detection module, which is used to perform anomaly detection on industrial pictures by using the trained anomaly detection model.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program codes for the device to execute, and the program codes include steps for executing the method according to any one of claims 1-7.
10. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the method according to any one of claims 1-7 is implemented.