Industrial multi-mode anomaly detection method based on wavelet transform
By using wavelet transformation technology and reconstructing the network in industrial multimodal anomaly detection, the memory usage and inference delay problems of the memory memory bank method are solved, and efficient and accurate industrial product anomaly detection is achieved.
Patent Information
- Application Number
- CN202510444800.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing industrial multimodal anomaly detection method based on memory memory library has problems in memory usage and inference delay, and is difficult to deploy on resource-constrained industrial production lines, and is difficult to meet real-time detection requirements.
The industrial multimodal anomaly detection method based on wavelet transform is adopted to restore the abnormal image to normal images by reconstructing the network, and multi-scale features are extracted using wavelet transform technology, combining multi-modal data for abnormality discrimination and positioning.
It effectively avoids the high memory footprint and inference delay caused by memory memory banks, improves the accuracy and speed of product abnormality detection, and can realize real-time detection in resource-constrained environments.
Smart Images

Figure CN120198413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial product anomaly detection, and particularly relates to an industrial multimodal anomaly detection method based on wavelet transform. Background Art
[0002] In traditional industrial manufacturing processes, defect detection mainly relies on inspectors to visually check whether products are abnormal. However, this method not only has a high labor intensity and low efficiency, but is also greatly affected by human factors. For example, it is prone to missed detections or false detections due to fatigue, lack of experience, or subjective judgment errors. In addition, in the face of a large-scale production environment, manual inspection is difficult to meet the requirements of high precision, high consistency, and real-time performance. Therefore, how to achieve automated, efficient, and high-precision industrial product anomaly detection through intelligent means is a hot issue in the field of industrial manufacturing.
[0003] In recent years, with the rapid development of deep neural networks (DNNs) and computer vision technology, anomaly detection methods based on deep neural networks have gradually replaced traditional detection means, promoting the development of industrial anomaly detection towards automation and intelligence. Deep learning technology can learn the normal patterns of industrial products using large-scale data and accurately capture anomaly features without complex manual intervention design, resulting in significant improvements in both the detection accuracy and speed of industrial product anomaly detection. In particular, the combination of multimodal data such as RGB images and depth images has enabled industrial product anomaly detection to move towards a more accurate and efficient intelligent stage. Industrial multimodal anomaly detection (IMAD) aims to use advanced visual perception technology and intelligent algorithms to automatically detect and classify industrial products, accurately identify and locate defect areas. For example, it can be used to detect scratches, dents, cracks, deformations, and other defects on the product surface that may affect product quality, thereby ensuring the stability of the production process and the consistency of products.
[0004] In the field of industrial multi-modal anomaly detection, although high-performance neural network models can achieve high-accuracy anomaly recognition and localization effects, there are still several key defects in practical applications. In industrial multi-modal anomaly detection, data collection itself is a huge challenge. First of all, obtaining high-quality data often comes with high costs. Secondly, the number of normal samples in industrial scenarios far exceeds that of abnormal samples, and this sample imbalance phenomenon makes it particularly difficult to construct a high-quality, labeled dataset. Therefore, unsupervised learning methods have become a common technical path to solve the problem of industrial multi-modal anomaly detection. Specifically, such methods only use normal samples in the training stage and use both normal and abnormal samples in the testing stage for detection. For this reason, researchers have proposed a method based on a memory bank, that is, in the training stage, first use a neural network to extract the features of normal samples and then store them in the memory bank; in the testing stage, determine whether there is an anomaly by comparing the features of the test samples with the features in the memory bank. However, although this method is effective, it brings significant memory overhead. For resource-constrained platforms, deploying such methods not only faces hardware limitations but also causes a significant drop in the model inference speed, making it difficult to meet the requirements of real-time industrial production lines, thus limiting its application in actual scenarios. Summary of the Invention
[0005] In view of the problem that the existing mainstream method for industrial multi-modal anomaly detection, namely the method based on a memory bank, often brings huge memory overhead and inference latency, resulting in difficulties in being deployed on industrial production lines, the purpose of the present invention is to provide an industrial multi-modal anomaly detection method based on wavelet transform, which can avoid problems such as high memory occupancy and inference latency brought by the memory bank while improving the accuracy of product anomaly detection.
[0006] The technical solution of the present invention is as follows:
[0007] An industrial multi-modal anomaly detection method based on wavelet transform, the method includes the following steps:
[0008] Step 1: Establish a multi-modal image dataset composed of RGB images and other modal images, and generate abnormal multi-modal images from the multi-modal images in the dataset;
[0009] Step 2: Use a reconstruction network to extract multi-scale features from the abnormal multi-modal images generated in Step 1, and restore the abnormal multi-modal images to normal images;
[0010] Step 3: Use the wavelet transform method to perform feature transformation on the multi-scale features extracted by the reconstruction network, and fuse the image features extracted by the reconstruction network with the image features after performing wavelet transform;
[0011] Step 4: Combine the fused features with the original RGB image features for anomaly discrimination and localization.
[0012] Further, according to the industrial multi-modal anomaly detection method based on wavelet transform, the other modal image is a depth image.
[0013] Further, according to the industrial multi-modal anomaly detection method based on wavelet transform, in step 1, an anomaly generation network is used to generate anomaly multi-modal images from the multi-modal images in the dataset, as follows:
[0014]
[0015] where represents the anomaly generation network; X R represents the normal RGB image; M a represents the binary anomaly mask, generated by Perlin noise; is the reciprocal of M a ; β represents a hyperparameter; A represents texture features; represents the anomaly RGB image; X D represents the normal depth image, represents the anomaly depth image.
[0016] Further, according to the industrial multi-modal anomaly detection method based on wavelet transform, step 2 includes the following steps:
[0017] Step 2.1: Train a reconstruction network that can restore the anomaly RGB image to a normal RGB image and enable it to effectively extract multi-scale features of the RGB image, with the formula as follows:
[0018]
[0019] where the anomaly RGB image serves as the input to the reconstruction network ; net_i represents the i-th neural network module of the reconstruction network , and there are n layers in total; is the output of the i-th neural network module of the reconstruction network , that is, the i-th RGB image feature extracted, and a total of n multi-scale RBG image features are extracted;
[0020] Step 2.2: Train a reconstruction network that can restore the anomaly depth image to a normal depth image and enable it to effectively extract multi-scale features of the depth image, with the formula as follows:
[0021]
[0022] where the anomaly depth image As the input of the depth image reconstruction network ; net_i represents the i-th layer neural network module of the reconstruction network There are a total of n layers; is the output of the i-th layer network module of the reconstruction network That is, the i-th depth image feature extracted, and a total of n multi-scale depth image features are extracted.
[0023] Furthermore, according to the industrial multi-modal anomaly detection method based on wavelet transform, in step 3, the wavelet transform method is used to perform feature transformation on the multi-scale features extracted by the reconstruction network from the spatial dimension, and the cascade operation is used to fuse the image features extracted by the reconstruction network with the image features after performing wavelet transform.
[0024] Furthermore, according to the industrial multi-modal anomaly detection method based on wavelet transform, step 3 includes the following steps:
[0025] Step 3.1: Use the wavelet transform method to perform feature transformation on the multi-scale RGB image features extracted by the reconstruction network from the spatial dimension, that is, four sub-band features are generated according to the height and width of the multi-scale RGB image features, as shown in formulas (5) to (8):
[0026]
[0027] Where and are the four sub-band features obtained by performing wavelet transform on the multi-scale RGB image features from the height and width; represents the output feature of the i-th layer neural network module of the reconstruction network; (p,q) represents the index value of the four sub-band features; K represents the length of the filter, k ∈ [0,..., K - 1], v ∈ [0,..., K - 1]; and μ(·) represent the coefficients of the filter;
[0028] Then, perform a cascade operation on these four sub-band features as the RGB image wavelet features output after performing wavelet transform, as shown in formula (9)
[0029]
[0030] Where represents the i-th RGB image wavelet feature, i ∈ [1,..., n];
[0031] Step 3.2: Use the wavelet transform method to perform feature transformation on the multi-scale depth image features extracted by the reconstruction network in the spatial dimension. Four sub-band features are also generated according to the height and width of the multi-scale depth image features, as shown in Formulas 10 to 13:
[0032]
[0033] where and are the four sub-band features obtained by performing wavelet transform on the multi-scale depth image feature from the height and width; represents the output feature of the i-th neural network module of the reconstruction network; (p, q) represents the index value of the four sub-band features; K represents the length of the filter, k ∈ [0, …, K - 1], v ∈ [0, …, K - 1]; and μ(·) represent the coefficients of the filter;
[0034] Then, perform a concatenation operation on these four sub-band features as the depth image wavelet features output after performing wavelet transform, as shown in Formula (14):
[0035]
[0036] where represents the wavelet feature of the i-th depth image, i ∈ [1, …, n];
[0037] Step 3.3: Fuse the first two layers of multi-scale RGB image features extracted by the reconstruction network with the RGB image wavelet features after wavelet transform in Step 3.1, as shown in Formula (15):
[0038]
[0039] where f RGB is the fused feature; is the output of the first-layer neural network module of the reconstruction network, that is, the first RGB image feature extracted; is the output of the second-layer neural network module of the reconstruction network, that is, the second RGB image feature extracted; is the first RGB image wavelet feature; is the second RGB image wavelet feature;
[0040] Step 3.4: Fuse the first two layers of multi-scale depth image features extracted by the reconstruction network with the depth image wavelet features after wavelet transform in Step 3.2, as shown in Formula (16):
[0041]
[0042] where f Depth is the fused feature; is the output of the first-layer network module of the reconstruction network, i.e., the first depth image feature extracted; is the output of the second-layer network module of the reconstruction network, i.e., the second depth image feature extracted; is the reconstruction network is the output of the second-layer network module of the reconstruction network, i.e., the second depth image feature extracted; represents the wavelet feature of the first depth image; represents the wavelet feature of the second depth image;
[0043] Step 3.5: Concatenate the f RGB obtained in Step 3.3 and the f Depth obtained in Step 3.4 as the final fused feature, i.e., f final .
[0044] Furthermore, according to the industrial multi-modal anomaly detection method based on wavelet transform, the wavelet transform method is Haar wavelet.
[0045] Furthermore, according to the industrial multi-modal anomaly detection method based on wavelet transform, in Step 4, the fused feature obtained in Step 3 and the original RGB image feature are sent into the anomaly detection network for anomaly discrimination and localization, as shown in Formulas (18) to (21)
[0046] P1 = Pool(E1) (18)
[0047] U1 = Upsample(E2) (19)
[0048]
[0049] U2 = Upsample(D1),
[0050] where represents the concatenation operation; and respectively represent convolutional neural networks with different convolutional kernel sizes; E1, P1, E2, U1, S1, D1, U2, and D2 respectively represent the features obtained after performing different operations; specifically, E1 represents the concatenation of f final and X RThe features obtained after performing a convolution operation on the splicing; P1 represents the features obtained after performing a pooling operation on E1; E2 represents the features obtained after performing a convolution operation on P1; U1 represents the features obtained after performing an upsampling operation on E2; D1 represents the features obtained after performing a convolution operation on S1; U2 represents the features obtained after performing an upsampling operation on D1; D2 represents the features obtained after performing a convolution operation on U2; finally, D2 is used for anomaly discrimination and localization.
[0051] Further, according to the industrial multi-modal anomaly detection method based on wavelet transform, the wavelet transform method is Haar wavelet.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] Different from the traditional method based on the memory bank, the present invention first adopts a reconstruction network as the core framework for industrial multi-modal anomaly detection. By restoring the synthesized abnormal product samples into normal product samples, the reconstruction network has the ability to identify product anomalies and does not require additional memory space to store features. Therefore, the present invention can effectively avoid problems such as high memory occupancy and inference latency caused by the memory bank. In addition, by introducing wavelet transform technology into the reconstruction network, the ability of the neural network to extract fine-grained features of multi-modal (i.e., RGB images and depth images) is further improved. This is because the wavelet transform technology can extract multi-level features of the product from different spatial dimensions, so as to more accurately capture the subtle anomalies on the product surface. Therefore, by combining multi-modal data, the present invention can more effectively extract product features to improve the accuracy of product anomaly detection. Description of the Drawings
[0054] Figure 1 It is a schematic flowchart of the industrial multi-modal anomaly detection method based on wavelet transform in this embodiment;
[0055] Figure 2 It is the overall framework diagram of the industrial multi-modal anomaly detection method based on wavelet transform in this embodiment; Detailed Embodiments
[0056] To facilitate the understanding of this application, the following will describe this application in a more comprehensive manner with reference to the relevant drawings.
[0057] The core idea of the present invention is as follows: First, an anomaly generation network is used to generate RGB images and depth images with anomalies from the RGB images and depth images in industrial multi-modal anomaly detection; then, a reconstruction network is trained to restore the abnormal RGB images and depth images to normal images, aiming to enable the reconstruction network to have the ability to capture image anomalies and extract normal features; next, the RGB image features and depth image features extracted by the reconstruction network are subjected to wavelet transform, and a cascading operation is used to fuse the RGB image features and depth image features extracted by the reconstruction network with the RGB image features and depth image features after wavelet transform; finally, the fused features and the original RGB image features are fed into the anomaly detection network to perform anomaly discrimination and localization.
[0058] Figure 1 It is a schematic flowchart of the industrial multi-modal anomaly detection method based on wavelet transform in this embodiment; Figure 2 It is the overall framework diagram of the industrial multi-modal anomaly detection method based on wavelet transform in this embodiment; As Figure 1 and Figure 2 shown, the industrial multi-modal anomaly detection method based on wavelet transform includes the following steps:
[0059] Step 1: Obtain a multi-modal image (i.e., RGB image and depth image) dataset, and use an anomaly generation network to generate abnormal multi-modal images from the multi-modal images in the dataset;
[0060] In this embodiment, the multi-modal images are RGB images and depth images. After obtaining the multi-modal dataset for industrial multi-modal anomaly detection, this embodiment uses an anomaly generation network to generate abnormal RGB images (X R ) and depth images (X D ) in the dataset into abnormal multi-modal images ( and ). The relevant formula representations are shown in Formulas (1) and (2)
[0061]
[0062] where X R represents the normal RGB image in the dataset; M a represents a binary anomaly mask generated by Perlin noise; is the reciprocal of M a ; β represents a hyperparameter; A represents texture features; represents the abnormal RGB image generated by using the anomaly generation network; X D represents the normal depth image in the dataset, represents the abnormal depth image generated by using the anomaly generation network .
[0063] Step 2: Based on the abnormal multimodal images (abnormal RGB image and abnormal depth image ) obtained in Step 1, use the reconstruction network to restore the abnormal RGB image and depth image to normal images respectively, aiming to enable the reconstruction network to have the ability to capture image abnormalities and extract normal features.
[0064] Step 2.1: Train a reconstruction network that can restore an abnormal RGB image to a normal RGB image and enable it to effectively extract the multi-scale features of the RGB image. The related formula is shown in Formula (3)
[0065]
[0066] where the abnormal RGB image serves as the input to the reconstruction network ; net_i represents the i-th neural network module of the reconstruction network , and there are a total of n layers; is the output of the i-th neural network module of the reconstruction network , that is, the i-th RGB image feature extracted, and a total of n multi-scale RBG image features are extracted.
[0067] Step 2.2: Train a reconstruction network that can restore an abnormal depth image to a normal depth image and enable it to effectively extract the multi-scale features of the depth image. The related formula is shown in Formula (4)
[0068]
[0069] where the abnormal depth image serves as the input to the depth image reconstruction network ; net_i represents the i-th neural network module of the reconstruction network , and there are a total of n layers; is the output of the i-th network module of the reconstruction network , that is, the i-th depth image feature extracted, and a total of n multi-scale depth image features are extracted.
[0070] Step 3: First, perform feature transformation on the multi-scale RGB image features and multi-scale depth image features extracted by the reconstruction network in Step 2 from the spatial dimension using a wavelet transform method such as the Haar wavelet, and then use a cascading operation to fuse the first two layers of multi-scale RGB image features and multi-scale depth image features extracted by the reconstruction network with the first two layers of RGB image features and depth image features after feature transformation using the Haar wavelet.
[0071] Step 3.1: Use the wavelet transform method such as Haar wavelet to perform feature transformation on the multi-scale RGB image features extracted by the reconstruction network from the spatial dimension, that is, generate four sub-band features according to the height and width of the multi-scale RGB image features. The relevant formula expressions are shown in Formulas (5) to (8):
[0072]
[0073] Where and are the four sub-band features obtained by performing Haar wavelet on the multi-scale RGB image features from the height and width (i.e., the LL, LH, HL, and HH dimensions); represents the output feature of the i-th neural network module of the reconstruction network; (p,q) represents the index values of the four sub-band features; K represents the length of the filter, k ∈ [0,…,K-1], v ∈ [0,…,K-1]; and μ(·) represent the coefficients of the filter.
[0074] Then perform a concatenation operation on these four sub-band features as the RGB image wavelet features output after performing the wavelet transform. The relevant formula expression is shown in Formula (9)
[0075]
[0076] Where represents the i-th RGB image wavelet feature, i ∈ [1,…,n].[[]END]]
[0077] Step 3.2: Use the wavelet transform method such as Haar wavelet to perform feature transformation on the multi-scale depth image features extracted by the reconstruction network from the spatial dimension. Similarly, generate four sub-band features according to the height and width of the multi-scale depth image features. The relevant formula expressions are shown in Formulas 10 to 13:
[0078]
[0079] Where and are the four sub-band features obtained by performing Haar wavelet on the multi-scale depth image features from the height and width (i.e., the LL, LH, HL, and HH dimensions); represents the output feature of the i-th neural network module of the reconstruction network; (p,q) represents the index values of the four sub-band features; K represents the length of the filter, k ∈ [0,…,K-1], v ∈ [0,…,K-1]; and μ(·) represent the coefficients of the filter.
[0080] Then, perform a concatenation operation on these four sub-band features, which are used as the depth image wavelet features output after performing wavelet transform. The relevant formula is shown in formula (14):
[0081]
[0082] Where represents the wavelet feature of the i-th depth image, and i ∈ [1, …, n].
[0083] Step 3.3: Fuse the first two layers of multi-scale RGB image features extracted by the reconstruction network in Step 2.1 with the RGB image wavelet features after performing the Haar wavelet transform in Step 3.1. The relevant formula is shown in formula (15):
[0084]
[0085] Where f RGB is the fused feature; is the output of the first-layer neural network module of the reconstruction network, that is, the first RGB image feature extracted; is the output of the second-layer neural network module of the reconstruction network, that is, the second RGB image feature extracted; is the reconstruction network The output of the second-layer neural network module, that is, the second RGB image feature extracted; represents the first RGB image wavelet feature; represents the second RGB image wavelet feature.
[0086] Step 3.4: Fuse the first two layers of multi-scale depth image features extracted by the reconstruction network in Step 2.2 with the depth image wavelet features after performing the Haar wavelet transform in Step 3.2. The relevant formula is shown in formula (16):
[0087]
[0088] Where f Depth is the fused feature; is the output of the first-layer network module of the reconstruction network, that is, the first depth image feature extracted; is the output of the second-layer network module of the reconstruction network, that is, the second depth image feature extracted; is the reconstruction network The output of the second-layer network module, that is, the second depth image feature extracted; represents the wavelet feature of the first depth image; represents the wavelet feature of the second depth image.
[0089] Step 3.5: Then, concatenate f RGB obtained in Step 3.3 and f Depth obtained in Step 3.4 as the final fused feature, that is, ffinal , the related formula is shown in Formula (17):
[0090]
[0091] Step 4: Feed the final fused feature (f final ) and the original RGB image feature (X R ) into the anomaly detection network to perform anomaly discrimination and localization. The related formula is shown in Formulas (18) to (21):
[0092] P1 = Pool(E1) (18)
[0093] U1 = Upsample(E2) (19)
[0094]
[0095] U2 = Upsample(D1),
[0096] where represents the concatenation operation; and respectively represent convolutional neural networks with different convolutional kernel sizes; E1, P1, E2, U1, S1, D1, U2, and D2 respectively represent the features obtained after performing different operations; specifically, E1 represents the feature obtained by performing a convolutional operation after concatenating f final and X R ; P1 represents the feature obtained by performing a pooling operation on E1; E2 represents the feature obtained by performing a convolutional operation on P1; U1 represents the feature obtained by performing an upsampling operation on E2; D1 represents the feature obtained by performing a convolutional operation on S1; U2 represents the feature obtained by performing an upsampling operation on D1; D2 represents the feature obtained by performing a convolutional operation on U2; finally, D2 is used for anomaly discrimination and localization.
[0097] The industrial multi-modal anomaly detection method based on wavelet transform provided by the present invention provides an efficient and reliable solution for industrial product anomaly detection platforms with limited resources. By innovatively combining a reconstruction network with wavelet transform technology, this method constructs a lightweight anomaly detection model with low memory occupancy, high detection accuracy, and fast inference capabilities. Specifically, the present invention uses multi-modal data of RGB images and depth images to co-train the reconstruction network, enabling the network to fully exploit the complementary information of these two modal data, thereby achieving precise identification of product surface anomalies. At the same time, the introduction of the wavelet transform module significantly improves the model's ability to judge subtle anomalies. Based on this, the industrial multi-modal anomaly detection model proposed by the present invention can achieve advantages such as high accuracy, high inference speed, and low memory overhead in complex industrial scenario environments.
[0098] It should be understood that those skilled in the art, inspired by the technical concept of the present invention and without departing from the content of the present invention, can also make various improvements or transformations based on the above content, and this still falls within the protection scope of the present invention.
Claims
1. An industrial multimodal anomaly detection method based on wavelet transform, characterized in that: The method comprises the following steps: Step 1: Establish a multimodal image dataset consisting of RGB images and other modal images, and generate abnormal multimodal images from the multimodal images in the dataset; Step 2: Use the reconstruction network to extract multi-scale features from the abnormal multimodal image generated in step 1, and restore the abnormal multimodal image to a normal image; Step 3: Use the wavelet transform method to transform the multi-scale features extracted by the reconstruction network, and fuse the image features extracted by the reconstruction network with the image features after performing the wavelet transform; Step 4: Combine the fusion features with the original RGB image features to identify and locate anomalies.
2. The industrial multimodal anomaly detection method based on wavelet transform according to claim 1 is characterized in that: The other modality image is a depth image.
3. The industrial multimodal anomaly detection method based on wavelet transform according to claim 2 is characterized in that: In step 1, the abnormal generation network is used to generate abnormal multimodal images from the multimodal images in the dataset, as shown in the following formula: in represents the anomaly generation network; X R Represents a normal RGB image; M a represents a binary anomaly mask, generated by Perlin noise; It is M a Inverse; β represents hyperparameter; A represents texture feature; Indicates abnormal RGB image; X D represents a normal depth image, Represents an abnormal depth image.
4. The industrial multimodal anomaly detection method based on wavelet transform according to claim 3 is characterized in that: Step 2 includes the following steps: Step 2.1: Train a reconstruction network that can restore abnormal RGB images to normal RGB images And it can effectively extract the multi-scale features of RGB images. The formula is as follows: Among them, the abnormal RGB image As a reconstruction network Input; net_i represents the reconstructed network The i-th layer neural network module has a total of n layers; Reconstructing the network The output of the i-th layer neural network module is the i-th RGB image feature extracted, and a total of n multi-scale RBG image features are extracted; Step 2.2: Train a reconstruction network that can restore abnormal depth images to normal depth images And it can effectively extract the multi-scale features of the depth image. The formula is as follows: The abnormal depth image As a deep image reconstruction network Input; net_i represents the reconstructed network The i-th layer neural network module has a total of n layers; Reconstructing the network The output of the i-th layer network module is the i-th extracted depth image feature, and a total of n multi-scale depth image features are extracted.
5. The industrial multimodal anomaly detection method based on wavelet transform according to claim 2 is characterized in that: In step 3, the multi-scale features extracted by the reconstruction network are transformed from the spatial dimension using the wavelet transform method, and the image features extracted by the reconstruction network are fused with the image features after the wavelet transform using a cascade operation.
6. The industrial multimodal anomaly detection method based on wavelet transform according to claim 5 is characterized in that: Described step 3 comprises the following steps: Step 3.1: Use the wavelet transform method to transform the multi-scale RGB image features extracted by the reconstruction network from the spatial dimension, that is, generate four sub-band features according to the height and width of the multi-scale RGB image features, as shown in formulas (5) to (8): in and is a multi-scale RGB image feature Four sub-band features obtained from wavelet transform of height and width; Represents the output features of the neural network module in the i-th layer of the reconstructed network; (p, q) represents the index value of the four sub-band features; K represents the length of the filter, k∈[0,…,K-1], v∈[0,…,K-1]; and μ(·) represent the coefficients of the filter; Then, these four sub-band features are concatenated to obtain the wavelet features of the RGB image after wavelet transform, as shown in formula (9): in represents the wavelet feature of the i-th RGB image, i∈[1,…,n]; Step 3.2: Use the wavelet transform method to transform the multi-scale depth image features extracted by the reconstruction network from the spatial dimension, and also generate four sub-band features based on the height and width of the multi-scale depth image features, as shown in Formulas 10 to 13: in and It is a multi-scale deep image feature Four sub-band features obtained from wavelet transform of height and width; Represents the output features of the neural network module in the i-th layer of the reconstructed network; (p, q) represents the index value of the four sub-band features; K represents the length of the filter, k∈[0,…,K-1], v∈[0,…,K-1]; and μ(·) represent the coefficients of the filter; Then, these four sub-band features are cascaded to obtain the wavelet features of the deep image output after wavelet transform, as shown in formula (14): in Represents the wavelet features of the i-th depth image, i∈[1,…,n]; Step 3.3: The first two layers of multi-scale RGB image features extracted by the reconstruction network are fused with the wavelet features of the RGB image after wavelet transformation in step 3.1, as shown in formula (15): where f RGB It is the fusion feature; Reconstructing the network The output of the first layer of neural network module is the first RGB image feature extracted; Reconstructing the network The output of the second layer of neural network module is the second RGB image feature extracted; Represents the first RGB image wavelet feature; Represents the wavelet features of the second RGB image; Step 3.4: Fuse the first two layers of multi-scale depth image features extracted by the reconstruction network with the wavelet features of the depth image after wavelet transformation in step 3.2, as shown in formula (16): where f Depth It is the fusion feature; Reconstructing the network The output of the first layer of network module is the first extracted deep image feature; Reconstructing the network The output of the second layer network module is the second extracted deep image feature; Represents the wavelet features of the first depth image; Representing the wavelet features of the second depth image; Step 3.5: Substitute the f obtained in step 3.3 RGB and f obtained in step 3.4 Depth Cascade as the final fusion feature, i.e. f final .
7. The industrial multimodal anomaly detection method based on wavelet transform according to claim 6 is characterized in that: In step 4, the fusion features obtained in step 3 are sent together with the original RGB image features into the anomaly detection network for anomaly identification and location, as shown in formulas (18) to (21). in Represents a splicing operation; and denote convolutional neural networks with different convolution kernel sizes; E1, P1, E2, I1, S1, D1, U2, and D2 denote the features obtained after performing different operations; specifically, E1 denotes f final With X R The features obtained by performing a convolution operation on E1 after splicing; P1 represents the features obtained by performing a pooling operation on E1; E2 represents the features obtained by performing a convolution operation on P1; U1 represents the features obtained by performing an upsampling operation on E2; D1 represents the features obtained by performing a convolution operation on S1; U2 represents the features obtained by performing an upsampling operation on D1; D2 represents the features obtained by performing a convolution operation on U2; and finally D2 is used for abnormality identification and positioning.
8. The industrial multimodal anomaly detection method based on wavelet transform according to any one of claims 1, 5-6, characterized in that: The wavelet transform method is Haar wavelet.
Citation Information
Cited By
Three-mode unsupervised industrial anomaly detection method based on reconstruction network
CN120388024A
A trimodal unsupervised industrial anomaly detection method based on reconstruction network
CN120388024B