Unsupervised anomaly detection method and device, computer device and readable storage medium
Patent Information
- Application Number
- CN202610935585.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-25
AI Technical Summary
但该类方法仅局限于空间域运算,未能利用异常的频域可分离特性,无法实现图像高低频特征的解耦建模;单一的递归压缩重建方式难以彻底分离频域特征,易受复杂纹理与背景干扰,降低了正常区域与异常区域的特征区分度,从而导致无监督异常检测的准确率偏低
[0011]上述无监督异常检测方法、装置、计算机设备及可读存储介质,首先由递归频域分解编码器处理待检测图像,得到信息维度完备的目标分解编码数据;再经频率感知先验重建模块对编码数据进行重建优化,规整正常特征并抑制异常干扰,获得第一重建结果;接着,通过跨域细节保留网络融合目标分解编码数据与第一重建结果,修正重建偏差、保留图像细节特征,生成第二重建结果;最后频率感知跨递归检测模块结合多源数据进行联合异常判别,相较于传统单一重建误差判别方式,可有效规避复杂背景与纹理干扰造成的误检、漏检问题,显著提高了无监督异常检测的准确率。
Smart Images

Figure CN122821184A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of anomaly detection technology, and in particular to an unsupervised anomaly detection method, apparatus, computer equipment, and readable storage medium. Background Technology
[0002] In the field of industrial quality inspection, unsupervised anomaly detection is a common method. Unsupervised anomaly detection refers to identifying abnormal areas that are inconsistent with the normal pattern by learning the distribution characteristics of normal data without the need to label abnormal samples.
[0003] Currently, traditional unsupervised anomaly detection methods rely on recursive coding reconstruction architectures to complete image feature processing and anomaly discrimination in the spatial domain. However, these methods are limited to spatial domain operations and fail to utilize the frequency domain separability of anomalies, thus failing to achieve decoupled modeling of high- and low-frequency features of the image. Furthermore, the single recursive compression reconstruction method is insufficient to completely separate frequency domain features and is susceptible to interference from complex textures and backgrounds, reducing the feature discrimination between normal and abnormal regions, resulting in low accuracy in unsupervised anomaly detection.
[0004] Therefore, improving the accuracy of unsupervised anomaly detection has become an urgent problem to be solved. Summary of the Invention
[0005] Therefore, it is necessary to provide an unsupervised anomaly detection method, apparatus, computer equipment, and readable storage medium to address the aforementioned technical problems, thereby improving the accuracy of unsupervised anomaly detection.
[0006] Firstly, this application provides an unsupervised anomaly detection method applied to an unsupervised anomaly detection system. The system includes: a recursive frequency domain decomposition encoder, a frequency-aware prior reconstruction module, a cross-domain detail preservation network, and a frequency-aware cross-recursive detection module. The system is trained solely on a normal sample dataset. The method includes: The target decomposition and coding data are obtained by processing the image to be detected through a recursive frequency domain decomposition encoder. The target decomposition and encoding data are processed by the frequency-aware prior reconstruction module to obtain the first reconstruction result; The target decomposition and encoding data and the first reconstruction result are processed by a cross-domain detail-preserving network to obtain the second reconstruction result; The second reconstruction result is processed by the frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected.
[0007] Secondly, this application provides an unsupervised anomaly detection device for use in an unsupervised anomaly detection system. The system includes: a recursive frequency domain decomposition encoder, a frequency-aware prior reconstruction module, a cross-domain detail preservation network, and a frequency-aware cross-recursive detection module. The system is trained solely on a normal sample dataset. The device includes: The encoding module is used to process the image to be detected through a recursive frequency domain decomposition encoder to obtain target decomposition encoding data; The first reconstruction module is used to process the target decomposition and coding data through the frequency-aware prior reconstruction module to obtain the first reconstruction result; The second reconstruction module is used to process the target decomposition and encoding data and the first reconstruction result through a cross-domain detail-preserving network to obtain the second reconstruction result; The anomaly detection module is used to process the second reconstruction result through the frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected.
[0008] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the method described above.
[0009] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.
[0010] Fifthly, this application provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described above.
[0011] The aforementioned unsupervised anomaly detection method, apparatus, computer equipment, and readable storage medium first process the image to be detected by a recursive frequency domain decomposition encoder to obtain target decomposition coding data with complete information dimensions. Then, the frequency-aware prior reconstruction module reconstructs and optimizes the coding data, regularizing normal features and suppressing abnormal interference to obtain a first reconstruction result. Next, the target decomposition coding data and the first reconstruction result are fused by a cross-domain detail-preserving network to correct reconstruction bias and preserve image detail features, generating a second reconstruction result. Finally, the frequency-aware cross-recursive detection module combines multi-source data for joint anomaly discrimination. Compared with the traditional single reconstruction error discrimination method, this can effectively avoid false detection and missed detection problems caused by complex background and texture interference, and significantly improve the accuracy of unsupervised anomaly detection. Attached Figure Description
[0012] Figure 1 This application provides an illustration of the application environment for an unsupervised anomaly detection method. Figure 2 This is a schematic diagram of the structure of an unsupervised anomaly detection system provided in an embodiment of this application; Figure 3 A flowchart illustrating an unsupervised anomaly detection method provided in this application embodiment; Figure 4 This is a schematic diagram of the structure of a recursive frequency domain decomposition encoder provided in an embodiment of this application; Figure 5 A structural block diagram of an unsupervised anomaly detection device provided in an embodiment of this application; Figure 6 An internal structural diagram of a computer device provided in an embodiment of this application; Figure 7 An internal structural diagram of another computer device provided in an embodiment of this application; Figure 8 This is an internal structural diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0014] The following explains some of the technical terms or phrases used in this application: Frequency domain decomposition refers to the process of decomposing an image signal into different frequency components through frequency domain analysis methods such as Fourier transform. Generally, it can be separated into low-frequency components representing the global structure and high-frequency components representing edge texture, so that features of different frequencies can be independently modeled and processed.
[0015] Unsupervised anomaly detection: an image detection method that does not require labeled anomaly samples. It can identify regions that are significantly different from normal patterns simply by learning the distribution patterns and feature representations of normal data. It is suitable for scenarios such as industrial quality inspection where there are few anomaly samples.
[0016] Cross-domain attention mechanism: An improved attention mechanism for achieving feature interaction and weighted fusion between two different feature domains (e.g., frequency domain and spatial domain), achieving cross-domain feature alignment and information complementarity through adaptive weight allocation.
[0017] Cross-level attention mechanism: An improved attention mechanism for realizing feature interaction and weighted fusion between features at different recursive levels under the same feature domain. For example, the fused features of the previous level are used as queries, and the spatial / frequency domain features of the current level are used as keys and values. The inter-level feature association weights are adaptively calculated to mine the progressive association information of multi-level reconstruction results, and to complete cross-level feature alignment, effective information filtering between levels, and multi-scale detail complementarity.
[0018] Multilayer Perceptron (MLP): A feedforward neural network composed of fully connected layers that achieves high-dimensional mapping and transformation of input features through multilayer linear transformations and nonlinear activation functions.
[0019] Decoding Multilayer Perceptron (MLP) inv ): A reverse mapping network symmetrically set with a multilayer perceptron maps the compressed latent features back to the original feature space.
[0020] Gradient map of an image: This refers to a feature map obtained by calculating the rate of change of image pixel values in the horizontal and vertical directions, which can highlight the edge and texture change information of the image.
[0021] Frequency domain: The domain that describes the characteristics of a signal with frequency as the independent variable. After converting a spatial domain image into a frequency domain signal through Fourier transform, the frequency distribution corresponding to the global structure and local details of the image can be analyzed intuitively.
[0022] Spatial domain: The domain that describes an image with pixel position as the independent variable, that is, the original pixel space of the image. All processing based directly on pixel values (e.g., filtering, edge detection) is performed in the spatial domain.
[0023] The unsupervised anomaly detection method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a communication network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0024] It should be explained that the terminal 102 or the server 104 can execute any of the implementation methods described in the unsupervised anomaly detection method provided in the embodiments of this application, and will not be repeated here.
[0025] The computer device described in the embodiments of this application may include at least one of terminal 102 or server 104.
[0026] Please see Figure 2 This application provides a schematic diagram of an unsupervised anomaly detection system. The unsupervised anomaly detection system (hereinafter referred to as the system) includes: a Recursive Frequency-Decomposition Encoder (RFDE), a Frequency-Aware Prior Reconstruction Module (FA-PRM), a Cross-Domain Detail Preservation Network (CD-DPN), and a Frequency-Aware Cross-Recursion Detection Module (FC-CRD), wherein: The recursive frequency domain decomposition encoder is used to perform multi-level frequency domain decomposition and feature compression on the input image, and outputs multi-scale high-frequency component sequences, low-frequency component sequences and latent representation sequences, providing multi-dimensional feature data in the frequency and spatial domains for subsequent modules.
[0027] The frequency-aware prior reconstruction module is used to receive the low-frequency features (i.e., low-frequency component sequences) output by RFDE, perform normalized reconstruction of the low-frequency features, suppress abnormal information, and output semantically correct low-frequency reconstruction results.
[0028] The cross-domain detail preservation network combines the original features output by RFDE with the reconstruction results of FA-PRM. It extracts details and recovers residuals through a cross-domain attention mechanism, and performs detail correction on the reconstruction results to obtain a high-precision reconstruction result that is closer to the structure of the original image.
[0029] The frequency-aware cross-recursive detection module receives the high-precision reconstruction results output by CD-DPN, performs spatial domain analysis, frequency domain analysis, and cross-domain scoring fusion on it, identifies regions inconsistent with the normal pattern by comparing reconstruction differences, and finally outputs the anomaly detection results.
[0030] Please see Figure 3 This application provides a flowchart of an unsupervised anomaly detection method, applied to an unsupervised anomaly detection system. The system includes: a recursive frequency domain decomposition encoder, a frequency-aware prior reconstruction module, a cross-domain detail-preserving network, and a frequency-aware cross-recursive detection module. The system is trained solely on a normal sample dataset. The method includes the following steps: S101. The image to be detected is processed by a recursive frequency domain decomposition encoder to obtain target decomposition coding data.
[0031] It's important to clarify that the system described above is trained solely on a normal sample dataset. This means that during the model training phase, each module uses only a sample dataset containing normal images for parameter learning and optimization, without introducing any samples containing anomalous regions or relying on any form of anomalous sample annotation or prior knowledge of anomalous patterns. During training, each module learns the inherent distribution characteristics of normal images in the frequency and spatial domains to establish a reconstruction and discrimination model that only reflects normal data patterns. In the inference phase, the system identifies regions inconsistent with the normal distribution simply by comparing the differences between the image to be detected and the learned normal pattern features, thus achieving unsupervised anomaly detection. This setup ensures that the system can complete model training and deployment without the participation of anomalous samples, perfectly meeting the practical application needs of industrial quality inspection and other scenarios where anomalous samples are scarce and difficult to annotate.
[0032] In some embodiments, the target decomposition and encoding data includes: target reconstruction results, low-frequency component sequences, and high-frequency component sequences; please refer to [link to relevant documentation]. Figure 4 This application provides a schematic diagram of a recursive frequency domain decomposition encoder. The recursive frequency domain decomposition encoder includes: a recursive compression decomposition unit and a recursive reconstruction synthesis unit. The recursive frequency domain decomposition encoder processes the image to be detected to obtain target decomposition and coding data, including: S11. The image to be detected is processed by a recursive compression decomposition unit to obtain a latent representation sequence; S12. The latent representation sequence is processed by recursively reconstructing the synthesis unit to obtain the target reconstruction result, the low-frequency component sequence, and the high-frequency component sequence.
[0033] Here, the latent representation sequence refers to the deep feature vector sequence obtained by the recursive compression decomposition unit encoding the image features of each level in the frequency domain decomposition and feature compression process of N first recursive levels, denoted as... ,in, This represents the deep feature vector output by the Nth recursive level; the target reconstruction result refers to the reconstructed image sequence output by the recursive reconstruction synthesis unit after level-by-level decoding and reconstruction based on the latent representation sequence, which has the same size as the input image to be detected; the low-frequency component sequence refers to the low-frequency signal components that represent the global structure and overall contour information of the image, obtained by separating the image in the frequency domain decomposition at each level, denoted as... ,in, This represents the low-frequency signal component output at the Nth recursive level; the low-frequency components mainly contain the basic background and main structural features of a normal image; the high-frequency component sequence refers to the high-frequency signal components representing image edges, textures, and detail changes obtained by separating the image in the frequency domain decomposition at each level, denoted as... in, This represents the high-frequency signal component output by the Nth recursive level; the high-frequency component mainly contains the texture, edge and other detailed features of the normal image, which are used for detail restoration and feature fusion of subsequent images.
[0034] Thus, by performing multi-level frequency domain decomposition and feature compression on the input image (i.e. the image to be detected), a potential representation sequence containing multi-scale abstract features is generated. Then, the sequence is decoded by a recursive reconstruction synthesis unit, and the target reconstruction result, low-frequency component sequence and high-frequency component sequence are output simultaneously. This achieves complete decoupling and multi-dimensional representation of the image frequency domain features, providing data support for subsequent processing.
[0035] In some embodiments, the recursive compression decomposition unit includes N first recursive levels connected in sequence; N is an integer greater than 1; the recursive compression decomposition unit processes the image to be detected to obtain a latent representation sequence, including: S21. In the first recursive level of the target, the image to be detected is processed to obtain the first high-frequency component and the first low-frequency component; the first recursive level of the target is the first in the N first recursive levels; S22. Compress and encode the first high-frequency component and the first low-frequency component respectively to obtain the high-frequency feature code and the low-frequency feature code; S23. Based on the cross-level attention mechanism, high-frequency feature encoding and low-frequency feature encoding are fused to obtain the first fused data; S24. Process the first fused data sequentially through N-1 first recursive levels (excluding the target first recursive level) to obtain the potential representation sequence.
[0036] Specifically, in the first recursive level of the target, the image to be detected... ( H, W, C The height, width, and number of channels of the image to be detected are respectively subjected to a two-dimensional discrete Fourier transform to convert it to the frequency domain, resulting in the first frequency domain data, as follows: ; in, This represents the frequency domain representation (i.e., the first frequency domain data) obtained after the image undergoes a two-dimensional discrete Fourier transform; DFT() represents the two-dimensional discrete Fourier transform function. This represents the input image of the first recursive level of the target, which is also the output image of the level preceding the first recursive level. Since the first recursive level of the target is the first recursive level in the sequence of N first recursive levels, it has no preceding level. Therefore, the input image of the first recursive level of the target is the image to be detected. Here, i is the level index, where i is an integer greater than 0 and less than or equal to N. Then, an ideal high-pass filter is used. Separate the high-frequency and low-frequency components from the first frequency domain data:
[0037] in, Indicates the current frequency point To the center of the frequency domain ( H / 2, W The Euclidean distance of / 2). It is the cutoff frequency threshold, determined by the hyperparameter. Control, in which different recursion levels can use different cutoff frequency thresholds. m The specific settings can be preset or defaulted; the separation process is achieved through point-by-point multiplication:
[0038]
[0039] in, This represents the high-frequency component of the i-th recursive level (i.e., the first high-frequency component). This represents the low-frequency component of the i-th recursive level (i.e., the first low-frequency component). This represents the high-pass filter mask for the i-th recursive level, which is related to... For matrices of the same size, high-frequency regions are preserved and low-frequency regions are suppressed by point-by-point multiplication. Specifically, the original frequency components are preserved at positions with a value of 1, and the frequency components at positions with a value of 0 are set to zero. CROP() represents element-wise multiplication; Operation functions, through The operation function removes high-frequency components from the frequency domain, resulting in low-frequency components.
[0040] It needs to be explained that, The operation function can be a user-defined function. It is a frequency domain cropping operation. This operation uses the center of the frequency domain as a reference and removes the high-frequency components of the frequency domain edge region by cropping or masking, retaining only the low-frequency components of the corresponding global image structure information, thereby achieving decoupling of high and low frequency features.
[0041] Next, the first high-frequency component and the first low-frequency component can be compressed and encoded separately to obtain high-frequency feature codes and low-frequency feature codes; then, the high-frequency feature codes and low-frequency feature codes can be fused based on a cross-level attention mechanism to obtain the first fused data, as follows: ; in, This represents the fusion feature of the i-th recursive level (i.e., the first fusion data); This represents the high-frequency feature (i.e., high-frequency feature encoding) of the i-th recursive level. This represents the low-frequency feature of the i-th recursive level (i.e., low-frequency feature encoding). This represents the fused feature of the (i-1)th recursive level. If the current recursive level is the first level and has no preceding level, then... It can be determined as a preset initial feature, which can be preset in advance or set by default; This represents a cross-level attention fusion operator, which achieves feature-weighted fusion based on an attention mechanism. Its core is Query-Key-Value (QKV) attention computation, specifically, the attention fusion of the previous level... For querying, the current level For key The value is calculated using an attention mechanism to adaptively determine the level of attention given to high-frequency details or low-frequency structures at the current level, thereby obtaining... .
[0042] Finally, the first fused data can be processed sequentially through N-1 first recursive levels (excluding the target first recursive level) to obtain a potential representation sequence. Specifically, after the target first recursive level outputs the first fused data, the subsequent N-1 first recursive levels receive and process the fused data from the previous level in sequence. After N rounds of recursive decomposition and feature fusion, a potential representation sequence containing multi-scale abstract features is generated.
[0043] Thus, through first-level frequency domain decomposition, high- and low-frequency coding and cross-domain attention fusion, and then through N-1 subsequent progressive processing, multi-scale decoupling and adaptive representation of image frequency domain features are achieved, taking into account both detailed information and global structure, providing more discriminative feature sequences for subsequent anomaly detection, and improving the system's ability to identify anomalies at different scales.
[0044] In some embodiments, the recursive compression decomposition unit further includes: a shared encoder and a multilayer perceptron; and performs compression encoding on the first high-frequency component and the first low-frequency component respectively to obtain high-frequency feature encoding and low-frequency feature encoding, including: S31. The first high-frequency component is compressed and encoded using a shared encoder to obtain the high-frequency feature code; S32. The first low-frequency component is compressed and encoded using a multilayer perceptron to obtain the low-frequency feature code.
[0045] Specifically, the shared encoder is composed of a multi-layer convolutional neural network, which can learn the edge and texture representation of the material from high-frequency components. By inputting the first high-frequency component into the shared encoder for compression encoding, the high-frequency feature code is output, as follows: ; in, Represents high-frequency feature encoding; Indicates the first high-frequency component; Indicates a shared encoder; This represents the encoder parameters, which are shared across all recursive levels. This is a key feature that distinguishes recursive architectures from traditional deep networks.
[0046] Then, the first low-frequency component is input into a multilayer perceptron for compression encoding to obtain the low-frequency feature encoding, as follows: ; in, This represents low-frequency feature encoding; Indicates the first low-frequency component; This refers to a multilayer perceptron, which is sufficient to capture the global structural information of low-frequency components because the dimensionality of low-frequency components is much lower than that of the original image.
[0047] It's important to explain that the recursive frequency domain decomposition encoder employs a recursive architecture. All recursive layers share the same shared encoder and decoder, and the encoder and decoder parameters are identical for each recursive layer. This means that each layer reuses the same set of learnable weights during feature encoding and decoding. This design eliminates the need to maintain independent network parameters for different layers, significantly reducing the number of system parameters and training costs. Furthermore, parameter sharing ensures consistency in feature learning across layers, enabling the system to uniformly learn a common pattern for extracting features from frequency domain components across different layers. This avoids parameter redundancy and overfitting risks between layers, improving the system's training efficiency and generalization performance.
[0048] Thus, by using a shared encoder to compress and encode the first high-frequency component, and simultaneously using a multilayer perceptron to compress and encode the first low-frequency component separately, differentiated encoding processing can be performed on the respective distribution characteristics of high-frequency detail features and low-frequency structural features. This approach can reduce the model size and maintain the consistency of high-frequency feature extraction logic by leveraging the parameter reuse advantage of the shared encoder, while also adapting the encoding requirements of low-frequency global structural features through the multilayer perceptron. This achieves feature decoupling and targeted representation of high and low frequency components, reducing feature information aliasing and detail loss.
[0049] In some embodiments, the recursive reconstruction synthesis unit and the recursive compression decomposition unit form a symmetrical structure. The recursive reconstruction synthesis unit includes N second recursive levels connected in sequence, but the information flow directions of the two are opposite. When decomposing, the recursive compression decomposition unit starts from the shallowest level 1 and decomposes step by step to the deeper level, all the way to the Nth level, and finally obtains the target decomposed encoded data. When reconstructing, the recursive reconstruction synthesis unit starts from the deepest level N and decodes step by step to the shallowest level, all the way to the first level, and finally restores the complete target reconstruction result.
[0050] In some embodiments, the latent representation sequence is processed by a recursive reconstruction synthesis unit to obtain the target reconstruction result, the low-frequency component sequence, and the high-frequency component sequence, including: S41. Reconstruct the image from deep to shallow layers according to the latent representation sequence through N second recursive levels to obtain the target reconstruction result; S42. Based on the target reconstruction results, determine the low-frequency component sequence and the high-frequency component sequence.
[0051] Specifically, using N second recursive levels, the latent representation sequence is decoded in reverse step by step and the image is restored layer by layer in order from deep features to shallow features, finally reconstructing the complete target reconstruction result. Then, based on the target reconstruction result, the low-frequency component sequence and the high-frequency component sequence are determined. Specifically, the target reconstruction result can be reprocessed by frequency domain decomposition. Through Fourier transform, frequency domain high and low frequency component splitting, and frequency domain clipping operations, the low-frequency components and high-frequency components corresponding to each recursive level are separated layer by layer and level by level. The low-frequency components and high-frequency components of each level are collected in the order of recursive levels and regularized to form the low-frequency component sequence and the high-frequency component sequence, respectively.
[0052] In some embodiments, during the image reconstruction process, the low-frequency components and high-frequency components of each level can be directly recorded and arranged in a recursive hierarchical order to form a low-frequency component sequence and a high-frequency component sequence. This eliminates the need to perform frequency domain decomposition on the target reconstruction result again, simplifies the calculation process, and ensures the one-to-one correspondence between the hierarchical components.
[0053] In this way, by reconstructing the target from deep to shallow layers, the reconstruction result is obtained. Then, the low-frequency and high-frequency component sequences are decoupled from the reconstruction result. This not only fully restores the global structure and texture details of the image, but also orderly separates the structural information from the detail information, providing a regular and complete frequency domain reference component for subsequent detail correction and anomaly feature comparison, thus ensuring accurate and reliable anomaly detection.
[0054] In some embodiments, the recursive reconstruction synthesis unit further includes: a shared decoder and a decoding multilayer perceptron; reconstructing the image from deep to shallow layers according to the latent representation sequence through N second recursive levels to obtain the target reconstruction result, including: S51. In the second recursive level of the target, the high-frequency components in the latent representation sequence are decoded by a shared decoder to obtain high-frequency reconstructed feature data; the second recursive level of the target is the last in the N second recursive levels. S52. Low-frequency components in the latent representation sequence are inversely projected by a decoding multilayer perceptron to obtain low-frequency reconstructed feature data. S53. Fuse the high-frequency reconstruction feature data and the low-frequency reconstruction feature data to obtain the target frequency spectrum; S54. Determine the intermediate reconstructed image based on the target frequency spectrum; S55. Determine the first residual information between the intermediate reconstructed image and the image to be detected; S56. Refine the intermediate reconstructed image based on the first residual information to obtain the target reconstructed image; S57. Process the target reconstruction image from deep to shallow layers through N-1 second recursive layers other than the target second recursive layer to obtain the target reconstruction result.
[0055] Specifically, in the second recursive level of the target, the latent representation sequence can be input into the shared decoder for decoding, and the high-frequency reconstructed feature data can be output as follows: ; in, This represents high-frequency reconstructed feature data; Indicates a shared decoder; Indicates decoder parameters; This represents the fusion features (i.e., the latent representation sequence).
[0056] It needs to be explained that the shared decoder and shared encoder have symmetrical structures, but their parameters are not shared and are independent of each other; the shared decoder is specifically responsible for upsampling the abstract latent representation step by step, restoring the image spatial structure, decoding in reverse step by step, and finally outputting high-frequency reconstructed features.
[0057] Then, the low-frequency components in the latent representation sequence can be inversely projected onto the decoding multilayer perceptron to obtain low-frequency reconstructed feature data, as follows: ; in, This represents low-frequency reconstruction feature data; This indicates decoding a multilayer perceptron; This represents the low-frequency components in the potential representation sequence; Then, the high-frequency reconstruction feature data and the low-frequency reconstruction feature data can be fused, as follows: ; in, Represents the target frequency spectrum; MERGE (,) represents the frequency domain merging operator, which follows the frequency domain addition principle, adding the high-frequency reconstruction features and the low-frequency reconstruction features at corresponding positions to reconstruct the complete frequency spectrum; according to the above formula, the target frequency spectrum can be obtained.
[0058] Furthermore, the intermediate reconstructed image can be determined based on the target frequency spectrum, as follows: ; in, The intermediate reconstructed image is represented by IDFT(), which represents the inverse discrete Fourier transform operator. Its function is to convert the frequency domain signal back to the spatial domain and restore the image with pixel values.
[0059] Then, the first residual information between the intermediate reconstructed image and the image to be detected can be determined. Specifically, the intermediate reconstructed image and the image to be detected can be aligned first, for example, by downsampling or spatial alignment preprocessing to ensure that the pixel space of the two corresponds one-to-one. Then, the pixel-by-pixel residual can be calculated. The difference between the intermediate reconstructed image and the image to be detected is calculated according to the pixel position to generate a residual map of the same size. This residual map is used as the first residual information. In some embodiments, the residual map can also be normalized or thresholded to highlight the error characteristics of abnormal regions and obtain the first residual information.
[0060] Next, the intermediate reconstructed image can be refined based on the first residual information, as follows: ; in, Represents the target reconstructed image; Indicates the first residual information; This represents the reconstructed image at the current level (i.e., the intermediate reconstructed image). This represents the input image (i.e., the image to be detected) at the current level; α represents the preset thinning coefficient; thus, the target reconstructed image can be obtained.
[0061] Finally, the target reconstruction image can be processed sequentially through the remaining N-1 second recursive levels, from deep to shallow (i.e. from the Nth level to the first level). Through step-by-step upsampling, detail restoration, and frequency domain synthesis, the complete target reconstruction result is finally obtained.
[0062] Thus, by differentially decoding and fusing the high and low frequency components of the latent representation sequence at the second recursive level of the target, the target frequency spectrum is generated and converted into an intermediate reconstructed image. Then, based on the residual information, it is refined and corrected, and processed step by step by the subsequent N-1 second recursive levels. This not only ensures the integrity of the frequency spectrum information, but also realizes error correction and detail enhancement, effectively improving the global structural consistency and local texture restoration of the reconstructed image.
[0063] S102. The target decomposition and coding data are processed by the frequency-aware prior reconstruction module to obtain the first reconstruction result.
[0064] In some embodiments, the frequency-aware prior reconstruction module includes: a learnable prior context library; and processes the target decomposition and encoding data through the frequency-aware prior reconstruction module to obtain a first reconstruction result, including: S61. Perform a linear transformation on the low-frequency component sequence to obtain the query vector; S62. Perform a linear transformation on all prior vectors in the learnable prior context library to obtain key vectors and value vectors; S63. Determine the attention weight distribution based on the query vector and key vector; S64. Determine the attention aggregation result based on the attention weight distribution, the learnable prior context library, and the value vector; S65. Based on the gating mechanism, the low-frequency component sequence and attention aggregation result are processed to obtain the reference reconstruction result; S66. Perform feature projection on the high-frequency component sequence to obtain the high-frequency projection result; the dimension of the high-frequency projection result is consistent with the reference reconstruction result. S67. The high-frequency projection results and the reference reconstruction results are fused to obtain the target fusion result; S68. Based on the target fusion result, determine the first reconstruction result.
[0065] Among them, the learnable prior context library can be denoted as P ={ p 1, p 2, ..., p k This library stores typical sample representations of low-frequency patterns for various normal materials; the library follows the following three design principles: First, each prior context The dimension of the signal space after low-frequency decomposition is kept in match to ensure that the prior representation can be directly compared with the actual low-frequency features. Second, the size of the library. k Determines the diversity of representation, k When the range of values is sufficient, it can cover the characteristic differences of normal materials; however, if the range of values is too large, it will increase the computational overhead and cause redundancy in prior characterization. Third, the prior context is initialized using the k-means++ algorithm, which can generate better initial cluster centers compared to random initialization, thus accelerating the network convergence speed. The k-means++ algorithm is an improved version of the classic k-means clustering algorithm, and its core optimization lies in the method of initializing cluster centers (centroids).
[0066] The learnable prior context library undergoes end-to-end optimization during training. Each prior vector can adaptively adjust according to the distribution of training data, gradually evolving into typical patterns that can characterize the low-frequency features of normal materials. This learning mechanism is fundamentally different from traditional clustering-based pre-computation methods: end-to-end optimization allows the library to directly learn the most valuable representations for the reconstruction task, rather than merely capturing the statistical properties of the training data.
[0067] Specifically, a linear transformation can be performed on the low-frequency component sequence, as follows: ; in, Q Represents the query vector; This represents a linear transformation layer used to generate a query vector, which maps the input low-frequency features (i.e., low-frequency component sequences) to a query vector of a specified dimension. This represents the low-frequency component sequence; the query vector can be obtained using the above formula.
[0068] Then, a linear transformation can be performed on all prior vectors in the learnable prior context library, as follows: ; ; Where K represents the key vector; V represents the value vector; and P represents all prior vectors in the learnable prior context library. This represents a linear transformation layer used to generate the key vector, mapping the prior vector to a key vector of the same dimension as the query vector; This represents a linear transformation layer used to generate value vectors, mapping prior vectors to value vectors of a specified dimension; thus, key vectors and value vectors can be obtained.
[0069] Next, the attention weight distribution can be determined based on the query vector and key vector, as follows: ; Where Attn represents the attention weight distribution; Softmax() represents the normalized exponential function; Indicates the scaling factor. The dimension of the key vector is represented by the formula, which is obtained by calculating the dot product of the query vector Q and the key vector K, and then normalized by Softmax() to obtain the probability distribution (i.e., the attention weight distribution). The attention weight distribution represents the degree of contribution of each prior vector to the current query.
[0070] Next, the attention aggregation result can be determined based on the attention weight distribution, the learnable prior context library, and the value vector, as follows: ; in, This represents the attention aggregation result, which is a weighted sum of all prior vectors.
[0071] Furthermore, based on a gating mechanism, the low-frequency component sequence and attention aggregation results can be processed to obtain a reference reconstruction result, as follows: Gating signal generation; Gated fusion output; in, The reference reconstruction result is the output after gated fusion, which combines prior knowledge and original information; gate represents the gate signal, which has a value range of 0 to 1 and is used to control the weight distribution of the two inputs in subsequent fusion; Sigmoid() represents the Sigmoid activation function, which compresses the output of the linear transformation to the range of 0 to 1 to ensure that the value of the gate signal meets the weight requirements. This represents a linear transform layer that generates the gated signal. The input is the concatenated features, and the output is used to calculate the gate value. Indicates will and Concatenated along the channel dimension, serving as a gating layer (i.e.) Input of ) This indicates element-wise multiplication.
[0072] It should be explained that when the gate value is close to 1, the output is closer to the attention aggregation result. (This indicates reconstruction based on prior knowledge); when the gate is close to 0, the output is closer to the original input. (This indicates that the original low-frequency signal is retained.)
[0073] Then, feature projection can be performed on the high-frequency component sequence to obtain the high-frequency projection result, as follows: ; in, This represents the high-frequency projection result; Represents high-frequency component sequences; Conv(); represents a convolutional layer. This represents the learnable parameters of the convolutional layer, which are optimized during training to achieve the best high-frequency information projection effect.
[0074] The high-frequency projection results and the reference reconstruction results are fused to obtain the target fusion result, as follows: ; in, This indicates a fusion feature containing dual-domain information (i.e., the target fusion result); Concat() represents the concatenation operation. Indicates the reference reconstruction results; Finally, the first reconstruction result can be determined based on the target fusion result, as follows: ; in, LN() represents the first reconstruction result; LN() represents the layer normalization operation, which normalizes the features, stabilizes the training process, and improves model convergence; Linear() represents the linear transformation layer, which normalizes the features... Adjust the channel dimensions and integrate information to output features that are compatible with subsequent modules.
[0075] Thus, by retrieving prior low-frequency patterns across attention levels, and then adaptively fusing the original low-frequency patterns with the prior reconstruction results via gating, while simultaneously introducing high-frequency projection feature constraints, a first reconstruction result is generated. This approach not only improves the standardization of the low-frequency patterns but also preserves image edge textures, avoiding structural distortion. It balances prior universality with the specificity of the original signal, effectively enhancing reconstruction consistency and detail restoration, and providing a high-precision benchmark for subsequent anomaly detection.
[0076] S103. The target decomposition coding data and the first reconstruction result are processed by a cross-domain detail-preserving network to obtain the second reconstruction result.
[0077] In some embodiments, the cross-domain detail-preserving network includes: a spatial domain branch network and a frequency domain branch network; the target decomposition and coding data and the first reconstruction result are processed by the cross-domain detail-preserving network to obtain a second reconstruction result, including: S71. Update the target reconstruction result based on the first reconstruction result and the high-frequency component sequence; S72. Determine the target gradient map corresponding to the image to be detected; S73. The spatial domain feature data is obtained by processing the target gradient map and the updated target reconstruction results through a spatial domain branch network. S74. The updated target reconstruction result is transformed in the frequency domain through a frequency domain branch network to obtain target frequency domain data; the target frequency domain data is processed to obtain frequency domain feature data; the frequency domain feature data is transformed inversely in the frequency domain to obtain target spatial domain data; the spatial structure of the target spatial domain data is aligned with the image to be detected. S75. Based on the cross-domain attention mechanism, spatial domain feature data and target spatial domain data are fused to obtain target fused data; S76. Determine the second residual information based on the target fusion data and the image to be detected; S77. Optimize the updated target reconstruction result based on the second residual information to obtain the second reconstruction result.
[0078] Specifically, the first reconstruction result is the reconstructed low-frequency component sequence optimized by the frequency-aware prior reconstruction module. This reconstructed low-frequency component sequence and the high-frequency component sequence are then re-input into the shared decoder in the recursive frequency domain decomposition encoder to perform multi-level recursive reconstruction. At each recursive level, the shared decoder merges the corresponding reconstructed low-frequency components and high-frequency components, then performs upsampling and spatial structure restoration to generate a new reconstructed image sequence. The image in the target reconstruction result is then updated with the new reconstructed image sequence to obtain the updated target reconstruction result.
[0079] Then, the target gradient map corresponding to the image to be detected can be determined. Specifically, if the image to be detected is color, it is first converted to grayscale. Then, the Sobel operator (or Scharr, Prewitt operator) is used to calculate the gradient components in the horizontal and vertical directions respectively. The gradient magnitude is calculated based on these gradient components to obtain the final target gradient map. Next, the target gradient map and the updated target reconstruction result can be input into the spatial domain branch network to output spatial domain feature data, as follows:
[0080]
[0081] in, Represents spatial domain feature data; Next, the updated target reconstruction results can be transformed in the frequency domain using a frequency domain branch network to obtain the target frequency domain data, as follows: ; in, The target frequency domain data is represented; then, the target frequency domain data is processed to obtain frequency domain feature data, as follows: ; in, Represents frequency domain characteristic data; Perform an inverse frequency domain transform on the frequency domain feature data to obtain the target spatial domain data, as follows:
[0082] in, The target spatial domain data is represented; then, based on a cross-domain attention mechanism, the spatial domain feature data and the target spatial domain data are fused to obtain the target fused data, as follows:
[0083] in, Indicates the target data fusion; Furthermore, the residual between the target fused data and the image to be detected can be calculated to obtain the second residual information, as follows: ; Wherein, Residual represents the second residual information; Finally, the updated target reconstruction result can be optimized based on the second residual information to obtain the second reconstruction result, as follows:
[0084] in, λ represents the second reconstruction result; λ represents the preset residual coefficient.
[0085] Thus, by updating the target reconstruction results and the spatial domain branching process guided by the gradient map, and the frequency domain branching process consisting of frequency domain transformation, feature extraction and inverse transformation, the complementary features of the two domains are achieved. Then, through cross-domain attention fusion of spatial and frequency domain features, the second residual information is generated based on the fused data and the reconstruction results are optimized. This not only strengthens the preservation of spatial details such as image edge texture, but also ensures the standardization of the frequency domain structure. The dual-domain collaboration improves the detail restoration and overall consistency of the reconstruction results, and significantly enhances the robustness of subsequent anomaly detection for the identification of defects of different scales and types.
[0086] S104. The second reconstruction result is processed by the frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected.
[0087] In some embodiments, the second reconstruction result includes: N optimized reconstruction results; each optimized reconstruction result corresponds to a second recursive level; the target anomaly detection result includes: target anomaly score and target anomaly localization map; the second reconstruction result is processed by a frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected, including: S81. Extract spatial domain features and frequency domain features from the N optimized reconstruction results respectively to obtain N spatial domain features and N frequency domain features; each optimized reconstruction result corresponds to one spatial domain feature and one frequency domain feature; S82. Based on the first preset formula and N spatial domain features, perform spatial domain consistency analysis to obtain spatial domain consistency parameters; S83. Based on the second preset formula, the third preset formula and N frequency domain features, perform frequency domain consistency analysis to obtain frequency domain difference parameters and frequency domain stability parameters; S84. Determine the target anomaly score corresponding to the image to be detected based on the spatial domain consistency parameter, frequency domain difference parameter, and frequency domain stability parameter. S85. Based on the target anomaly score, determine the target anomaly localization map corresponding to the image to be detected.
[0088] The second reconstruction result can be denoted as: ={ , ,..., }, This represents the reconstructed image output at the Nth recursive level (i.e., the optimized reconstruction result).
[0089] Specifically, spatial domain features and frequency domain features are extracted from the N optimized reconstruction results, as follows: ; ; in, This represents the frequency domain characteristics of the nth optimized reconstruction result; This represents the spatial domain feature of the nth optimized reconstruction result; This represents the nth optimized reconstruction result among N optimized reconstruction results; n is the result index, an integer greater than 0 and less than or equal to N; using the above formula, N spatial domain features and N frequency domain features can be calculated.
[0090] Next, based on the first preset formula and N spatial domain features, spatial domain consistency analysis is performed to obtain spatial domain consistency parameters; wherein, the first preset formula is as follows: ; in, The spatial distribution of differences between levels (i.e., spatial domain consistency parameter). Then, based on the second preset formula, the third preset formula, and N frequency domain features, frequency domain consistency analysis is performed to obtain frequency domain difference parameters and frequency domain stability parameters; the second preset formula is as follows: ; in, It represents the difference between the frequency domain features of adjacent levels (i.e., the frequency domain difference parameter). The third preset formula is as follows: ; in, This represents the frequency domain stability parameter, which measures frequency stability by calculating the variance of the frequency characteristics at different recursion levels for each spatial location. The physical meaning of variance is that for normal regions, the frequency characteristics remain relatively consistent across all recursion levels, resulting in low variance; for anomalous regions, the frequency characteristics may change significantly between different levels, leading to high variance. This frequency domain stability analysis provides anomaly criteria that complement those in the spatial domain.
[0091] Next, the target anomaly score can be determined based on the spatial domain consistency parameter, frequency domain difference parameter, and frequency domain stability parameter. Specifically, the spatial domain score can be determined first based on the spatial domain consistency parameter. For example, a pre-stored mapping relationship between the consistency parameter and the score can be used to determine the spatial domain score corresponding to the spatial domain consistency parameter. Then, the frequency domain score can be determined based on the frequency domain difference parameter and the frequency domain stability parameter. For example, a first weight corresponding to the frequency domain difference parameter and a second weight corresponding to the frequency domain stability parameter can be determined. A weighted calculation is then performed based on the frequency domain difference parameter, the first weight, the frequency domain stability parameter, and the second weight to obtain the frequency domain score. The sum of the first weight and the second weight is 1; both the first weight and the second weight can be 0.5.
[0092] Then, the target anomaly score corresponding to the image to be detected is determined based on the spatial domain score and the frequency domain score, as follows: ; in, Indicates the target anomaly score; Indicates spatial domain score; γ represents the frequency domain score; γ represents the adaptive fusion coefficient.
[0093] In some embodiments, the process of obtaining the adaptive fusion coefficient γ is as follows: ; ; in, Represents the global spatial domain score; MLP() represents the multilayer perceptron; Finally, based on the target anomaly score, the target anomaly localization map corresponding to the image to be detected is determined. Specifically, the target anomaly score can be mapped to a two-dimensional space that matches the size of the image to be detected to generate an initial anomaly heatmap. Subsequently, the initial anomaly heatmap is thresholded, and areas with scores higher than a preset score threshold are marked as potential anomaly regions. Then, morphological operations (e.g., erosion, dilation) are used to denoise and analyze the connected components of the potential anomaly regions, removing isolated noise points and merging neighboring anomaly regions. Finally, the processed anomaly regions are mapped back to the pixel coordinate space of the image to be detected to generate a target anomaly localization map, which visually identifies the location and extent of the defect.
[0094] Thus, by simultaneously extracting spatial and frequency domain features from multiple reconstruction results, spatial domain consistency analysis, frequency domain difference and stability analysis are carried out respectively. The reconstruction differences are characterized from multiple dimensions such as image texture structure, frequency domain distribution deviation and feature fluctuation degree. Then, multiple parameters are integrated to comprehensively determine the target anomaly score and generate an anomaly location map accordingly. This breaks through the limitations of single-dimensional evaluation, takes into account the integrity of spatial details and the regularity of frequency domain features, improves the objectivity and comprehensiveness of anomaly scoring, and can accurately locate the anomaly location range, effectively enhancing the accuracy and robustness of defect identification and location in complex scenarios.
[0095] In summary, this method first processes the image to be detected using a recursive frequency domain decomposition encoder to obtain target decomposition coding data with complete information dimensions. Then, the frequency-aware prior reconstruction module reconstructs and optimizes the coding data, regularizing normal features and suppressing abnormal interference to obtain the first reconstruction result. Next, the target decomposition coding data and the first reconstruction result are fused through a cross-domain detail preservation network to correct reconstruction bias and preserve image detail features, generating the second reconstruction result. Finally, the frequency-aware cross-recursive detection module combines multi-source data for joint anomaly discrimination. Compared with the traditional single reconstruction error discrimination method, this method can effectively avoid false detection and missed detection problems caused by complex background and texture interference, and significantly improve the accuracy of unsupervised anomaly detection.
[0096] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0097] Based on the same inventive concept, this application also provides an unsupervised anomaly detection device. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more unsupervised anomaly detection device embodiments provided below can be found in the limitations of the unsupervised anomaly detection method above, and will not be repeated here.
[0098] like Figure 5As shown, this application provides an unsupervised anomaly detection device applied to an unsupervised anomaly detection system. The system includes: a recursive frequency domain decomposition encoder, a frequency-aware prior reconstruction module, a cross-domain detail preservation network, and a frequency-aware cross-recursive detection module. The system is trained solely on a normal sample dataset. The unsupervised anomaly detection device 500 includes: The encoding module 501 is used to process the image to be detected through a recursive frequency domain decomposition encoder to obtain target decomposition encoding data; The first reconstruction module 502 is used to process the target decomposition and coding data through the frequency-aware prior reconstruction module to obtain the first reconstruction result; The second reconstruction module 503 is used to process the target decomposition coding data and the first reconstruction result through a cross-domain detail-preserving network to obtain the second reconstruction result; The anomaly detection module 504 is used to process the second reconstruction result through the frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected.
[0099] Each module in the aforementioned unsupervised anomaly detection device 500 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0100] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data related to unsupervised anomaly detection. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the aforementioned unsupervised anomaly detection method.
[0101] In some embodiments, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it performs the steps in the aforementioned unsupervised anomaly detection method. The display unit of the computer device forms a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen; the input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs or touchpads set on the casing of the computer device, or external keyboards, touchpads or mice, etc.
[0102] Those skilled in the art will understand that Figure 6 or Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0103] In some embodiments, a computer device is provided, the computer device including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.
[0104] In some embodiments, such as Figure 8 The diagram shows the internal structure of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the above-described method embodiments.
[0105] In some embodiments, a computer program product is provided, which includes a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0106] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0107] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0109] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An unsupervised anomaly detection method, characterized in that, An unsupervised anomaly detection system is applied, comprising: a recursive frequency domain decomposition encoder, a frequency-aware prior reconstruction module, a cross-domain detail preservation network, and a frequency-aware cross-recursive detection module; the system is trained solely on a normal sample dataset; the method includes: The recursive frequency domain decomposition encoder processes the image to be detected to obtain target decomposition and encoding data. The target decomposition and encoding data are processed by the frequency-aware prior reconstruction module to obtain a first reconstruction result; The target decomposition and encoding data and the first reconstruction result are processed by the cross-domain detail-preserving network to obtain the second reconstruction result; The frequency-aware cross-recursive detection module processes the second reconstruction result to obtain the target anomaly detection result corresponding to the image to be detected.
2. The method according to claim 1, characterized in that, The target decomposition and encoding data includes: target reconstruction result, low-frequency component sequence, and high-frequency component sequence; the recursive frequency domain decomposition encoder includes: recursive compression decomposition unit and recursive reconstruction synthesis unit; The process of processing the image to be detected through the recursive frequency domain decomposition encoder to obtain target decomposition coding data includes: The image to be detected is processed by the recursive compression decomposition unit to obtain a latent representation sequence; The latent representation sequence is processed by the recursive reconstruction synthesis unit to obtain the target reconstruction result, the low-frequency component sequence, and the high-frequency component sequence.
3. The method according to claim 2, characterized in that, The recursive compression decomposition unit includes N first recursive levels connected in sequence; where N is an integer greater than 1. The process of processing the image to be detected through the recursive compression decomposition unit to obtain the latent representation sequence includes: In the first recursive level of the target, the image to be detected is processed to obtain the first high-frequency component and the first low-frequency component; the first recursive level of the target is the first in the N first recursive levels; The first high-frequency component and the first low-frequency component are compressed and encoded respectively to obtain high-frequency feature code and low-frequency feature code; The high-frequency feature encoding and the low-frequency feature encoding are fused based on a cross-level attention mechanism to obtain the first fused data; The first fused data is processed sequentially through N-1 first recursive levels (excluding the target first recursive level) among the N first recursive levels to obtain a potential representation sequence.
4. The method according to claim 3, characterized in that, The recursive compression decomposition unit also includes: a shared encoder and a multilayer perceptron; The step of compressing and encoding the first high-frequency component and the first low-frequency component respectively to obtain high-frequency feature codes and low-frequency feature codes includes: The first high-frequency component is compressed and encoded using the shared encoder to obtain high-frequency feature codes; The first low-frequency component is compressed and encoded using the multilayer perceptron to obtain the low-frequency feature code.
5. The method according to claim 3 or 4, characterized in that, The recursive reconstruction synthesis unit includes N second recursive levels connected in sequence; The process of processing the latent representation sequence through the recursive reconstruction synthesis unit to obtain the target reconstruction result, the low-frequency component sequence, and the high-frequency component sequence includes: The target reconstruction result is obtained by reconstructing the image from deep to shallow layers through the N second recursive levels according to the latent representation sequence; Based on the target reconstruction results, the low-frequency component sequence and the high-frequency component sequence are determined.
6. The method according to claim 5, characterized in that, The recursive reconstruction synthesis unit also includes: a shared decoder and a decoding multilayer perceptron; The process of reconstructing the image from deep to shallow based on the latent representation sequence through the N second recursive levels to obtain the target reconstruction result includes: In the target second recursive level, the high-frequency components in the latent representation sequence are decoded by the shared decoder to obtain high-frequency reconstructed feature data; the target second recursive level is the last in the N second recursive levels; The low-frequency components in the latent representation sequence are inversely projected using the decoding multilayer perceptron to obtain low-frequency reconstructed feature data. The high-frequency reconstruction feature data and the low-frequency reconstruction feature data are fused to obtain the target frequency spectrum; Based on the target frequency spectrum, determine the intermediate reconstructed image; Determine the first residual information between the intermediate reconstructed image and the image to be detected; The intermediate reconstructed image is refined based on the first residual information to obtain the target reconstructed image; The target reconstruction image is processed from deep to shallow through N-1 second recursive levels, excluding the target second recursive level, to obtain the target reconstruction result.
7. The method according to any one of claims 2-4, characterized in that, The frequency-aware prior reconstruction module includes: a learnable prior context library; The process of processing the target decomposition and encoding data through the frequency-aware prior reconstruction module to obtain a first reconstruction result includes: A linear transformation is performed on the low-frequency component sequence to obtain the query vector; A linear transformation is performed on all prior vectors in the learnable prior context library to obtain key vectors and value vectors; The attention weight distribution is determined based on the query vector and the key vector; The attention aggregation result is determined based on the attention weight distribution, the learnable prior context library, and the value vector; Based on the gating mechanism, the low-frequency component sequence and the attention aggregation result are processed to obtain the reference reconstruction result; The high-frequency component sequence is subjected to feature projection to obtain a high-frequency projection result; the dimension of the high-frequency projection result is consistent with the reference reconstruction result. The high-frequency projection result and the reference reconstruction result are fused to obtain the target fusion result; Based on the target fusion result, the first reconstruction result is determined.
8. The method according to any one of claims 2-4, characterized in that, The cross-domain detail-preserving network includes: a spatial domain branch network and a frequency domain branch network; The process of processing the target decomposition and encoding data and the first reconstruction result through the cross-domain detail-preserving network to obtain the second reconstruction result includes: The target reconstruction result is updated based on the first reconstruction result and the high-frequency component sequence; Determine the target gradient map corresponding to the image to be detected; The spatial domain branch network is used to process the target gradient map and the updated target reconstruction result to obtain spatial domain feature data. The updated target reconstruction result is transformed in the frequency domain by the frequency domain branch network to obtain target frequency domain data; the target frequency domain data is processed to obtain frequency domain feature data; the frequency domain feature data is transformed in the frequency domain to obtain target spatial domain data; the spatial structure of the target spatial domain data is aligned with the image to be detected. Based on the cross-domain attention mechanism, the spatial domain feature data and the target spatial domain data are fused to obtain the target fused data; The second residual information is determined based on the target fusion data and the image to be detected; The updated target reconstruction result is optimized based on the second residual information to obtain the second reconstruction result.
9. The method according to claim 5, characterized in that, The second reconstruction result includes: N optimized reconstruction results; each optimized reconstruction result corresponds to a second recursive level; the target anomaly detection result includes: target anomaly score and target anomaly localization map; The step of processing the second reconstruction result through the frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected includes: Spatial domain features and frequency domain features are extracted from the N optimized reconstruction results respectively, resulting in N spatial domain features and N frequency domain features; each optimized reconstruction result corresponds to one spatial domain feature and one frequency domain feature; Based on the first preset formula and the N spatial domain features, spatial domain consistency analysis is performed to obtain spatial domain consistency parameters. Based on the second preset formula, the third preset formula and the N frequency domain features, frequency domain consistency analysis is performed to obtain frequency domain difference parameters and frequency domain stability parameters. The target anomaly score corresponding to the image to be detected is determined based on the spatial domain consistency parameter, the frequency domain difference parameter, and the frequency domain stability parameter. Based on the target anomaly score, the target anomaly localization map corresponding to the image to be detected is determined.
10. An unsupervised anomaly detection device, characterized in that, An unsupervised anomaly detection system is applied, comprising: a recursive frequency domain decomposition encoder, a frequency-aware prior reconstruction module, a cross-domain detail preservation network, and a frequency-aware cross-recursive detection module; the system is trained solely on a normal sample dataset; the apparatus includes: The encoding module is used to process the image to be detected through the recursive frequency domain decomposition encoder to obtain target decomposition encoding data; The first reconstruction module is used to process the target decomposition and coding data through the frequency-aware prior reconstruction module to obtain the first reconstruction result; The second reconstruction module is used to process the target decomposition and encoding data and the first reconstruction result through the cross-domain detail-preserving network to obtain the second reconstruction result; An anomaly detection module is used to process the second reconstruction result through the frequency-aware cross-recursive detection module to obtain the target anomaly detection result corresponding to the image to be detected.
11. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.