Anomaly Image Detection Method, Device and Equipment Based on Image Detection Model
Through feature fusion network and feature processing at multiple scales, the problem of inaccurate abnormal detection caused by incomplete encoding features in the prior art is solved, and the generation of high-quality reconstruction features and more accurate abnormal detection are achieved.
Patent Information
- Application Number
- CN202311756327.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-12-18
AI Technical Summary
The existing image detection method has the incomplete representation of image information by the encoder network, and the reconstruction features generated by the decoder are rough, which makes the abnormal detection results in inaccurate.
The feature fusion network is used to fuse the encoded features of multiple scales, combine the upsampling network, style conversion network and feature decoding network, and generate high-quality reconstruction features for abnormal detection through feature processing of multiple scales.
By fusion of coding features at multiple scales, high-quality reconstruction features are generated, which improves the accuracy and efficiency of abnormal image detection and reduces misjudgment.
Smart Images

Figure CN117710785B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an abnormal image detection method, device and equipment based on an image detection model. Background Art
[0002] The unsupervised abnormal image detection task in the field of computer vision can realize the detection of abnormal images only by training with normal images. Due to the low-cost training method of unsupervised and its important significance in practical applications, abnormal image detection algorithms have gradually received more and more attention and are widely used in industrial defect detection, medical image lesion detection, video anomaly detection and other fields.
[0003] In related technologies, a Reconstruction-based image detection method is used to achieve this goal. Specifically, first, sample images for training the image detection model are collected. The encoder maps the sample images to a high-dimensional latent space to obtain encoded features, and the decoder restores the encoded features to obtain reconstructed features. With the goal of minimizing the difference between the encoded features and the reconstructed features, the model parameters are gradually optimized, thereby obtaining a trained image detection model. We can input an image to be detected, and through the trained image detection model, the abnormal detection of the image can be realized according to the difference between the encoded features and the reconstructed features.
[0004] In the above method, since the existing encoder network usually adopts a CNN (Convolutional Neural Network) model, only the encoded features of a single scale of the input image are extracted. This makes the representation of the encoded features for the image information not complete and comprehensive. In this case of incomplete information, the decoder can only obtain relatively rough reconstructed features. At this time, the difference between the encoded features and the reconstructed features not only comes from the existence of abnormalities, but also from the information loss in the encoding-decoding process. For the abnormal detection of the image to be detected, directly performing abnormal image detection based on the difference between the encoded features and the reconstructed features is likely to introduce misjudgment. That is, the poor effect of obtaining the reconstructed features through the decoder will lead to inaccurate detection results. Summary of the Invention
[0005] Embodiments of this application provide an abnormal image detection method, device and equipment based on an image detection model. The technical solutions provided by the embodiments of this application are as follows:
[0006] According to one aspect of the embodiments of this application, an abnormal image detection method based on an image detection model is provided. The image detection model includes a feature fusion network, an upsampling network, a style conversion network and a feature decoding network; the method includes:
[0007] For the first image to be detected, extract the encoded features of N scales of the first image, where N is an integer greater than 1;
[0008] Fuse the encoded features of the N scales through the feature fusion network to obtain fused features;
[0009] Upsample the fused features through the upsampling network to obtain the upsampled features of the N scales;
[0010] Perform channel transformation on the upsampled features of the N scales through the style conversion network to obtain the adapted features of the N scales, where the channel transformation is channel compression or channel expansion;
[0011] Obtain the reconstructed features of the N scales through the feature decoding network according to the adapted features of the N scales;
[0012] Determine the anomaly detection result of the first image according to the reconstructed features of the N scales and the encoded features of the N scales.
[0013] According to one aspect of the embodiments of the present application, a method for training an image detection model is provided. The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network; the method includes:
[0014] Obtain a sample image for training the image detection model, where there is no abnormal area in the sample image;
[0015] Extract the encoded features of N scales of the sample image, where N is an integer greater than 1;
[0016] Fuse the encoded features of the N scales through the feature fusion network to obtain fused features;
[0017] Upsample the fused features through the upsampling network to obtain the upsampled features of the N scales;
[0018] Perform channel transformation on the upsampled features of the N scales through the style conversion network to obtain the adapted features of the N scales, where the channel transformation is channel compression or channel expansion;
[0019] Obtain the reconstructed features of the N scales through the feature decoding network according to the adapted features of the N scales;
[0020] With the goal of minimizing the difference between the reconstructed features of the N scales and the encoded features of the N scales, adjust the parameters of the image detection model to obtain the trained image detection model.
[0021] According to one aspect of the embodiments of the present application, an abnormal image detection device based on an image detection model is provided. The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network. The device includes:
[0022] An extraction module, configured to extract encoded features of the first image at N scales for the first image to be detected, where N is an integer greater than 1;
[0023] A first obtaining module, configured to fuse the encoded features of the N scales through the feature fusion network to obtain a fused feature;
[0024] A second obtaining module, configured to upsample the fused feature through the upsampling network to obtain upsampled features of the N scales;
[0025] A third obtaining module, configured to perform channel transformation on the upsampled features of the N scales through the style conversion network to obtain adapted features of the N scales, where the channel transformation is channel compression or channel expansion;
[0026] A fourth obtaining module, configured to obtain reconstructed features of the N scales according to the adapted features of the N scales through the feature decoding network;
[0027] A determination module, configured to determine an abnormal detection result of the first image according to the reconstructed features of the N scales and the encoded features of the N scales.
[0028] According to one aspect of the embodiments of the present application, a training device for an image detection model is provided. The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network. The device includes:
[0029] An acquisition module, configured to acquire a sample image for training the image detection model, where there is no abnormal area in the sample image;
[0030] An extraction module, configured to extract encoded features of the sample image at N scales, where N is an integer greater than 1;
[0031] A first obtaining module, configured to fuse the encoded features of the N scales through the feature fusion network to obtain a fused feature;
[0032] A second obtaining module, configured to upsample the fused feature through the upsampling network to obtain upsampled features of the N scales;
[0033] A third obtaining module, configured to perform channel transformation on the upsampled features of the N scales through the style conversion network to obtain the adapted features of the N scales, where the channel transformation is channel compression or channel expansion;
[0034] A fourth obtaining module, configured to obtain the reconstructed features of the N scales through the feature decoding network according to the adapted features of the N scales;
[0035] An adjustment module, configured to adjust the parameters of the image detection model with the goal of minimizing the difference between the reconstructed features of the N scales and the encoded features of the N scales, so as to obtain a trained image detection model.
[0036] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned abnormal image detection method based on an image detection model, or the above-mentioned training method of the image detection model.
[0037] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided. A computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the above-mentioned abnormal image detection method based on an image detection model, or the above-mentioned training method of the image detection model.
[0038] According to one aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes a computer program. The computer program is stored in a computer-readable storage medium, and a processor reads and executes the computer program from the computer-readable storage medium to implement the above-mentioned abnormal image detection method based on an image detection model, or the above-mentioned training method of the image detection model.
[0039] The technical solutions provided by the embodiments of the present application at least include the following beneficial effects:
[0040] By extracting the encoded features of the input image at multiple scales, the feature information of multiple scales of the input image is effectively retained. The feature fusion network fuses the encoded features of multiple scales to obtain fused features. The upsampling network upsamples the fused features to obtain upsampled features of multiple scales. Further, the style conversion network transforms the channel dimension of the upsampled features of different scales to obtain adapted features of multiple scales, which can increase the diversity of feature expression. According to the adapted features of different scales through the feature decoding network, high-quality reconstruction features of different scales can be obtained. On the one hand, the above network proposed in this application only contains convolutional operations, which is simple and efficient. On the other hand, through the fusion and utilization of encoded features of multiple scales, high-quality reconstruction features of multiple scales can be obtained, thereby achieving more accurate and efficient abnormal image detection. Description of the Drawings
[0041] Figure 1 is a schematic diagram of the implementation environment of the solution provided by an embodiment of this application;
[0042] Figure 2 is a flowchart of an abnormal image detection method based on an image detection model provided by an embodiment of this application;
[0043] Figure 3 is a schematic diagram of an image detection model provided by an embodiment of this application;
[0044] Figure 4 is a schematic diagram of an abnormal image provided by an embodiment of this application;
[0045] Figure 5 is a schematic diagram of an abnormal image provided by another embodiment of this application;
[0046] Figure 6 is a schematic diagram of a training method of an image detection model provided by an embodiment of this application;
[0047] Figure 7 is a schematic diagram of the comparison of experimental results of different resolutions provided by an embodiment of this application;
[0048] Figure 8 is a schematic diagram of the comparison of experimental results of different training rounds provided by an embodiment of this application;
[0049] Figure 9 is a block diagram of an abnormal image detection device based on an image detection model provided by an embodiment of this application;
[0050] Figure 10 is a block diagram of a training device of an image detection model provided by an embodiment of this application;
[0051] Figure 11It is a structural block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0052] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the implementation manners of the present application in detail with reference to the accompanying drawings.
[0053] Artificial Intelligence (AI for short) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.
[0054] Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0055] Machine Learning (ML for short) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.
[0056] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, AIGC (Artificial Intelligence Generated Content), conversational interaction, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0057] The solution provided in the embodiments of this application relates to machine learning and deep learning technologies of artificial intelligence, and will be specifically described through the following embodiments.
[0058] Please refer to Figure 1 , which shows a schematic diagram of the solution implementation environment provided by an embodiment of this application. The solution implementation environment may include a model training device 110 and a model using device 120.
[0059] The model training device 110 may be an electronic device such as a mobile phone, a desktop computer, a tablet computer, a laptop computer, a vehicle-mounted terminal, a server, a smart robot, a smart TV, a multimedia playback device, etc., or some other electronic device with strong computing power. This application does not make any limitations in this regard. The model training device 110 is used to train an image detection model.
[0060] In the embodiments of this application, the image detection model is a deep neural network model. Optionally, the model training device 110 may adopt a machine learning method to train the image detection model to make it have better performance. Optionally, the training process of the image detection model is as follows (only a brief description here, and the specific training process is shown in the following embodiments): Obtain sample images for training the image detection model, extract the encoded features of N scales of the sample images, fuse the encoded features of N scales through a feature fusion network to obtain fused features, upsample the fused features through an upsampling network to obtain upsampled features of N scales, perform channel transformation on the upsampled features of N scales through a style conversion network to obtain adapted features of N scales, and obtain reconstructed features of N scales through a feature decoding network according to the adapted features of N scales. With the goal of minimizing the difference between the reconstructed features of N scales and the encoded features of N scales, adjust the parameters of the image detection model to obtain the trained image detection model.
[0061] The model-using device 120 can be an electronic device such as a mobile phone, a desktop computer, a tablet computer, a laptop computer, a vehicle-mounted terminal, a server, an intelligent robot, an intelligent TV, a multimedia playback device, or some other electronic device with strong computing capabilities. The present application does not limit this. The model-using device 120 can adopt the trained image detection model to perform anomaly detection on the image to be detected.
[0062] The model-training device 110 and the model-using device 120 can be two independent devices or the same device.
[0063] In the method provided by the embodiments of the present application, the execution subject of each step can be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. Among them, when the computer device is a server, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The computer device can be Figure 1 the model-training device 110 in
[0064] In some embodiments, the technical solution proposed in the present application can be applied to the traffic detection scenario. Exemplarily, it can help monitor abnormal situations at traffic intersections or traffic sections (such as illegal driving of key vehicles, traffic accidents, etc.), effectively improving the traffic supervision ability. It can also be applied to the industrial product anomaly detection scenario. Exemplarily, in the industrial production process, it can perform anomaly detection on product images, immediately discover product quality problems, and help improve production efficiency and product quality stability. It can also be applied to the medical image lesion detection scenario, which can effectively assist doctors in quickly and accurately discovering lesions in medical images such as human organs. The present application does not limit the application scenarios of the solution.
[0065] Please refer to Figure 2 which shows a flowchart of an abnormal image detection method based on an image detection model provided by an embodiment of the present application. The execution subject of each step of this method can be a computer device. For example, the computer device can be Figure 1 the model-using device 120 in the solution implementation environment shown. This method can include at least one of the following steps 210 to 260.
[0066] Step 210, for the first image to be detected, extract the encoded features of N scales of the first image, where N is an integer greater than 1.
[0067] The first image to be detected refers to the original image input into the image detection model for anomaly detection. Anomaly detection refers to a technique or method for a given image dataset, which constructs a model to learn the feature information of normal image samples and then detects abnormal images based on the deviation of the data distribution. Scale refers to resolution. N scales refer to N different scales.
[0068] In some possible implementation manners, the first image to be detected can be obtained through local selection and input by the user, or can be obtained by downloading the retrieved image from the network platform. The first image can be any image to be detected for anomalies, and this application does not limit the manner of obtaining the image.
[0069] In some embodiments, the features of N scales of the first image can be extracted by a pre-trained model first, and then the features of the above N scales are respectively downsampled to obtain the encoded features of N scales of the first image. Exemplarily, the pre-trained model can be the WideResNet50 model pre-trained using ImageNet-1K, or other pre-trained models that output the features of multiple scales of the original image by inputting the original image, and this application does not limit this.
[0070] The features of N scales of the first image are extracted by the above pre-trained model. Exemplarily, the features of 3 scales of the first image are extracted by the above pre-trained model, such as the features of the 512×512 scale, the features of the 256×256 scale, and the features of the 128×128 scale. Then the features of the 512×512 scale are downsampled to obtain the encoded features of the 256×256 scale; the features of the 256×256 scale are downsampled to obtain the encoded features of the 128×128 scale; the features of the 128×128 scale are downsampled to obtain the encoded features of the 64×64 scale.
[0071] Step 220, fuse the encoded features of N scales through a feature fusion network to obtain a fused feature.
[0072] Fusion refers to combining the encoded features of different scales to obtain more comprehensive or higher-quality features. The fused feature refers to a new feature vector obtained by fusing the encoded features from different scales.
[0073] In some embodiments, the encoded features of N scales are converted into encoded features of the same scale through a feature fusion network to obtain N encoded features of the same scale; the N encoded features of the same scale are concatenated in channels to obtain a fused feature.
[0074] The feature fusion network maps the encoded features of different scales into the feature space of the same scale through the spatial transformation of the convolutional layer, so that the features of different scales can be aligned in the same feature space.
[0075] The channel refers to the channel dimension. The encoded features of N scales correspond to N channel dimensions. Among them, the channel dimensions corresponding to the encoded features of N scales can be all the same, or at least two of them can be different. As Figure 3 shown, it shows the encoded features of 3 different scales. Assume that the 3 different scales are 256×256, 128×128, and 64×64 respectively. The channel dimension corresponding to the encoded feature 31 of the 256×256 scale is K1, the channel dimension corresponding to the encoded feature 32 of the 128×128 scale is K2, and the channel dimension corresponding to the encoded feature 33 of the 64×64 scale is K3. K1, K2, and K3 are integers greater than 1. K1, K2, and K3 can be all the same, or K1, K2, and K3 can be all different, or K1 and K2 are the same but different from K3, or K1 and K3 are the same but different from K2, or K2 and K3 are the same but different from K1. This application does not make any limitations on this.
[0076] Channel concatenation means connecting the encoded features of the same scale after spatial alignment in the channel dimension to form a higher-dimensional fused feature. Exemplarily, the channel dimension corresponding to the encoded feature of the 256×256 scale is 64, the channel dimension corresponding to the encoded feature of the 128×128 scale is 64, and the channel dimension corresponding to the encoded feature of the 64×64 scale is 64. After downsampling, the encoded features of the above 3 scales can all be converted into encoded features of the 32×32 scale. Assume that the number of channels of the encoded features remains unchanged. Then, after channel concatenation, a fused feature with a scale of 32×32 can be obtained, and the number of channels of this fused feature is 192.
[0077] The above method can fuse the features of different levels by fusing the encoded features of different scales, so as to obtain a more comprehensive and richer feature representation.
[0078] In some embodiments, the encoded features of the larger scale among the N scales can be downsampled to the smallest scale, and the smallest scale refers to the smallest scale among the N scales of the encoded features. After downsampling, all the encoded features have the same scale size. Please refer to Figure 3 where subfigure (a) is a feature transformation diagram for obtaining reconstructed features of N scales based on the encoded features of N scales, and subfigure (b) is a specific implementation process for obtaining reconstructed features based on the adaptive features. In subfigure (a), F f is the fused feature, and F I is the encoded feature. As Figure 3As shown, it shows encoded features at three different scales. Assume that the encoded features at the three different scales are the encoded feature 31 at the scale of 256×256, the encoded feature 32 at the scale of 128×128, and the encoded feature 33 at the scale of 64×64. The encoded features at the scales of 256×256 and 128×128 can be respectively converted to the encoded features at the scale of 64×64 through downsampling. Then, the three encoded features at the scale of 64×64 are concatenated in channels to obtain the fused feature F f .
[0079] In some embodiments, the encoded features at a larger scale among the N scales can be downsampled to a target scale, where the target scale is smaller than the smallest scale among the N scales of the encoded features. As shown in the above example, assume that the target scale here is 32×32. The encoded features at the scales of 256×256, 128×128, and 64×64 can be respectively converted to the encoded features at the scale of 32×32 through downsampling. Then, the three encoded features at the scale of 32×32 are concatenated in channels to obtain the fused feature F f .
[0080] In some embodiments, after obtaining the fused feature, the fused feature can also be subjected to feature transformation through at least one Bottleneck layer to obtain the transformed fused feature. Among them, the transformed fused feature has the same scale and number of channels as the fused feature, and the transformed fused feature has a different numerical representation from the fused feature. The transformed fused feature is used to perform upsampling in step 230 below to obtain the upsampled features at N scales
[0081] The transformed fused feature has a different numerical representation from the fused feature, which can be achieved through different convolutional kernels or functions. For example, convolutional kernels with different weight parameters can be used, or a non-linear function such as ReLU (Rectified Linear Unit) can be applied to introduce non-linear transformation, thereby changing the representation of the features
[0082] In the above method, by performing feature transformation on the fused feature through the Bottleneck layer, the ability to represent abstract features in the fused feature can be enhanced, and the learning ability of the model for complex features can be improved
[0083] Step 230, perform upsampling on the fused feature through an upsampling network to obtain the upsampled features at N scales
[0084] An upsampling network refers to a neural network structure used to upsample low-resolution features to high-resolution to obtain more refined feature information
[0085] In some embodiments, the upsampling network includes N upsampling sub-networks; for the i-th upsampling sub-network among the N upsampling sub-networks, the input data of the i-th upsampling sub-network is upsampled through the i-th upsampling sub-network to obtain the upsampling feature output by the i-th upsampling sub-network, where i is a positive integer less than or equal to N. Among them, when i equals 1, the input data of the i-th upsampling sub-network is the fusion feature, and when i is greater than 1, the input data of the i-th upsampling sub-network is the upsampling feature output by the (i - 1)-th upsampling sub-network; the upsampling features output by the N upsampling sub-networks are determined as the upsampling features of N scales.
[0086] Exemplarily, as Figure 3 shown, it shows upsampling features of 3 different scales, F R refers to the upsampling feature. Assuming that the upsampling features of 3 different scales are the upsampling feature 34 of 64×64 scale, the upsampling feature 35 of 128×128 scale, and the upsampling feature 36 of 256×256 scale respectively. The fusion feature F f can be converted to the upsampling feature of 64×64 scale through upsampling respectively. Then, the upsampling feature of 64×64 scale is upsampled to obtain the upsampling feature of 128×128 scale. Then, the upsampling feature of 128×128 scale is upsampled to obtain the upsampling feature of 256×256 scale.
[0087] In the above method, by using the upsampling sub-networks to perform upsampling step by step, the refinement and amplification of the fusion feature can be realized. Exemplarily, assuming that the input data is the fusion feature of 64×64 scale, through the upsampling operation of the first upsampling sub-network, the upsampling feature of 128×128 scale can be obtained; then, through the upsampling operation of the second upsampling sub-network, the upsampling feature of 256×256 scale is obtained. These upsampling features with different scales and increasingly high resolutions provide multiple scales and more detailed feature information, which helps to improve the model's understanding and perception of images.
[0088] In some embodiments, the i-th upsampling sub-network includes an upsampling convolutional layer and at least one ordinary convolutional layer; the input data of the i-th upsampling sub-network is upsampled through the upsampling convolutional layer to obtain the first intermediate feature; the first intermediate feature is convolved through at least one ordinary convolutional layer to obtain the upsampling feature output by the i-th upsampling sub-network.
[0089] The first intermediate feature refers to the feature representation obtained from the encoded feature of a smaller scale through the upsampling convolutional layer in the intermediate process. Exemplarily, the upsampling feature of 64×64 scale can be upsampled through the upsampling convolution to obtain the upsampling feature of 128×128 scale, and the upsampling feature of 128×128 scale here is used as the first intermediate feature.
[0090] The upsampling convolutional layer is used for scale expansion. The upsampling features of a smaller scale can be transformed into upsampling features of a larger scale through upsampling convolution. The ordinary convolutional layer is used for channel dimension transformation. The upsampling features obtained by upsampling convolution can be subjected to channel dimension transformation through the ordinary convolutional layer, where the channel dimension transformation includes channel compression or channel expansion. Exemplarily, the first intermediate features obtained above, such as the upsampling features of the 128×128 scale, can be subjected to channel dimension transformation to obtain the upsampling features after channel dimension transformation.
[0091] Through the upsampling convolutional layer and the ordinary convolutional layer in the upsampling subnet, the above method can achieve scale expansion and channel dimension transformation of the input data. In this way, richer feature expressions can be obtained, which helps to improve the model's analysis and understanding of images.
[0092] Step 240: Perform channel transformation on the upsampling features of N scales through the style conversion network to obtain the adaptive features of N scales, where the channel transformation is channel compression or channel expansion.
[0093] Channel compression refers to reducing the channel dimension. Exemplarily, the upsampling features with a channel dimension of 256 can be compressed into adaptive features with a channel dimension of 128. This can reduce feature redundancy and improve computational and storage efficiency. Channel expansion refers to increasing the channel dimension. Exemplarily, the upsampling features with a channel dimension of 128 can be expanded into adaptive features with a channel dimension of 256. This can increase the richness of feature expression.
[0094] Through appropriate channel compression and expansion operations, the network structure can be better adjusted to make the features adapt to the requirements of the current task in the channel dimension, ensuring the efficiency and accuracy of the subsequent abnormal image detection task.
[0095] In some embodiments, for the i-th scale among the N scales, perform convolutional processing on the upsampling features of the i-th scale in the channel dimension through the style conversion network to obtain the adaptive features of the i-th scale, where the convolutional processing in the channel dimension is used to achieve channel compression or channel expansion, and i is a positive integer less than or equal to N.
[0096] Performing convolution processing in the channel dimension means achieving channel compression or channel expansion by changing the number of channels of the convolution kernel during the convolution process. In the convolution operation, element-wise multiplication is performed between each channel of the convolution kernel and the corresponding channel of the input upsampled feature, and the results are summed to obtain a single channel of the output. In the convolution processing in the channel dimension, we can control the number of output channels by changing the number of channels of the convolution kernel. Specifically, when the number of channels of the convolution kernel in the channel dimension is less than the number of channels of the input upsampled feature, channel compression can be achieved, reducing the original number of channels; when the number of channels of the convolution kernel in the channel dimension is greater than the number of channels of the input upsampled feature, channel expansion can be achieved, increasing the original number of channels. In this way, we can flexibly adjust and optimize the channel dimension of the upsampled feature to meet the requirements of different tasks.
[0097] As Figure 3 shown, F S refers to the adapted feature, and the channel dimension transformation can be performed on the upsampled feature 34 at the 64×64 scale to obtain its corresponding adapted feature 37. The channel dimension transformation can be performed on the upsampled feature 35 at the 128×128 scale to obtain its corresponding adapted feature 38. The channel dimension transformation can be performed on the upsampled feature 36 at the 256×256 scale to obtain its corresponding adapted feature 39.
[0098] Step 250: Obtain the reconstructed features at N scales through the feature decoding network based on the adapted features at N scales.
[0099] The reconstructed feature refers to the feature representation that has the same scale as the original encoded feature after a specific decoding process based on the original encoded feature. Exemplarily, the encoded feature at the 64×64 scale corresponds to the reconstructed feature at the 64×64 scale.
[0100] In some embodiments, for the i-th scale among the N scales, the reconstructed feature at the i-th scale is obtained through the feature decoding network based on the adapted feature at the i-th scale and the initial recovery feature at the i-th scale, where i is a positive integer less than or equal to N; among them, when i is equal to 1, the initial recovery feature at the i-th scale is a preset fixed feature, and when i is greater than 1, the initial recovery feature at the i-th scale is the reconstructed feature at the i - 1-th scale.
[0101] As Figure 3 shown, F O refers to the reconstructed feature, F Const is the fixed feature. In subfigure (b), the reconstructed feature at the i-th scale is obtained through the adapted feature at the i-th scale and the initial recovery feature at the i-th scale where, when i is 1, is equal to the fixed feature F Const。
[0102] Exemplarily, as Figure 3 shown, the first scale is 64×64. The reconstruction feature 310 of the 64×64 scale can be obtained according to the adaptation feature 37 and the fixed feature F of the 64×64 scale Const wherein the fixed feature F Const is randomly generated, with a scale of 64×64 and the same number of channels as the adaptation feature of this 64×64 scale. To obtain the reconstruction feature of the 128×128 scale, first, the reconstruction feature of the 64×64 scale is upsampled to obtain the initial restored feature of the 128×128 scale, as shown in subfigure (b). Then, the reconstruction feature 311 of the 128×128 scale is obtained according to the adaptation feature 38 of the 128×128 scale and the initial restored feature of this 128×128 scale. The reconstruction feature 312 of the 256×256 scale can be obtained according to the same method.
[0103] According to the above method, through the feature decoding network, according to the adaptation features and the initial restored features of multiple scales, multiple scales of reconstruction features can be generated accordingly.
[0104] In some embodiments, the feature decoding network includes M feature decoding subnets, where M is a positive integer; for the j-th feature decoding subnet among the M feature decoding subnets, through the j-th feature decoding subnet, the input data of the j-th feature decoding subnet is processed to obtain the restored feature output by the j-th feature decoding subnet, where j is a positive integer less than or equal to M. Among them, when j equals 1, the input data of the j-th feature decoding subnet includes the adaptation feature of the i-th scale and the initial restored feature of the i-th scale. When j is greater than 1, the input data of the j-th feature decoding subnet includes the adaptation feature of the i-th scale and the restored feature output by the (j - 1)-th feature decoding subnet; the restored feature output by the M-th feature decoding subnet among the M feature decoding subnets is determined as the reconstruction feature of the i-th scale.
[0105] The restored feature refers to the intermediate feature representation generated by the feature decoding subnet through a step-by-step decoding and reconstruction process. It restores some of the detail information of the original image, but has less information content than the encoded feature.
[0106] For the solution of the reconstructed features at each of the N scales (such as the i-th scale), M feature decoding subnets need to be passed through. The output result of each feature decoding subnet is the restored feature, that is, the feature that gradually restores some details of the original image before encoding. The restored feature output by the M-th feature decoding subnet is used as the reconstructed feature that finally fully restores the detail information of the original image. That is to say, for the encoded features at each scale, by concatenating M feature decoding subnets, the details of the original image are gradually restored in reverse from the adapted features. The output of each layer of the subnet contains some restored detail information, that is, the restored feature. When reaching the M-th subnet, the image detail information is fully restored, and the output at this time is the final reconstructed feature.
[0107] In the above method, the reconstructed features are obtained through M feature decoding subnets, which can improve the accuracy and stability of image reconstruction.
[0108] In some embodiments, through the j-th feature decoding subnet, two different linear processes are performed on the adapted features at the i-th scale. As shown in subfigure (b) of Figure 3, through linear layer 1, the second intermediate feature is obtained. Through linear layer 2, the third intermediate feature is obtained. Subtract the mean value of each value of the restored feature included in the input data of the j-th feature decoding subnet from each value of the restored feature to obtain the fourth intermediate feature; multiply the corresponding values in the second intermediate feature and the fourth intermediate feature to obtain the fifth intermediate feature; divide the fifth intermediate feature by the standard deviation of each value of the restored feature to obtain the sixth intermediate feature; add the corresponding values in the sixth intermediate feature and the third intermediate feature to obtain the restored feature output by the j-th feature decoding subnet.
[0109] Please refer to Formula 1:
[0110]
[0111] Where μ is the mean value and σ is the variance. and are the second intermediate feature and the third intermediate feature respectively after two different linear processes. is the restored feature obtained by passing the i-th scale through the j-th feature decoding subnet. The fourth intermediate feature refers to The fifth intermediate feature refers to The sixth intermediate feature refers to
[0112] In the above method, by introducing two linear layers, these two linear layers are two independent linear layers, and each linear layer has appropriate weights and biases, which can further optimize the representation ability of the restored features.
[0113] Step 260: Determine the anomaly detection result of the first image based on the reconstruction features and the encoded features at N scales.
[0114] In some embodiments, for the i-th scale among the N scales, a difference image at the i-th scale is obtained according to the difference between the values at the corresponding positions in the reconstruction feature and the encoded feature at the i-th scale. The value of each pixel in the difference image at the i-th scale is the difference between a group of values at the corresponding positions in the reconstruction feature and the encoded feature at the i-th scale, where i is a positive integer less than or equal to N. The N difference images are converted to the scale of the first image to obtain N difference images of the same scale. The N difference images of the same scale are fused to obtain a final difference image. Based on the final difference image, the anomaly detection result of the first image is determined.
[0115] The reconstruction feature and the encoded feature at the i-th scale can be regarded as two images with the same scale, i.e., the same resolution. Each pixel position corresponds to a vector. If the number of channels of the reconstruction feature and the encoded feature at the i-th scale is 64, then each pixel position of the reconstruction feature and the encoded feature corresponds to a 64-dimensional vector. Then, the numerical difference between the two images at the corresponding positions is calculated pixel by pixel. This difference forms a new difference image. The difference image highlights the differences in the values of the reconstruction feature and the encoded feature at each corresponding pixel position. The value of each pixel in the difference image reflects the degree of difference in the direction of the pixel value in the reconstruction feature and the pixel value in the encoded feature in the vector space at that pixel position. If the two are highly consistent and the vector gap is small, the pixel value in the difference image is also small. If there is a large deviation, a large vector distance is formed, and the corresponding pixel value in the difference image is higher. By analyzing the size of the difference image, it can be detected which pixel regions have large losses in the encoding-decoding process, thereby reflecting abnormal situations.
[0116] The above method can effectively detect abnormal situations in an image by analyzing and fusing the differences between the reconstruction features and the encoded features.
[0117] In some embodiments, this difference can be quantified by the cosine distance between each pixel position. Other distance metric methods can also be used for quantification, such as Euclidean distance, Manhattan distance, or Chebyshev distance, which are not limited in this application.
[0118] Exemplarily, as Figure 3As shown, the difference between the reconstructed features at the 64×64 scale and the corresponding position values in the encoded features at the 64×64 scale can be calculated to obtain a difference image at the 64×64 scale. According to the same method, difference images at the 128×128 scale and the 256×256 scale can be obtained.
[0119] In some embodiments, after obtaining the difference images at N scales respectively, the difference images at N scales need to be converted into the scale of the first image respectively. The difference images at N scales of the same scale are fused to obtain a final difference image. The scale of the first image usually refers to the resolution scale of the original input image, such as 512×512. The difference images at other smaller scales can be gradually enlarged to the 512×512 scale by upsampling. In this way, the difference information at different scales is mapped into a unified scale coordinate space, and the N difference images can be fused by simple operations such as simple superposition. Finally, a difference image integrating the difference information of each scale is output. The fused difference image is the final difference image, and the abnormal situation of the first image can be determined based on this final difference image.
[0120] In some embodiments, the values at the corresponding positions in the N difference images at the same scale are added to obtain the final difference image; alternatively, the values at the corresponding positions in the N difference images at the same scale are averaged to obtain the final difference image.
[0121] For the N difference images at the same scale, at each pixel position, the corresponding pixel values in the N scale results are directly added, which can enhance the suspected abnormal regions that are larger in multiple scale results. It is also possible to average the values at the corresponding positions in the N difference images at the same scale. The principle is similar to summation, except that the average value of the N scale results is output finally. This application does not limit this.
[0122] Through the above method, by fusing the difference images at multiple scales, the stability and accuracy of anomaly detection can be improved, and it is easier to identify and locate abnormal regions.
[0123] In some embodiments, if there are abnormal pixels in the final difference image, it is determined that the first image is an abnormal image. An abnormal pixel refers to a pixel whose value belongs to a set value range; based on the abnormal pixels, the abnormal region in the first image is determined.
[0124] A threshold can be set. When a certain pixel value in the difference image exceeds this threshold, it is determined as an abnormal pixel. The threshold can be determined by statistically analyzing the difference distribution of normal samples. For example, it is set within the range of P times the standard deviation above and below the mean of normal samples, where P is an integer greater than 1.
[0125] If there are abnormal pixels in the difference image, directly determine the corresponding original image as an abnormal image. Based on the position coordinates of the pixels determined to be abnormal, delimit the connected region where the abnormal pixels are located to form an abnormal region, and perform abnormal marking on the first image according to the determined abnormal region to obtain the first image after abnormal marking. As Figure 4 and Figure 5 shown, they respectively show the abnormal maps obtained by the GT (Ground Truth, annotation value), RD (Recursive Deep Learning), UniAD (Unified Anomaly Detection), and InvAD (Inversion Anomaly Detection, anomaly detection based on feature inversion) methods according to the input image, where InvAD is the method proposed in this application.
[0126] Among them, the area enclosed by the dashed box in each abnormal map is the abnormal region, and there is an obvious abnormal region 41 in the input image. It can be seen from the result of InvAD that this method correctly detects and marks the abnormal region, that is, region 42 in the figure. This verifies that the technical solution proposed in this application can effectively identify the abnormal part in the image. Exemplarily, there is an obvious abnormal region 43 in the input image. It can be seen from the result of InvAD that this method correctly detects and marks the abnormal region, that is, region 44 in the figure. Exemplarily, there is an obvious abnormal region 45 in the input image. It can be seen from the result of InvAD that this method correctly detects and marks the abnormal region, that is, region 46 in the figure. Exemplarily, as Figure 5 shown, there is an obvious abnormal region 51 in the input image. It can be seen from the result of InvAD that this method correctly detects and marks the abnormal region, that is, region 52 in the figure.
[0127] In some embodiments, post-processing can be performed on the determined abnormal region, such as smoothing, morphological operations, etc., which can eliminate the discrete scattered points caused by noise being misjudged as abnormal regions.
[0128] In some embodiments, it is also possible to first mark the respective abnormal candidate regions according to N difference images of the same scale, and finally determine the abnormal region in the first image by combining the abnormal candidate regions corresponding to the N difference images.
[0129] The technical solution proposed in this application effectively retains the feature information of the input image at multiple scales by extracting the encoded features of the input image at multiple scales. The feature fusion network fuses the encoded features at multiple scales to obtain fused features. The upsampling network upsamples the fused features to obtain upsampled features at multiple scales. Further, the style conversion network transforms the channel dimension of the upsampled features at different scales to obtain adapted features at multiple scales, which can increase the diversity of feature expression. According to the adapted features at different scales, the feature decoding network can obtain high-quality reconstructed features at different scales. On the one hand, the above network proposed in this application only includes convolutional operations, which are simple and efficient. On the other hand, by fusing and utilizing the encoded features at multiple scales, high-quality reconstructed features at multiple scales can be obtained, thereby achieving more accurate and efficient abnormal image detection.
[0130] The above embodiments introduced the abnormal image detection method based on the image detection model. Next, the training process of the image detection model will be introduced through embodiments. For the application and training of the image detection model, the two are related. Details not described in detail in one side of the embodiments can be referred to the description in the other side of the embodiments.
[0131] Please refer to Figure 6 , which shows a flowchart of the training method of the image detection model provided by an embodiment of the present application. The execution subject of each step of this method can be a computer device. For example, this computer device can be the model training device 110 in the solution implementation environment shown in Figure 1 . This method may include at least one of the following steps 610 to 670.
[0132] Step 610, obtain a sample image for training the image detection model, and there is no abnormal area in the sample image.
[0133] The sample image refers to a normal image sample for training the image detection model. There are no abnormal areas such as damage, defects, and stains in it. All contents in the sample image are normal. These normal samples can come from the passed product images taken by the quality inspection cameras on the production line, or from the abnormal-free sample library screened by manual annotation. As long as the image content is intact and meets the expected normal situation, it can be used as a sample image for training to provide the model with a normal discrimination benchmark. Only learning the distribution characteristics of normal image samples can avoid the model from "remembering" specific abnormal situations in the training samples, thereby improving the ability to detect unknown new types of defects.
[0134] To improve the model's detection effect on abnormal images, we introduce the concept of GAN Inversion (Generative Adversarial Network Inversion), aiming to achieve more accurate abnormal region localization by restoring high-quality reconstructed features. The image detection model proposed in this application adopts a series of key network components, including a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network. For the specific implementation methods of each network, please refer to the following text.
[0135] Step 620: Extract the encoded features of the sample image at N scales, where N is an integer greater than 1.
[0136] In some embodiments, the N-scale features of the sample image can be extracted first through a pre-trained model. Exemplarily, the pre-trained model can be a WideResNet50 model pre-trained using ImageNet-1K, or other pre-trained models that output multiple-scale features of the original image by inputting the original image. This application does not make any limitations in this regard.
[0137] Step 630: Fuse the encoded features at N scales through a feature fusion network to obtain a fused feature.
[0138] In some embodiments, the encoded features at N scales are converted into encoded features of the same scale through a feature fusion network to obtain N encoded features of the same scale; the N encoded features of the same scale are concatenated in channels to obtain a fused feature.
[0139] In some embodiments, after obtaining the fused feature, the fused feature can also be subjected to feature transformation through at least one Bottleneck layer to obtain a transformed fused feature, where the transformed fused feature has the same scale and number of channels as the fused feature, and the transformed fused feature has a different numerical representation from the fused feature. The transformed fused feature is used to perform upsampling in the following step 640 to obtain N-scale upsampled features.
[0140] Step 640: Upsample the fused feature through an upsampling network to obtain N-scale upsampled features.
[0141] In some embodiments, the upsampling network includes N upsampling sub-networks; for the i-th upsampling sub-network among the N upsampling sub-networks, the input data of the i-th upsampling sub-network is upsampled through the i-th upsampling sub-network to obtain the upsampling feature output by the i-th upsampling sub-network, where i is a positive integer less than or equal to N. Among them, when i equals 1, the input data of the i-th upsampling sub-network is the fused feature; when i is greater than 1, the input data of the i-th upsampling sub-network is the upsampling feature output by the (i - 1)-th upsampling sub-network; the upsampling features output by the N upsampling sub-networks are determined as the upsampling features of N scales.
[0142] In some embodiments, the i-th upsampling sub-network includes an upsampling convolutional layer and at least one ordinary convolutional layer; the input data of the i-th upsampling sub-network is upsampled through the upsampling convolutional layer to obtain a first intermediate feature; the first intermediate feature is convolved through at least one ordinary convolutional layer to obtain the upsampling feature output by the i-th upsampling sub-network.
[0143] Step 650, the upsampling features of N scales are subjected to channel transformation through the style conversion network to obtain the adapted features of N scales, where the channel transformation is channel compression or channel expansion.
[0144] In some embodiments, for the i-th scale among the N scales, the upsampling feature of the i-th scale is subjected to convolutional processing in the channel dimension through the style conversion network to obtain the adapted feature of the i-th scale, where the convolutional processing in the channel dimension is used to implement channel compression or channel expansion, and i is a positive integer less than or equal to N.
[0145] Step 660, the reconstruction features of N scales are obtained through the feature decoding network according to the adapted features of N scales.
[0146] In some embodiments, for the i-th scale among the N scales, the reconstruction feature of the i-th scale is obtained through the feature decoding network according to the adapted feature of the i-th scale and the initial recovery feature of the i-th scale, where i is a positive integer less than or equal to N; among them, when i equals 1, the initial recovery feature of the i-th scale is a preset fixed feature; when i is greater than 1, the initial recovery feature of the i-th scale is the reconstruction feature of the (i - 1)-th scale.
[0147] In some embodiments, the feature decoding network includes M feature decoding subnets, where M is a positive integer. For the j-th feature decoding subnet among the M feature decoding subnets, the input data of the j-th feature decoding subnet is processed through the j-th feature decoding subnet to obtain the recovered feature output by the j-th feature decoding subnet, where j is a positive integer less than or equal to M. Among them, when j equals 1, the input data of the j-th feature decoding subnet includes the adapted feature of the i-th scale and the initially recovered feature of the i-th scale. When j is greater than 1, the input data of the j-th feature decoding subnet includes the adapted feature of the i-th scale and the recovered feature output by the (j - 1)-th feature decoding subnet. The recovered feature output by the M-th feature decoding subnet among the M feature decoding subnets is determined as the reconstructed feature of the i-th scale.
[0148] In some embodiments, through the j-th feature decoding subnet, two different linear processes are performed on the adapted feature of the i-th scale. As shown in subfigure (b) of Figure 3, through linear layer 1, a second intermediate feature is obtained. Through linear layer 2, a third intermediate feature is obtained. The values of each recovered feature included in the input data of the j-th feature decoding subnet are subtracted from the mean value of the values of each recovered feature to obtain a fourth intermediate feature. The values at the corresponding positions in the second intermediate feature and the fourth intermediate feature are multiplied to obtain a fifth intermediate feature. The fifth intermediate feature is divided by the standard deviation of the values of each recovered feature to obtain a sixth intermediate feature. The values at the corresponding positions in the sixth intermediate feature and the third intermediate feature are added to obtain the recovered feature output by the j-th feature decoding subnet.
[0149] Step 670: With the goal of minimizing the difference between the reconstructed features of N scales and the encoded features of N scales, the parameters of the image detection model are adjusted to obtain the trained image detection model.
[0150] Taking the difference between the reconstructed features of N scales and the encoded features of N scales as the loss function, calculate the value of the loss function. With the goal of minimizing the value of the loss function, adjust the parameters of the image detection model. When the value of the loss function is lower than the preset threshold, stop training, thereby obtaining the trained image detection model. The loss function can be L1 (absolute error), or MSE (Mean Squared Error), or cosine distance. This application does not make any limitations in this regard.
[0151] In some embodiments, such as Figure 4As shown, the InvAD proposed in this application is compared with GT, RD, and UniAD. In contrast, the InvAD proposed in this paper produces more accurate anomaly localization results. Other methods may misjudge normal regions as abnormal regions. Exemplarily, the RD method mislocates regions 47 and 48 as abnormal regions. The UniAD method mislocates regions 49 and 410 as abnormal regions, etc.
[0152] In some embodiments, please refer to Figure 7 , which shows a schematic diagram of the comparison of experimental results with different resolutions provided in an embodiment of this application. Figure 7 shows the results on the mAD I , mAD P , mAU-PRO and mIoU-max metrics for different sample image resolutions from 64 to 512. As the resolution increases, our method can obtain better experimental results. When the resolution of the sample image is greater than or equal to 256×256, compared with other methods, the method proposed in this application shows obvious advantages. Considering the computational load and performance, the default resolution 256×256 is adopted during training in this paper.
[0153] In some embodiments, please refer to Figure 8 , which shows a schematic diagram of the comparison of experimental results with different training epochs provided in an embodiment of this application. Figure 8 shows the results on the mAD I , mAD P , mAU-PRO and mIoU-max metrics for different training epochs of 100, 200, 300, 600, and 1000. The convergence speed of our method is faster than other methods, and stable results can be obtained with only 100 iterations.
[0154] In some embodiments, this application conducts a comparative test on the MVTec AD dataset. N R and N S are the number of channels of the upsampled features at N scales and the number of channels of the adapted features at N scales respectively. N B , N C and N L are the number of layers of the Bottleneck, the number of ordinary convolutional layers in the upsampling subnet, and the number of convolutional layers of the feature decoding network respectively. mAD I , mAD P , mAU-PRO and MIoU-max are used as evaluation metrics. Channel Configure refers to the configuration of the number of channels, and Stack Number Configure refers to the configuration of the number of layers. As shown in Table 1:
[0155] Table 1
[0156]
[0157]
[0158] According to Table 1, it shows that a relatively small number of channels is sufficient to obtain satisfactory results, while an excessive number of channels does not significantly improve the results but instead incurs additional computational costs. And for N B 、N C and N L with the number of layers configured as 1 layer, 1 layer, and 2 layers respectively, it is sufficient to obtain satisfactory results, and more layers may have an adverse impact on the model performance.
[0159] In some embodiments, for mAD I 、mAD P 、 mAU-PRO and MIoU-max, these 5 metrics, the present application analyzes the results under different loss function constraints. Among them, LOSS refers to the loss function, Cos f represents using the cosine distance between the encoded feature and the decoded feature as the loss function, Cos p represents using the partial cosine distance between the encoded feature and the decoded feature as the loss function. L1 represents using the absolute distance between the encoded feature and the decoded feature as the loss function, and MSE represents using the mean square error between the encoded feature and the decoded feature as the loss function. Sch. represents the scheduler, where Cosine refers to the cosine annealing scheduler, and Step refers to the step decay strategy of reducing the learning rate, and they are two different schedulers respectively. Among them, as shown in Table 2 below:
[0160] Table 2
[0161]
[0162] According to Table 2, the method of the present application has strong robustness under the constraints of Cos f 、Cos p 、L1 and MSE, and there is no significant difference in the results of the metrics of the present application under the two schedulers of Cosine and Step.
[0163] In some embodiments, on three mainstream datasets, namely the COCO-AD dataset, the MVTec AD dataset, and the VisA dataset. Our method has significantly higher advantages compared to DRAEM (Dynamic Routing for Aspect Extraction and Modulation), RD, UniAD, DeSTSeg (Deep Spatio-Temporal Semantic Segmentation), and SimpleNet methods in 14 metrics. Among these metrics, the prefix "m" in the metric represents the average measurement result for all classes.
[0164] In Table 3, Image-level (Classification) refers to the level of classifying the entire image in computer vision tasks. This means that the model classifies the input image as a whole without considering specific regions or pixels within the image. Region-level refers to the level of analyzing and processing specific regions within an image in computer vision tasks. At this level, the model focuses on specific regions in the image. Pixel-level (Segmentation) refers to the level of segmenting and classifying each pixel in an image in computer vision tasks. This means that the model predicts the class or label to which each pixel in the image belongs, thereby segmenting the image into multiple regions. Pixel-level segmentation usually requires more refined, pixel-level annotation or label information. Averaged Metrics refers to the metric that comprehensively considers or averages multiple metrics. When evaluating the performance of a model, multiple metrics can be used to measure different aspects of performance, such as accuracy, recall, precision, etc. Averaged Metrics combines these metrics, usually by calculating the mean or weighted average to obtain a comprehensive evaluation metric.
[0165] Metrics represent different indicators. Among them, the meanings of different indicators are as follows: mAU-ROC (mean Average-precision of area under ROC curve): the average ROC-AUC (area under the ROC curve) indicator; mAP (mean Average Precision): the average precision indicator; mF1-max: the maximum F1 value indicator, a metric that considers both precision and recall; mAU-PRO (mean Average-precision of Parzen-Rosenblatt estimator Output) is a performance indicator used to detect problems, similar to mAU-ROC. It uses a large-area probability density function (Parzen-Rosenblatt estimator) to simulate the sample probability distribution of each class and uses these probability values to calculate the average accuracy; MIoU-max (maximum mean Intersection over Union): the maximum average IoU value indicator; mAD I (mean Average Distance Intersection): the average intersection distance indicator; mAD P (mean AverageDistance Precision): the average distance precision indicator. As shown in Table 3 below:
[0166] Table 3
[0167]
[0168]
[0169] According to Table 3, the bold numbers in the table are the optimal results. From the results, it can be seen that on the COCO-AD dataset in the general complex scenario, the industrial MVTec AD dataset, and the small object defect detection dataset VisA, the proposed InvAD framework can achieve significantly better results in all metrics.
[0170] The technical solution proposed in this application effectively retains the feature information of the input image at multiple scales by extracting the encoded features of the input image at multiple scales. The feature fusion network fuses the encoded features at multiple scales to obtain fused features. The upsampling network upsamples the fused features to obtain upsampled features at multiple scales. Further, the style conversion network transforms the channel dimensions of the upsampled features at different scales to obtain adapted features at multiple scales, which can increase the diversity of feature expressions. According to the adapted features at different scales through the feature decoding network, high-quality reconstruction features at different scales can be obtained. On the one hand, the above network proposed in this application only includes convolutional operations, which is simple and efficient. On the other hand, through the fusion and utilization of the encoded features at multiple scales, high-quality reconstruction features at multiple scales can be obtained, thereby achieving more accurate and efficient abnormal image detection.
[0171] The following is an embodiment of the device of this application, which can be used to execute the method embodiment of this application. For the details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.
[0172] Please refer to Figure 9 , which shows a block diagram of an abnormal image detection device based on an image detection model provided by an embodiment of this application. The device has the function of implementing the above-mentioned abnormal image detection method based on the image detection model, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be a computer device or can be set in a computer device. The device 900 may include: an extraction module 910, a first obtaining module 920, a second obtaining module 930, a third obtaining module 940, a fourth obtaining module 950, and a determination module 960.
[0173] The extraction module 910 is configured to extract the encoded features of the first image at N scales for the first image to be detected, where N is an integer greater than 1.
[0174] The first obtaining module 920 is configured to fuse the encoded features at the N scales through the feature fusion network to obtain fused features.
[0175] The second obtaining module 930 is configured to upsample the fused features through the upsampling network to obtain the upsampled features at the N scales.
[0176] The third obtaining module 940 is configured to perform channel transformation on the upsampled features at the N scales through the style conversion network to obtain the adapted features at the N scales, where the channel transformation is channel compression or channel expansion.
[0177] The fourth obtaining module 950 is configured to obtain the reconstruction features at the N scales according to the adapted features at the N scales through the feature decoding network.
[0178] A determination module 960, configured to determine an anomaly detection result of the first image according to the reconstruction features of the N scales and the encoded features of the N scales.
[0179] In some embodiments, the fourth obtaining module 950 includes: a first obtaining unit ( Figure 9 not shown in the figure).
[0180] The first obtaining unit is configured to, for the i-th scale among the N scales, obtain the reconstruction feature of the i-th scale through the feature decoding network according to the adaptation feature of the i-th scale and the initially recovered feature of the i-th scale, where i is a positive integer less than or equal to N; wherein, when i is equal to 1, the initially recovered feature of the i-th scale is a preset fixed feature, and when i is greater than 1, the initially recovered feature of the i-th scale is the reconstruction feature of the (i - 1)-th scale.
[0181] In some embodiments, the feature decoding network includes M feature decoding subnets, where M is a positive integer; the first obtaining unit includes: an obtaining subunit and a determining subunit ( Figure 9 not shown in the figure).
[0182] The obtaining subunit is configured to, for the j-th feature decoding subnet among the M feature decoding subnets, process the input data of the j-th feature decoding subnet through the j-th feature decoding subnet to obtain the recovered feature output by the j-th feature decoding subnet, where j is a positive integer less than or equal to M, and wherein, when j is equal to 1, the input data of the j-th feature decoding subnet includes the adaptation feature of the i-th scale and the initially recovered feature of the i-th scale, and when j is greater than 1, the input data of the j-th feature decoding subnet includes the adaptation feature of the i-th scale and the recovered feature output by the (j - 1)-th feature decoding subnet.
[0183] The determining subunit is configured to determine the recovered feature output by the M-th feature decoding subnet among the M feature decoding subnets as the reconstruction feature of the i-th scale.
[0184] In some embodiments, the obtaining subunit is configured to: perform two different linear processes on the adaptation feature of the i-th scale through the j-th feature decoding subnet to obtain a second intermediate feature and a third intermediate feature; subtract the mean value of each value of the restored feature from each value of the restored feature included in the input data of the j-th feature decoding subnet to obtain a fourth intermediate feature; multiply the corresponding values of the second intermediate feature and the fourth intermediate feature to obtain a fifth intermediate feature; divide the fifth intermediate feature by the standard deviation of each value of the restored feature to obtain a sixth intermediate feature; and add the corresponding values of the sixth intermediate feature and the third intermediate feature to obtain the restored feature output by the j-th feature decoding subnet.
[0185] In some embodiments, the determining module 960 includes: a second obtaining unit, a third obtaining unit, a fourth obtaining unit, and a determining unit ( Figure 9 not shown in the figure).
[0186] The second obtaining unit is configured to, for the i-th scale among the N scales, obtain a difference image of the i-th scale according to the difference between the reconstruction feature of the i-th scale and the corresponding position values in the encoding feature of the i-th scale, where the value of each pixel in the difference image of the i-th scale is the difference between the reconstruction feature of the i-th scale and a set of corresponding position values in the encoding feature of the i-th scale, and i is a positive integer less than or equal to N.
[0187] The third obtaining unit is configured to convert the difference images of the N scales into the scale of the first image to obtain N difference images of the same scale.
[0188] The fourth obtaining unit is configured to perform a fusion process on the N difference images of the same scale to obtain a final difference image.
[0189] The determining unit is configured to determine an anomaly detection result of the first image based on the final difference image.
[0190] In some embodiments, the fourth obtaining unit is configured to: add the corresponding position values in the N difference images of the same scale to obtain the final difference image; or average the corresponding position values in the N difference images of the same scale to obtain the final difference image.
[0191] In some embodiments, the determining unit is configured to: if there are abnormal pixels in the final difference image, determine that the first image is an abnormal image, where the abnormal pixels refer to pixels whose values belong to a set value range; and determine an abnormal area in the first image based on the abnormal pixels.
[0192] In some embodiments, the first obtaining module 920 includes: a fusion unit ( Figure 9 not shown in
[0193] The fusion unit is configured to convert the encoded features of the N scales into encoded features of the same scale through the feature fusion network, obtaining N encoded features of the same scale; and perform channel splicing on the N encoded features of the same scale to obtain the fusion feature.
[0194] In some embodiments, the first obtaining module 920 further includes: a transformation unit ( Figure 9 not shown in
[0195] The transformation unit is configured to perform feature transformation on the fusion feature through at least one Bottleneck layer to obtain a transformed fusion feature, where the transformed fusion feature has the same scale and number of channels as the fusion feature, and the transformed fusion feature has a different numerical representation from the fusion feature, and the transformed fusion feature is used for upsampling to obtain the upsampled features of the N scales.
[0196] In some embodiments, the upsampling network includes N upsampling subnets; the second obtaining module 930 includes: an upsampling unit and a determination unit ( Figure 9 not shown in
[0197] The upsampling unit is configured to, for the i-th upsampling subnet among the N upsampling subnets, perform upsampling on the input data of the i-th upsampling subnet through the i-th upsampling subnet to obtain the upsampled feature output by the i-th upsampling subnet, where i is a positive integer less than or equal to N. When i equals 1, the input data of the i-th upsampling subnet is the fusion feature; when i is greater than 1, the input data of the i-th upsampling subnet is the upsampled feature output by the (i - 1)-th upsampling subnet.
[0198] The determination unit is configured to determine the upsampled features output by the N upsampling subnets as the upsampled features of the N scales.
[0199] In some embodiments, the i-th upsampling subnet includes an upsampling convolutional layer and at least one ordinary convolutional layer; the upsampling unit is configured to: perform upsampling on the input data of the i-th upsampling subnet through the upsampling convolutional layer to obtain a first intermediate feature; and perform convolutional processing on the first intermediate feature through the at least one ordinary convolutional layer to obtain the upsampled feature output by the i-th upsampling subnet.
[0200] In some embodiments, the third obtaining module 940 is configured to: for the i-th scale among the N scales, perform convolution processing on the upsampled features of the i-th scale in the channel dimension through the style conversion network to obtain the adapted features of the i-th scale, where the convolution processing in the channel dimension is used to implement channel compression or channel expansion, and i is a positive integer less than or equal to N.
[0201] The technical solution proposed in this application effectively retains the feature information of multiple scales of the input image by extracting the encoded features of the input image at multiple scales. The feature fusion network fuses the encoded features of multiple scales to obtain fused features. The upsampling network upsamples the fused features to obtain upsampled features of multiple scales. Further, the style conversion network transforms the channel dimension of the upsampled features of different scales to obtain adapted features of multiple scales, which can increase the diversity of feature expression. According to the adapted features of different scales through the feature decoding network, high-quality reconstruction features of different scales can be obtained. On the one hand, only convolution operations are included in the above network proposed in this application, which is simple and efficient. On the other hand, through the fusion and utilization of encoded features of multiple scales, high-quality reconstruction features of multiple scales can be obtained, thereby realizing more accurate and efficient abnormal image detection.
[0202] Please refer to Figure 10 , which shows a block diagram of a training device for an image detection model provided by an embodiment of this application. This device has the function of implementing the training method of the above image detection model, and this function can be implemented by hardware or by hardware executing corresponding software. This device can be a computer device or can be set in a computer device. The device 1000 may include: an acquisition module 1010, an extraction module 1020, a first obtaining module 1030, a second obtaining module 1040, a third obtaining module 1050, a fourth obtaining module 1060, and an adjustment module 1070.
[0203] The acquisition module 1010 is configured to acquire a sample image for training the image detection model, and there is no abnormal area in the sample image.
[0204] The extraction module 1020 is configured to extract encoded features of N scales of the sample image, where N is an integer greater than 1.
[0205] The first obtaining module 1030 is configured to fuse the encoded features of the N scales through the feature fusion network to obtain fused features.
[0206] The second obtaining module 1040 is configured to upsample the fused features through the upsampling network to obtain the upsampled features of the N scales.
[0207] A third obtaining module 1050 is configured to perform channel transformation on the upsampled features of the N scales through the style transformation network to obtain the adapted features of the N scales, where the channel transformation is channel compression or channel expansion.
[0208] A fourth obtaining module 1060 is configured to obtain the reconstructed features of the N scales through the feature decoding network according to the adapted features of the N scales.
[0209] An adjustment module 1070 is configured to adjust the parameters of the image detection model with the goal of minimizing the difference between the reconstructed features of the N scales and the encoded features of the N scales, so as to obtain a trained image detection model.
[0210] The technical solution proposed in this application effectively retains the feature information of multiple scales of the input image by extracting the encoded features of the input image at multiple scales. The feature fusion network fuses the encoded features of multiple scales to obtain a fused feature. The upsampling network upsamples the fused feature to obtain upsampled features of multiple scales. Further, the style transformation network transforms the channel dimension of the upsampled features of different scales to obtain adapted features of multiple scales, which can increase the diversity of feature expression. According to the adapted features of different scales through the feature decoding network, high-quality reconstructed features of different scales can be obtained. On the one hand, the above-mentioned network proposed in this application only includes convolutional operations, which is simple and efficient. On the other hand, by fusing and utilizing the encoded features of multiple scales, high-quality reconstructed features of multiple scales can be obtained, thereby realizing more accurate and efficient abnormal image detection.
[0211] It should be noted that for the device provided in the above embodiment, when implementing its functions, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0212] Please refer to Figure 11 , which shows a structural block diagram of a computer device 1100 provided in an embodiment of the present application.
[0213] Generally, the computer device 1100 includes a processor 1110 and a memory 1120.
[0214] The processor 1110 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1110 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1110 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1110 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1110 may further include an AI processor, which is used to process computational operations related to machine learning.
[0215] The memory 1120 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1120 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1120 is used to store a computer program, and the computer program is configured to be executed by one or more processors to implement the above-mentioned abnormal image detection method based on the image detection model or the training method of the above-mentioned image detection model.
[0216] Those skilled in the art can understand that Figure 11 the structure shown in does not constitute a limitation on the computer device 1100, and it may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.
[0217] In some embodiments, a computer-readable storage medium is also provided. A computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the above-mentioned abnormal image detection method based on the image detection model or the training method of the above-mentioned image detection model.
[0218] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0219] In some embodiments, a computer program product is also provided. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor reads and executes the computer program to implement the above-mentioned abnormal image detection method based on the image detection model or the above-mentioned training method of the image detection model.
[0220] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described in this article only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of the present application do not limit this.
[0221] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An abnormal image detection method based on an image detection model, characterized in that, The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network; the method includes: For a first image to be detected, extract encoded features of N scales of the first image, where N is an integer greater than 1; Fuse the encoded features of the N scales through the feature fusion network to obtain a fused feature; Upsample the fused feature through the upsampling network to obtain upsampled features of the N scales; Perform channel transformation on the upsampled features of the N scales through the style conversion network to obtain adapted features of the N scales, where the channel transformation is channel compression or channel expansion; the channel compression means reducing the channel dimension, and the channel expansion means increasing the channel dimension; For the i-th scale among the N scales, through the feature decoding network, based on the adapted feature of the i-th scale and the initial recovery feature of the i-th scale, obtain the reconstructed feature of the i-th scale, where i is a positive integer less than or equal to N; when i equals 1, the initial recovery feature of the i-th scale is a preset fixed feature, and when i is greater than 1, the initial recovery feature of the i-th scale is the reconstructed feature of the (i - 1)-th scale; the reconstructed feature refers to a feature representation that is generated after a decoding process based on the original encoded feature and has the same scale as the original encoded feature; Determine the anomaly detection result of the first image based on the reconstructed features of the N scales and the encoded features of the N scales.
2. The method according to claim 1, wherein The feature decoding network includes M feature decoding subnets, where M is a positive integer; The step of obtaining the reconstructed feature of the i-th scale through the feature decoding network based on the adapted feature of the i-th scale and the initial recovery feature of the i-th scale includes: For the j-th feature decoding subnet among the M feature decoding subnets, process the input data of the j-th feature decoding subnet through the j-th feature decoding subnet to obtain the recovered feature output by the j-th feature decoding subnet, where j is a positive integer less than or equal to M, and when j equals 1, the input data of the j-th feature decoding subnet includes the adapted feature of the i-th scale and the initial recovery feature of the i-th scale, and when j is greater than 1, the input data of the j-th feature decoding subnet includes the adapted feature of the i-th scale and the recovered feature output by the (j - 1)-th feature decoding subnet; Determine the recovered feature output by the M-th feature decoding subnet among the M feature decoding subnets as the reconstructed feature of the i-th scale.
3. The method according to claim 2, characterized in that, The step of processing the input data of the j-th feature decoding subnet through the j-th feature decoding subnet to obtain the recovered feature output by the j-th feature decoding subnet includes: Perform two different linear processes on the adapted feature of the i-th scale through the j-th feature decoding subnet to obtain a second intermediate feature and a third intermediate feature; Subtract each value of the restored feature included in the input data of the j-th feature decoding subnet from the mean value of each value of the restored feature to obtain a fourth intermediate feature; Multiply the corresponding values in the second intermediate feature and the fourth intermediate feature to obtain a fifth intermediate feature; Divide the fifth intermediate feature by the standard deviation of each value of the restored feature to obtain a sixth intermediate feature; Add the corresponding values in the sixth intermediate feature and the third intermediate feature to obtain the restored feature output by the j-th feature decoding subnet.
4. The method according to claim 1, characterized in that, The determining the anomaly detection result of the first image according to the reconstructed features of the N scales and the encoded features of the N scales includes: For the i-th scale among the N scales, obtain a difference image of the i-th scale according to the difference between the reconstructed feature of the i-th scale and the corresponding position values in the encoded feature of the i-th scale, where the value of each pixel in the difference image of the i-th scale is the difference between a group of corresponding position values in the reconstructed feature of the i-th scale and the encoded feature of the i-th scale, and i is a positive integer less than or equal to N; Convert the difference images of the N scales to the scale of the first image to obtain N difference images of the same scale; Perform a fusion process on the N difference images of the same scale to obtain a final difference image; Based on the final difference image, determine the anomaly detection result of the first image.
5. The method according to claim 4, wherein The performing a fusion process on the N difference images of the same scale to obtain a final difference image includes: Add the corresponding position values in the N difference images of the same scale to obtain the final difference image; Or, Average the corresponding position values in the N difference images of the same scale to obtain the final difference image.
6. The method according to claim 5, wherein The determining the anomaly detection result of the first image based on the final difference image includes: If there are abnormal pixels in the final difference image, determine that the first image is an abnormal image, where the abnormal pixels are pixels whose values belong to a set value range; Based on the abnormal pixels, determine the abnormal region in the first image.
7. The method according to claim 1, characterized in that, The fusing the encoded features of the N scales through the feature fusion network to obtain a fused feature includes: Convert the encoded features of the N scales into encoded features of the same scale through the feature fusion network to obtain N encoded features of the same scale; Perform channel concatenation on the N encoded features of the same scale to obtain the fused feature.
8. The method according to claim 7, wherein After performing channel concatenation on the N encoded features of the same scale to obtain the fused feature, it further includes: Perform feature transformation on the fused feature through at least one Bottleneck layer to obtain a transformed fused feature, where the transformed fused feature has the same scale and number of channels as the fused feature, and the transformed fused feature has a different numerical representation from the fused feature, and the transformed fused feature is used for upsampling to obtain the upsampled features of the N scales.
9. The method according to claim 1, wherein The upsampling network includes N upsampling subnets; Upsampling the fusion feature through the upsampling network to obtain the upsampling features at the N scales, including: For the i-th upsampling subnet among the N upsampling subnets, upsampling the input data of the i-th upsampling subnet through the i-th upsampling subnet to obtain the upsampling feature output by the i-th upsampling subnet, where i is a positive integer less than or equal to N. Among them, when i equals 1, the input data of the i-th upsampling subnet is the fusion feature; when i is greater than 1, the input data of the i-th upsampling subnet is the upsampling feature output by the (i - 1)-th upsampling subnet; Determine the upsampling features output by the N upsampling subnets as the upsampling features at the N scales.
10. The method according to claim 9, characterized in that The i-th upsampling subnet includes an upsampling convolutional layer and at least one ordinary convolutional layer; Upsampling the input data of the i-th upsampling subnet through the i-th upsampling subnet to obtain the upsampling feature output by the i-th upsampling subnet, including: Upsampling the input data of the i-th upsampling subnet through the upsampling convolutional layer to obtain a first intermediate feature; Performing convolutional processing on the first intermediate feature through the at least one ordinary convolutional layer to obtain the upsampling feature output by the i-th upsampling subnet.
11. The method according to claim 1, characterized in that, Performing channel transformation on the upsampling features at the N scales through the style conversion network to obtain the adapted features at the N scales, including: For the i-th scale among the N scales, performing convolutional processing on the upsampling feature at the i-th scale through the style conversion network in the channel dimension to obtain the adapted feature at the i-th scale, where the convolutional processing in the channel dimension is used to achieve channel compression or channel expansion, and i is a positive integer less than or equal to N.
12. A training method for an image detection model, characterized in that, The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network; the method includes: Obtain a sample image for training the image detection model, where there is no abnormal area in the sample image; Extract the encoded features at N scales of the sample image, where N is an integer greater than 1; Fuse the encoded features at the N scales through the feature fusion network to obtain a fusion feature; Upsample the fusion feature through the upsampling network to obtain the upsampling features at the N scales; Perform channel transformation on the upsampling features at the N scales through the style conversion network to obtain the adapted features at the N scales, where the channel transformation is channel compression or channel expansion; channel compression means reducing the channel dimension, and channel expansion means increasing the channel dimension; For the i-th scale among the N scales, through the feature decoding network, based on the adaptive feature of the i-th scale and the initial restored feature of the i-th scale, a reconstructed feature of the i-th scale is obtained, where i is a positive integer less than or equal to N; when i equals 1, the initial restored feature of the i-th scale is a preset fixed feature, and when i is greater than 1, the initial restored feature of the i-th scale is the reconstructed feature of the (i - 1)-th scale; the reconstructed feature refers to a feature representation that is generated after the decoding process based on the original encoded feature and has the same scale as the original encoded feature. With the goal of minimizing the difference between the reconstructed features of the N scales and the encoded features of the N scales, the parameters of the image detection model are adjusted to obtain a trained image detection model.
13. An abnormal image detection device based on an image detection model, characterized in that, The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network; the device includes: An extraction module, configured to extract encoded features of N scales of a first image to be detected, where N is an integer greater than 1. A first obtaining module, configured to fuse the encoded features of the N scales through the feature fusion network to obtain a fused feature. A second obtaining module, configured to upsample the fused feature through the upsampling network to obtain upsampled features of the N scales. A third obtaining module, configured to perform channel transformation on the upsampled features of the N scales through the style conversion network to obtain adaptive features of the N scales, where the channel transformation is channel compression or channel expansion; channel compression refers to reducing the channel dimension, and channel expansion refers to increasing the channel dimension. The fourth obtaining module includes a first obtaining unit, configured to, for the i-th scale among the N scales, through the feature decoding network, based on the adaptive feature of the i-th scale and the initial restored feature of the i-th scale, obtain a reconstructed feature of the i-th scale, where i is a positive integer less than or equal to N; when i equals 1, the initial restored feature of the i-th scale is a preset fixed feature, and when i is greater than 1, the initial restored feature of the i-th scale is the reconstructed feature of the (i - 1)-th scale; the reconstructed feature refers to a feature representation that is generated after the decoding process based on the original encoded feature and has the same scale as the original encoded feature. A determination module, configured to determine an anomaly detection result of the first image according to the reconstructed features of the N scales and the encoded features of the N scales.
14. A training device for an image detection model, characterized in that, The image detection model includes a feature fusion network, an upsampling network, a style conversion network, and a feature decoding network; the device includes: An acquisition module, configured to acquire a sample image for training the image detection model, where there is no anomaly region in the sample image. An extraction module, configured to extract encoded features of N scales of the sample image, where N is an integer greater than 1. A first obtaining module, configured to fuse the encoded features of the N scales through the feature fusion network to obtain a fused feature. A second obtaining module, configured to perform upsampling on the fused feature through the upsampling network to obtain upsampled features at the N scales; A third obtaining module, configured to perform channel transformation on the upsampled features at the N scales through the style conversion network to obtain adapted features at the N scales, where the channel transformation is channel compression or channel expansion; the channel compression means reducing the channel dimension, and the channel expansion means increasing the channel dimension; The fourth obtaining module includes a first obtaining unit, configured to, for the i-th scale among the N scales, obtain a reconstructed feature at the i-th scale through the feature decoding network according to the adapted feature at the i-th scale and the initially restored feature at the i-th scale, where i is a positive integer less than or equal to N; when i is equal to 1, the initially restored feature at the i-th scale is a preset fixed feature, and when i is greater than 1, the initially restored feature at the i-th scale is the reconstructed feature at the (i - 1)-th scale; the reconstructed feature refers to a feature representation that is generated after the original encoded feature undergoes a decoding process and has the same scale as the original encoded feature; An adjustment module, configured to adjust the parameters of the image detection model with the goal of minimizing the difference between the reconstructed features at the N scales and the encoded features at the N scales to obtain a trained image detection model.
15. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 11, or the method according to claim 12.
16. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 11, or the method according to claim 12.
17. A computer program product, characterized in that The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and the processor reads and executes the computer program from the computer-readable storage medium to implement the method according to any one of claims 1 to 11, or the method according to claim 12.
Citation Information
Patent Citations
Image detection method and device, medium and equipment
CN116958025A
Method for constructing image anomaly detection model based on mask multi-modal generative adversarial network
CN116994044A