Image Anomaly Detection Method, Device, Electronic Device and Storage Medium
By performing image correction and feature extraction on the original image, combined with two-stage training of neural networks, the problem of low abnormal positioning accuracy in the prior art is solved, and high-precision abnormality detection is achieved at the pixel level.
Patent Information
- Application Number
- CN202111147844.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-09-28
AI Technical Summary
The existing image abnormality detection methods do not perform well in abnormal positioning accuracy, especially in industrial quality inspection and medical image detection, and it is difficult for the prior art to achieve high-precision abnormality detection.
By performing image correction and feature extraction on the original image, the subnet model of the neural network is used for two-stage training, including image-level coarse alignment and feature-level fine alignment, combined with Gaussian distribution and feature vector calculation, pixel-level abnormal positioning is achieved.
It improves the accuracy of abnormal detection, realizes pixel-level abnormal positioning, and can more obviously capture fine-grained abnormalities, improving the accuracy of detection results.
Smart Images

Figure CN113888498B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to an image anomaly detection method, apparatus, electronic device, and storage medium. Background Art
[0002] As a popular research topic in recent years, computer vision technology has been widely applied in various industries in society, and is particularly prominent in the industrial and medical fields. Anomaly or defect detection is an inevitable problem in the industrial and medical fields, such as defective product detection in industrial quality inspection, auxiliary medical image detection, etc. The core of anomaly detection lies in distinguishing normal images from abnormal images through computer vision technology. The current mainstream approach is to learn the generation method of normal samples based on a reconstruction scheme to identify abnormal samples in detection, or to treat anomaly detection as a single-classification problem for anomaly classification and localization. However, these schemes do not perform well in the accuracy of anomaly localization, and the accuracy of anomaly detection is relatively low. Summary of the Invention
[0003] Embodiments of this application provide an image anomaly detection method, apparatus, electronic device, and storage medium.
[0004] In a first aspect of the embodiments of this application, an image anomaly detection method is provided. The method includes:
[0005] Perform image correction on the original image to obtain a first image to be detected;
[0006] Perform feature extraction and correction on the first image to be detected to obtain a first target feature map;
[0007] Obtain an anomaly detection result based on the first target feature map and the feature map of the positive sample image.
[0008] In combination with the first aspect, in a possible implementation manner, obtaining an anomaly detection result based on the first target feature map and the feature map of the positive sample image includes:
[0009] Obtain the target Gaussian distribution corresponding to each position in the feature map of the positive sample image;
[0010] Obtain the anomaly score of each position in the first target feature map according to the feature of each position in the first target feature map and the target Gaussian distribution;
[0011] Determine the anomaly score map composed of the anomaly scores of each position in the first target feature map as the anomaly detection result.
[0012] In combination with the first aspect, in a possible implementation manner, obtaining the anomaly score of each position in the first target feature map according to the feature of each position in the first target feature map and the target Gaussian distribution includes:
[0013] Calculate the first distance between the feature at each position in the first target feature map and the target Gaussian distribution corresponding to that position;
[0014] Determine the first distance as the anomaly score at each position in the first target feature map.
[0015] Combined with the first aspect, in a possible implementation manner, image correction is performed on the original image to obtain a first image to be detected, including:
[0016] Obtain a first transformation matrix;
[0017] Align the original image using the first transformation matrix to obtain a first image to be detected.
[0018] Combined with the first aspect, in a possible implementation manner, feature extraction and correction are performed on the first image to be detected to obtain a first target feature map, including:
[0019] Perform feature extraction on the first image to be detected to obtain a feature map to be corrected;
[0020] Obtain a second transformation matrix;
[0021] Align the feature map to be corrected using the second transformation matrix to obtain a first feature map;
[0022] Obtain the first target feature map based on the first feature map.
[0023] Combined with the first aspect, in a possible implementation manner, feature extraction and correction are performed on the first image to be detected to obtain a first target feature map, including:
[0024] Perform the operation of the first feature extraction and correction on the first image to be detected to obtain the first first feature map;
[0025] Perform the operation of the (r + 1)-th feature extraction and correction on the r-th first feature map to obtain the (r + 1)-th first feature map, where r is an integer greater than or equal to 1;
[0026] Perform at least 2 operations of feature extraction and correction to obtain M first feature maps, where M is an integer greater than or equal to 2, and the M first feature maps include the first first feature map and the (r + 1)-th first feature map;
[0027] Obtain the first target feature map according to the M first feature maps.
[0028] Combined with the first aspect, in a possible implementation manner, the first image to be detected and the first target feature map are obtained through a first sub-network model, and the first sub-network model is obtained by training the first sub-network of a neural network;
[0029] The first sub-network model is trained by the following steps:
[0030] The first sub-network corrects N positive sample images in the sample image set to obtain N second images to be detected, where the N positive sample images correspond one-to-one to the N second images to be detected, and N is an integer greater than 1;
[0031] The first sub-network extracts and corrects features from the N second images to be detected to obtain N second feature maps, and the N second images to be detected correspond one-to-one to the N second feature Figure 1 maps;
[0032] Determine the target loss according to the first loss determined by the N second images to be detected and the second loss determined by the N second feature maps;
[0033] Adjust the parameters of the neural network, iterate the positive sample image set to make the target loss converge, and obtain the first sub-network model.
[0034] Combined with the first aspect, in a possible implementation, the steps for determining the first loss include:
[0035] Randomly select two second images to be detected from the N second images to be detected;
[0036] Calculate the second distance between the two second images to be detected, and determine the second distance as the first loss.
[0037] Combined with the first aspect, in a possible implementation, the neural network further includes a second sub-network and a third sub-network, and the steps for determining the second loss include:
[0038] Encode any second feature map a in the N second feature maps through the second sub-network to obtain a first set of feature vectors;
[0039] Encode any second feature map b in the shuffled N second feature maps through the third sub-network to obtain a second set of feature vectors;
[0040] Determine the second loss according to the first set of feature vectors and the second set of feature vectors.
[0041] Combined with the first aspect, in a possible implementation, encoding any second feature map a in the N second feature maps through the second sub-network to obtain a first set of feature vectors includes:
[0042] For each first position in any second feature map a, encode and project the feature at the first position through the second sub-network to obtain a first feature vector;
[0043] A first feature vector set is composed of first feature vectors corresponding to each first position;
[0044] Encoding any second feature map b among the N shuffled second feature maps through a third sub-network to obtain a second feature vector set, including:
[0045] For each second position in any second feature map b, encoding the feature of the second position through a third sub-network to obtain a second feature vector;
[0046] A second feature vector set is composed of second feature vectors corresponding to each second position.
[0047] Combined with the first aspect, in a possible implementation manner, determining a second loss according to the first feature vector set and the second feature vector set, including:
[0048] Determining a target second feature map with the same sequence identifier as any second feature map a from the N shuffled second feature maps, and any second feature map b includes the target second feature map;
[0049] Obtaining a target second feature vector set corresponding to the target second feature map, and the second feature vector set includes the target second feature vector set;
[0050] Determining a target second feature vector corresponding to the first feature vector in the first feature vector set from the target second feature vector set;
[0051] Determining the second loss according to the first feature vector and the corresponding target second feature vector.
[0052] Combined with the first aspect, in a possible implementation manner, performing feature extraction and correction on N second images to be detected through a first sub-network to obtain N second feature maps, including:
[0053] For any one of the N second images to be detected, performing at least 2 operations of feature extraction and correction through the first sub-network to obtain the second feature map of any one of the N second images to be detected;
[0054] The N second feature maps are composed of the second feature maps of each of the N second images to be detected.
[0055] Combined with the first aspect, in a possible implementation manner, after obtaining the first sub-network model, the method further includes:
[0056] Processing N positive sample images through the first sub-network model to obtain N second target feature maps, and the N positive sample images and the N second target features Figure 1 are in one-to-one correspondence;
[0057] Perform Gaussian fitting on the features corresponding to the positions in the N second target feature maps to obtain the target Gaussian distribution of the positive sample image.
[0058] Combined with the first aspect, in a possible implementation, the N positive sample images are processed by the first sub-network model to obtain N second target feature maps, including:
[0059] Perform image correction on the N positive sample images through the first sub-network model to obtain N third images to be detected, and the N positive sample images and the N third images to be detected correspond one by one;
[0060] Perform feature extraction and correction on the N third images to be detected through the first sub-network model to obtain N second target feature maps, and the N second target feature maps and the N third images to be detected correspond one by one;
[0061] Among them, the feature extraction and correction of any one of the N third images to be detected includes:
[0062] Perform the operation of the first feature extraction and correction on any one of the third images to be detected through the first sub-network model to obtain the first third feature map;
[0063] Perform the operation of the (r + 1)th feature extraction and correction on the rth third feature map through the first sub-network model to obtain the (r + 1)th third feature map;
[0064] Perform at least 2 operations of feature extraction and correction through the first sub-network model to obtain R third feature maps, where R is an integer greater than or equal to 2, and the R third feature maps include the first third feature map and the (r + 1)th third feature map;
[0065] Obtain the second target feature map corresponding to any one of the third images to be detected through the first sub-network model according to the R third feature maps.
[0066] A second aspect of the embodiments of the present application provides an image anomaly detection device, and the device includes an acquisition unit and a processing unit;
[0067] The acquisition unit is used to perform image correction on the original image to obtain the first image to be detected;
[0068] The processing unit is used to perform feature extraction and correction on the first image to be detected to obtain the first target feature map;
[0069] The processing unit is further used to obtain an anomaly detection result according to the first target feature map and the feature map of the positive sample image.
[0070] In a third aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes an input device and an output device, and further includes a processor and a computer storage medium. The processor is adapted to implement one or more instructions; and the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the steps in the image anomaly detection method described in the first aspect above.
[0071] In a fourth aspect of the embodiments of the present application, a computer storage medium is provided. The computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by a processor to perform the steps in the image anomaly detection method described in the first aspect above.
[0072] In a fifth aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes a computer program, and the computer program is operable to cause a computer to perform the steps in the image anomaly detection method described in the first aspect above. The computer program product may be a software installation package.
[0073] It can be seen that in the embodiments of the present application, the original image is corrected to obtain a first image to be detected; feature extraction and correction are performed on the first image to be detected to obtain a first target feature map; and an anomaly detection result is obtained based on the first target feature map and the feature map of the positive sample image. In this way, the original image is corrected so as to align the features at each position of the original image with the feature distribution of the normal image, which is beneficial to capturing fine-grained anomalies from each position subsequently, realizing pixel-level anomaly localization. In addition, the first target feature map is obtained by fusing feature maps of different layers, which can aggregate the anomalies at each position in the feature map, making the pixel-level anomalies more obvious. Pixel-level anomaly localization and more obvious pixel-level anomalies mean higher accuracy of anomaly localization. Based on the first target feature map with higher localization accuracy and the feature map of the normal sample image, it is beneficial to improve the accuracy of the anomaly detection result. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0075] Figure 1 It is an architecture diagram of an application environment provided by an embodiment of the present application;
[0076] Figure 2 It is a schematic structural diagram of a neural network provided by an embodiment of the present application;
[0077] Figure 3 Schematic flow chart of an image anomaly detection method provided by an embodiment of the present application;
[0078] Figure 4 Schematic diagram of obtaining a first target feature map provided by an embodiment of the present application;
[0079] Figure 5 Schematic comparison diagram of an uncoarsely aligned image and a coarsely aligned image provided by an embodiment of the present application;
[0080] Figure 6 Schematic flow chart of another image anomaly detection method provided by an embodiment of the present application;
[0081] Figure 7 Schematic flow chart of the processing in the fine alignment stage provided by an embodiment of the present application;
[0082] Figure 8 Schematic diagram of selecting a second feature map provided by an embodiment of the present application;
[0083] Figure 9 Schematic structural diagram of an image anomaly detection device provided by an embodiment of the present application;
[0084] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0085] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0086] The terms "including" and "having" and any variations thereof that appear in the specification, claims and drawings of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. In addition, the terms "first", "second", "third", etc. are used to distinguish different objects, rather than to describe a specific order.
[0087] Reference to "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application may be combined with other embodiments.
[0088] An embodiment of this application proposes an image anomaly detection method, which can be implemented based on Figure 1 the application environment shown, such as Figure 1 shown, the application environment includes an image acquisition device 101, a terminal device 102, and an electronic device 103. Among them, the image acquisition device 101, the terminal device 102, and the electronic device 103 are connected through a network. The electronic device 103 involved in the embodiments of this application may include various devices with the ability to run program code and communicate. For example, the electronic device 103 may be an independent physical server, an embedded device, or a server cluster or a distributed system. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms, and so on.
[0089] Among them, the image acquisition device 101 is used to send the original images collected in various fields such as industry and medical to the electronic device 103, so that the electronic device 103 performs anomaly detection on the original images, and the anomaly detection results can be displayed through the terminal device 102. For example, the anomaly detection may be defect detection of high-speed rail catenaries, anomaly detection of workpieces, lesion segmentation in medical images, etc. Among them, the terminal device 102 is also used to provide a positive sample image set to the electronic device 103, so that the electronic device 103 trains the neural network through the positive sample image set and deploys the trained neural network model locally or on other devices. Then, the electronic device 103 or the other device can perform image anomaly detection through the neural network model. That is, the training device and the calling device of the neural network model can be the same device or different devices. Since the electronic device 103 or the other device can achieve pixel-level anomaly localization during anomaly detection, it is beneficial to improve the accuracy of image anomaly detection when the anomaly localization accuracy is higher.
[0090] The neural network architecture of the embodiments of this application is briefly described below in conjunction with the relevant drawings.
[0091] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a neural network provided by an embodiment of this application, such as Figure 2As shown, the neural network includes a first sub-network, a second sub-network, and a third sub-network. The entire framework can be divided into a coarse alignment stage and a fine alignment stage. Among them, the first sub-network is used for the processing in the coarse alignment stage, and the second sub-network and the third sub-network are used for the processing in the fine alignment stage. The purpose of providing the entire framework is to obtain a feature extractor that can be used for pixel-level anomaly localization through training the neural network, that is, the first sub-network model obtained by training the first sub-network.
[0092] Among them, the first sub-network mainly includes an Image-level Coarse Alignment (ICA) module and a feature extraction network. The Image-level Coarse Alignment module is used to locate the positive sample images in the input positive sample image set to similar directions and positions, and obtain the to-be-detected images after coarse alignment of each positive sample image to regularize the pixel distribution of the positive sample images. The feature extraction network uses a pre-trained ResNet-18 network as the basic structure. It should be understood that the ResNet-18 network usually includes 4 feature extraction layers (layer1, layer2, layer3, and layer4). In this application, a Feature-Level Coarse Alignment (FCA) module is added after each feature extraction layer in the feature extraction network. The Feature-Level Coarse Alignment module further aligns the features extracted by each feature extraction layer to enable the features of the positive sample images to complete global depth contraction, facilitating the comparison of pixel-level dense features in the subsequent fine alignment stage. Exemplarily, in this application, corresponding layers can be selected from the ResNet-18 network as feature extraction layers. For example, the pre-trained layer1, layer2, and layer3 are used as the feature extraction layers in this application, and Feature-Level Coarse Alignment modules are added after layer1, layer2, and layer3 respectively.
[0093] Among them, the second sub-network and the third sub-network are two branches after the first sub-network. The direct input of the second sub-network is the feature map output by the last Feature-Level Coarse Alignment module in the first sub-network. The second sub-network maps the directly input feature map to a feature vector through an encoder f and a predictor g. The input of the third sub-network is the feature map with the order of the feature map output by the last Feature-Level Coarse Alignment module in the first sub-network scrambled. The third sub-network maps the scrambled feature map to a feature vector through an encoder f. In the fine alignment stage, finally, based on the feature vectors extracted by the second sub-network and the third sub-network, the pixel-level (or fine-grained) comparison between positive sample images is completed, enabling the Feature-Level Coarse Alignment module to learn the feature distribution of normal images in a self-supervised manner. Exemplarily, in order to avoid the situation of network collapse or breakdown caused by training the neural network only with positive sample images, this application also adds a stop gradient operation in the third sub-network, that is, only the second sub-network is allowed to participate in the gradient backpropagation.
[0094] It can be seen that for the neural network architecture provided in this application, through two-stage training from coarse alignment to fine alignment, it is beneficial to supervise the first sub-network to learn the pixel-level feature distribution of positive sample images, so that pixel-level anomaly localization can be achieved during feature extraction, which in turn is beneficial to improving the accuracy of anomaly detection.
[0095] The following elaborates in detail on the image anomaly detection method provided in the embodiments of this application in conjunction with relevant accompanying drawings.
[0096] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an image anomaly detection method provided in the embodiments of this application. This method is applied to an electronic device. As Figure 3 shown, it includes steps 301-304:
[0097] Step 301: Perform image correction on the original image to obtain a first image to be detected.
[0098] In the embodiments of this application, for the original image collected in the actual scene, as Figure 4 shown, it is processed using the first sub-network model obtained through training. The original image is corrected through the image-level coarse alignment module to obtain a first image to be detected. Specifically, the image-level coarse alignment module completes image-level coarse alignment by obtaining the transformation matrix (i.e., the first transformation matrix) on the original image and regressing the transformation matrix in a manner similar to the spatial transformation network. Its formula is as follows:
[0099]
[0100] Among them, represents the pixel in the original image, represents the affine transformation matrix. represents the pixel after coarse alignment, and Z i represents the original image, and τ θ represents the image-level coarse alignment module. Among them, the first transformation matrix includes an affine transformation matrix. As Figure 5 shown, for the abnormal screw image, before coarse alignment, its abnormal area is relatively loose on the heat map, which will lead to low accuracy of subsequent anomaly localization. After coarse alignment, due to the contraction of the feature distribution, its abnormal area is relatively compact on the heat map, which will relatively improve the accuracy of subsequent anomaly localization.
[0101] Step 302: Perform feature extraction and correction on the first image to be detected to obtain a first target feature map.
[0102] In the embodiments of this application, please continue to refer toFigure 4 For the first image to be detected obtained in step 301, feature extraction is performed using a feature extractor. Exemplarily, feature extraction and correction are performed on the first image to be detected to obtain a first target feature map, including:
[0103] Feature extraction is performed on the first image to be detected to obtain a feature map to be corrected;
[0104] A transformation matrix is obtained;
[0105] The feature map to be corrected is aligned using the transformation matrix to obtain a first feature map;
[0106] The first target feature map is obtained based on the first feature map.
[0107] Specifically, for the first image to be detected, the layer of the ResNet-18 network is used to perform feature extraction on it, and the obtained feature map is the feature map to be corrected. Similar to image correction, FCA obtains the transformation matrix (i.e., the second transformation matrix) on the feature map to be corrected, and then performs regression on the second transformation matrix to complete feature-level rough alignment, obtaining the feature map output by FCA (i.e., the first feature map), and the final first target feature map is obtained based on the first feature map. For example: when only one layer + FCA is used, the first feature map is the first target feature map; when multiple layers + FCA are used, the fused feature map of the feature maps output by each FCA is the first target feature map.
[0108] Exemplarily, feature extraction and correction are performed on the first image to be detected to obtain a first target feature map, including:
[0109] The operation of performing the first feature extraction and correction on the first image to be detected is performed to obtain the first first feature map;
[0110] The operation of performing the (r + 1)-th feature extraction and correction on the r-th first feature map is performed to obtain the (r + 1)-th first feature map, where r is an integer greater than or equal to 1;
[0111] The operation of performing at least two feature extraction and correction operations is performed to obtain M first feature maps, where M is an integer greater than or equal to 2, and the M first feature maps include the first first feature map and the (r + 1)-th first feature map;
[0112] The first target feature map is obtained according to the M first feature maps.
[0113] Taking three feature extraction layers (assumed to be layer1, layer2, and layer3 in sequence) as an example, there will be a feature-level coarse alignment module (assumed to be FCA1, FCA2, and FCA3 in sequence) after each feature extraction layer, and a total of 3 operations of feature extraction and correction are performed. Specifically, layer1 is used to extract features from the first image to be detected. For the feature map output by layer1, FCA1 is used to perform coarse alignment on it to obtain the first feature map of the first one; layer2 is used to extract features from the first feature map of the first one. For the feature map output by layer2, FCA2 is used to perform coarse alignment on it to obtain the second first feature map; layer3 is used to extract features from the second first feature map. For the feature map output by layer3, FCA3 is used to perform coarse alignment on it to obtain the third first feature map. The first first feature map, the second first feature map, and the third first feature map are the above-mentioned M first feature maps. In this embodiment, during the process of feature extraction from the first image to be detected, a feature-level coarse alignment process is incorporated to align the features at each position of the extracted first feature map with the feature distribution of the normal image, realizing further contraction of the feature distribution, which is beneficial to achieving fine-grained anomaly localization.
[0114] Exemplarily, the processing of the feature maps output by the corresponding layer by FCA1, FCA2, and FCA3 includes: obtaining the transformation matrix on the feature map and regressing the transformation matrix, that is, completing the coarse alignment of the feature map.
[0115] For the obtained M first feature maps, the first feature maps other than the first first feature map in the M first feature maps (such as the (r + 1)-th first feature map) are scaled to obtain a to-be-fused feature map with the same size as the first first feature map, and the to-be-fused feature map is fused with the first first feature map to obtain a first target feature map. For example, the size of the above-mentioned third first feature map is scaled to be the same as the size of the first first feature map to obtain the corresponding to-be-fused feature map, the size of the above-mentioned second first feature map is scaled to be the same as the size of the first first feature map to obtain the corresponding to-be-fused feature map, and all to-be-fused feature maps are fused with the first first feature map to obtain the first target feature map for anomaly score calculation.
[0116] In this embodiment, since the first target feature map is obtained by fusing the first feature maps of different layers, it can aggregate the anomalies at each position in the first feature map, making the pixel-level anomalies more obvious, and further enabling the position of the anomaly to be clearly observed according to the calculated anomaly score.
[0117] It should be noted that the processing method of the feature-level rough alignment module is the same as that of the image-level rough alignment module, both of which perform regression on the affine transformation matrix. For the number of feature extraction layers and feature-level rough alignment modules, it can be adaptively adjusted according to the actual scenario. For example, 4 layers of the ResNet-18 network can also be used for feature extraction. The above examples do not limit the embodiments of the present application.
[0118] In this embodiment, a feature-level rough alignment module is added after each feature extraction layer to perform rough alignment of features, which is beneficial to align the extracted features with the affine transformation matrix learned by each feature-level rough alignment module. Then, the features on normal pixels will tend to follow a normal distribution. Relatively speaking, the difference between the feature distribution on abnormal pixels and the normal distribution will become more obvious, which is more conducive to realizing pixel-level anomaly localization.
[0119] Step 303: Obtain an anomaly detection result according to the first target feature map and the feature map of the positive sample image.
[0120] In the embodiments of the present application, after the first sub-network model is trained, the first sub-network model is used to infer the positive sample image set to obtain N second target feature maps. The electronic device performs Gaussian fitting on the features corresponding to the positions in the N second target feature maps, and establishes a Gaussian distribution model N(μ ij , ∑ ij ) for the features corresponding to the positions in the N second target feature maps. The formula is as follows:
[0121]
[0122] Among them, μ ij represents the mean of the features at the position of the i-th row and j-th column in the k-th second target feature map among the N second target feature maps, k represents the k-th second target feature map among the N second target feature maps, represents the feature at the position of the i-th row and j-th column in the k-th second target feature map, Σ ij represents the covariance of the features at the position of the i-th row and j-th column in the N second target feature maps, and T represents the transpose. Among them, position correspondence means the position of the i-th row and j-th column on each of the N second target feature maps, and N is an integer greater than 1.
[0123] It should be understood that the N second target feature maps have the same size as the first target feature map, so there is a corresponding Gaussian distribution model (i.e., the target Gaussian distribution) for each position of the first target feature map. Exemplarily, obtaining an anomaly detection result according to the first target feature map and the feature map of the positive sample image includes:
[0124] Obtain the target Gaussian distribution corresponding to each position in the feature map of the positive sample image;
[0125] According to the features at each position in the first target feature map and the target Gaussian distribution, obtain the anomaly scores at each position in the first target feature map;
[0126] Determine the anomaly score map composed of the anomaly scores at each position in the first target feature map as the anomaly detection result.
[0127] In the embodiments of the present application, the positive sample image refers to the image in the above positive sample image set. By calculating the first distance between the features at each position in the first target feature map and the target Gaussian distribution corresponding to this position, the anomaly scores at each position in the first target feature map are obtained. Exemplarily, the Mahalanobis distance can be used as the measure of the anomaly score, that is, the first distance includes the Mahalanobis distance. By calculating the Mahalanobis distance between the features at each position in the first target feature map and the target Gaussian distribution corresponding to this position, the anomaly scores at each position in the first target feature map are obtained, and its formula is expressed as follows:
[0128]
[0129] where d(e ij ) represents the anomaly score at the position of the i-th row and j-th column in the first target feature map, and e ij represents the feature at the position of the i-th row and j-th column in the first target feature map.
[0130] Finally, the anomaly scores at each position in the first target feature map are composed into a pixel-level anomaly score map. This anomaly score map is a distance matrix, and the anomaly score map is used as the anomaly detection result for output. It should be understood that the area where the anomaly score is greater than the threshold is the detected anomaly area, and the higher the anomaly score, the higher the degree of anomaly.
[0131] It can be seen that in the embodiments of the present application, the first image to be detected is obtained by performing image correction on the original image; the first target feature map is obtained by performing feature extraction and correction on the first image to be detected; and the anomaly detection result is obtained according to the first target feature map and the feature map of the positive sample image. In this way, image correction is performed on the original image to align the features at each position of the original image with the feature distribution of the normal image, which is beneficial to capturing fine-grained anomalies from each position later and realizing pixel-level anomaly localization. In addition, the first target feature map is obtained by fusing feature maps of different layers, which can aggregate the anomalies at each position in the feature map, making the pixel-level anomalies more obvious. Pixel-level anomaly localization and more obvious pixel-level anomalies mean higher accuracy of anomaly localization. Based on the first target feature map with higher localization accuracy and the feature map of the normal sample image, it is beneficial to improve the accuracy of the anomaly detection result.
[0132] Please refer to Figure 6 , Figure 6 , which is a schematic flowchart of another image anomaly detection method provided by an embodiment of the present application. As Figure 6 shown, it includes steps 601-605:
[0133] Step 601: Perform image correction on the original image to obtain the first image to be detected;
[0134] Step 602: Perform feature extraction and correction on the first image to be detected to obtain the first target feature map;
[0135] Step 603: Obtain the target Gaussian distribution corresponding to each position in the feature map of the positive sample image;
[0136] Step 604: According to the features of each position in the first target feature map and the target Gaussian distribution, obtain the anomaly score of each position in the first target feature map;
[0137] Step 605: Determine the anomaly score map composed of the anomaly scores of each position in the first target feature map as the anomaly detection result.
[0138] Among them, the specific implementation manners of steps 601-605 are already described in the embodiments Figure 3 shown, and will not be elaborated here.
[0139] Exemplarily, before performing image correction on the original image to obtain the first image to be detected, the method further includes:
[0140] Perform image correction on N positive sample images in the sample image set through the first sub-network to obtain N second images to be detected, where the N positive sample images and the N second images to be detected correspond one by one, and N is an integer greater than 1;
[0141] Perform feature extraction and correction on the N second images to be detected through the first sub-network to obtain N second feature maps, and the N second images to be detected and the N second features Figure 1 correspond one by one;
[0142] Determine the target loss according to the first loss determined by the N second images to be detected and the second loss determined by the N second feature maps;
[0143] Adjust the parameters of the neural network, iterate the positive sample image set, so that the target loss converges, and obtain the first sub-network model.
[0144] In the embodiments of the present application, please continue to refer to Figure 2, for the N positive sample images in the batch of positive sample images, input them into the first sub-network. First, the image-level coarse alignment module performs coarse alignment on each positive sample image to obtain the second image to be detected after coarse alignment for each positive sample image, and correspondingly, N second images to be detected will be generated. Among them, for the implementation of coarse alignment, reference can be made to the implementation manner in step 301 above, that is, through the image-level coarse alignment module τ θ regresses the transformation matrix on each positive sample image. In this implementation manner, no standard direction or standard position is assigned to the N positive sample images to achieve coarse alignment. The image-level coarse alignment module learns the pixel-level feature distribution of normal images in a self-supervised manner during the training stage, which is beneficial for capturing subtle anomalies in abnormal images, thereby improving the accuracy and performance of anomaly localization.
[0145] Exemplarily, the first sub-network performs feature extraction and correction on the N second images to be detected to obtain N second feature maps, including:
[0146] For any one of the N second images to be detected, at least 2 operations of feature extraction and correction are performed through the first sub-network to obtain the second feature map of any one of the second images to be detected;
[0147] The second feature maps of each of the N second images to be detected form N second feature maps.
[0148] Among them, the at least 2 operations of feature extraction and correction include:
[0149] Perform the first feature extraction and correction on any one of the second images to be detected to obtain the fourth feature map;
[0150] Perform the second feature extraction and correction on the fourth feature map, and repeat the operation of performing the current feature extraction and correction on the feature map obtained from the previous feature extraction and correction. After at least 2 operations of feature extraction and correction, the second feature map of any one of the second images to be detected is obtained.
[0151] In the embodiments of the present application, the feature extraction network in the first sub-network is used to process N second images to be detected. For any one of the N second images to be detected, in the feature extraction network, through at least two operations of feature extraction and correction, the corresponding second feature map is obtained. Specifically, the at least two operations of feature extraction and correction correspond to the number of feature extraction layers in the feature extraction network. For example, if layer1, layer2, and layer3 of the ResNet-18 network are used as feature extraction layers, there are correspondingly three operations of feature extraction and rough alignment. Using layer1 to perform feature extraction on the any one of the second images to be detected, for the feature map output by layer1, using FCA1 to perform rough alignment on it to obtain the feature map output by the first feature extraction and rough alignment; using layer2 to perform feature extraction on the feature map output by the first feature extraction and rough alignment, for the feature map output by layer2, using FCA2 to perform rough alignment on it to obtain the feature map output by the second feature extraction and rough alignment; using layer3 to perform feature extraction on the feature map output by the second feature extraction and rough alignment, for the feature map output by layer3, using FCA3 to perform rough alignment on it to obtain the feature map output by the third feature extraction and rough alignment, and determining the feature map output by the third feature extraction and rough alignment as the second feature map. Thus, the above-mentioned N second feature maps are obtained from N second images to be detected. In this embodiment, the at least two operations of feature extraction and correction indicate that there are at least two feature extraction layers. Correspondingly, after each of the at least two feature extraction layers, there is a feature-level rough alignment module connected. In this way, the at least two feature extraction layers and the at least two feature-level rough alignment modules can be trained to improve the capabilities of the feature extraction layers and the feature-level rough alignment modules, which is convenient for fusing the feature maps output by each feature-level rough alignment module in the application stage to achieve the purpose of abnormal signal enhancement. Based on the image-level rough alignment, the feature-level rough alignment module also further learns the feature distribution of normal images in a self-supervised manner, and enhances the learning process of the feature distribution of normal images through layer-by-layer rough alignment processing, so as to improve the ability of the feature extraction network to capture abnormal features.
[0152] Exemplarily, the step of determining the first loss includes:
[0153] Randomly select two second images to be detected from the N second images to be detected;
[0154] Calculate the second distance between the two second images to be detected, and determine the second distance as the first loss.
[0155] In the embodiments of the present application, in order to train the image-level rough alignment module, two second images to be detected corresponding to the positive sample images P and Q are randomly selected from N second images to be detected, and the distance between the second images to be detected corresponding to the positive sample images P and Q (i.e., the second distance) is calculated. Exemplarily, the L2 norm of the second images to be detected corresponding to the positive sample images P and Q can be used as the second distance, and the second distance is determined as the first loss, and its formula is expressed as follows:
[0156]
[0157] Wherein, L ICA represents the first loss, represents the second image to be detected corresponding to the positive sample image P, represents the second image to be detected corresponding to the positive sample image Q, D represents the set of positive sample images, and θ h represents the parameters in the affine transformation matrix, H represents the height of the positive sample images P and Q, and W represents the width of the positive sample images P and Q. In this embodiment, the first loss is used to supervise the alignment of the positive sample images P and Q in the direction of reducing entropy. After a certain number of iterations, the entropy will be reduced to a certain value, and the image-level rough alignment module can learn how to align the input images.
[0158] Exemplarily, the steps for determining the second loss include:
[0159] Encoding any second feature map a in the N second feature maps through a second sub-network to obtain a first set of feature vectors;
[0160] Encoding any second feature map b in the shuffled N second feature maps through a third sub-network to obtain a second set of feature vectors;
[0161] Determining the second loss according to the first set of feature vectors and the second set of feature vectors.
[0162] Wherein, encoding any second feature map a in the N second feature maps through a second sub-network to obtain a first set of feature vectors includes:
[0163] For each first position in any second feature map a, encoding and projecting the feature at the first position through the second sub-network to obtain a first feature vector;
[0164] The first set of feature vectors is composed of the first feature vectors corresponding to each first position;
[0165] Wherein, encoding any second feature map b in the shuffled N second feature maps through a third sub-network to obtain a second set of feature vectors includes:
[0166] For each second position in any second feature map b, encode the feature at the second position through a third sub-network to obtain a second feature vector;
[0167] Form a second feature vector set from the second feature vectors corresponding to each second position.
[0168] In the embodiments of the present application, please refer to Figure 7 , form a second feature map sequence A by arranging the N second feature maps in the original order and directly input it into the second sub-network for processing. For any second feature map a in the second feature map sequence A, encode any second feature map a through an encoder f to obtain a first feature vector set. Shuffle the N second feature maps, form a second feature map sequence B from the shuffled N second feature maps, input the second feature map sequence B into the third sub-network for processing, and for any second feature map b in the second feature map sequence B, encode any second feature map b through an encoder f to obtain a second feature vector set.
[0169] Among them, for each first position in any second feature map a, the encoder f in the second sub-network encodes the feature w ij at this first position into a to-be-projected feature vector m ij , form a to-be-projected feature vector set from the to-be-projected feature vectors m ij corresponding to each first position, and project the to-be-projected feature vectors m ij in the to-be-projected feature vector set into the vector space of the second feature vectors n ij in the second feature vector set through the predictor g in the second sub-network. The projected to-be-projected feature vector is the first feature vector, denoted as g(m ij ). In other words, each second feature map in the second feature map sequence A corresponds to a first feature vector set. Among them, the encoder f uses 3 1*1 convolutional layers to encode w ij into m ij . Among them, the predictor g uses 2 1*1 convolutional layers to map m ij to g(m ij ). For each second position in any second feature map b, the encoder f in the third sub-network encodes the feature v ij at this second position into a second feature vector n ij , form a second feature vector set from the second feature vectors n ij corresponding to each second position. In other words, each second feature map in the second feature map sequence B corresponds to a second feature vector set. Among them, the encoder f uses 3 1*1 convolutional layers to encode v ij into n ijAmong them, the encoder f in the second sub-network and the encoder f in the third sub-network are shared encoders. In this embodiment, the encoder f is used to encode the features at each position in any second feature map a into feature vectors, and the encoder f is used to encode the features at each position in any second feature map b into feature vectors, so as to achieve dense supervision for each position of any second feature map a and any second feature map b in the fine alignment stage. By shuffling the order, the similarity between N second feature maps is maximized.
[0170] Exemplarily, determining the second loss according to the first feature vector set and the second feature vector set includes:
[0171] Determine the target second feature map with the same sequence identifier as any second feature map a from the shuffled N second feature maps, and any second feature map b includes the target second feature map;
[0172] Obtain the target second feature vector set corresponding to the target second feature map, and the second feature vector set includes the target second feature vector set;
[0173] Determine the target second feature vector corresponding to the first feature vector in the first feature vector set from the target second feature vector set;
[0174] Determine the second loss according to the first feature vector and the corresponding target second feature vector.
[0175] Among them, determining the target loss by determining the second loss according to the first feature vector and the corresponding target second feature vector includes:
[0176] Calculate the negative cosine similarity between the first feature vector and the target second feature vector;
[0177] Determine the second loss according to the negative cosine similarity and the stop gradient operation in the third sub-network.
[0178] In the embodiments of the present application, please refer to Figure 8 , assuming that the number of second feature maps in the second feature map sequence A and the second feature map sequence B is 3, and any second feature map a is the second second feature map in the second feature map sequence A, then its sequence identifier is 2. Correspondingly, the target second feature map is also the second second feature map in the second feature map sequence B. It should be understood that after the encoding by the encoder f, there is a corresponding target second feature vector set for the target second feature map. For each feature vector m to be projected ij , there is a corresponding target second feature vector n in the target second feature vector set ij , and the negative cosine similarity between the first feature vector g(m ij ) and the target second feature vector n ij is used to represent the feature vector m to be projectedij The similarity with the target second feature vector n ij is as follows:
[0179]
[0180] where L ij (m ij , n ij , θ g , θ f ) represents the negative cosine similarity between the first feature vector g(m ij ) and the target second feature vector n ij , θ g represents the parameter of the predictor g, and θ f represents the parameter of the encoder f. In the fine alignment stage, by minimizing the negative cosine similarity between any second feature map a and the corresponding target second feature map at the position of the i-th row and j-th column, the feature alignment between any second feature map a and the target second feature map is densely supervised, which is beneficial to improving the performance of the feature-level coarse alignment module in feature-level coarse alignment.
[0181] It can be seen from the structure of the neural network that in order to prevent the network from crashing, a stop gradient operation (stop_grad) is added to the third sub-network in this solution. That is, after each iteration, the third sub-network does not perform the gradient backpropagation operation, and only the second sub-network is allowed to perform the gradient backpropagation. That is, the above adjustment of the parameters of the neural network is based on the gradient backpropagated by the second sub-network, which can enhance the robustness of the model. Based on the purpose of minimizing the negative cosine similarity and the stop gradient operation, the total loss at each position of the N second feature maps can be determined, that is, the second loss, and its formula is as follows:
[0182]
[0183] where L FAS represents the second loss, which means that the neural network does not receive gradient information from n ij , m represents the first feature vector set, and n represents the second feature vector set.
[0184] L FAS is the main loss function for supervising all parameters of the neural network, and L ICA is the auxiliary loss function for supervising the convergence of the image-level feature extraction module. Using the coefficients λ1 and λ2 to adjust the weights of the coarse alignment stage and the fine alignment stage, the target loss of the entire neural network framework is defined as:
[0185] L total (·; θ h , f, g, τ θ ) = λ1·L ICA+λ2·L FAS ;
[0186] Based on the target loss L total The rough alignment stage and the fine alignment stage are optimized as a whole, and the positive sample image set is continuously iterated until the target loss converges, that is, the first sub-network model is obtained.
[0187] It can be seen that before applying the first sub-network model in the embodiments of the present application, a two-stage training method is also used to train the neural network framework. In the rough alignment stage, the image-level rough alignment module and the feature-level rough alignment module are mainly supervised to perform image-level alignment and feature-level alignment on the feature distribution of normal images. In the fine alignment stage, the entire neural network is mainly supervised to learn the feature distribution of normal images through pixel-level comparison, so as to improve the abnormal capture ability of the first sub-network, and thus the positioning effect of the first sub-network model on subtle abnormalities can be improved.
[0188] Exemplarily, after obtaining the first sub-network model, the method further includes:
[0189] Processing N positive sample images through the first sub-network model to obtain N second target feature maps. The N positive sample images and the N second target features Figure 1 are in one-to-one correspondence. The sizes of the N second target feature maps are the same, and the sizes of the N second target feature maps are the same as those of the first target feature map;
[0190] Performing Gaussian fitting on the features corresponding to the positions in the N second target feature maps to obtain the target Gaussian distribution of the positive sample images.
[0191] In the embodiments of the present application, processing N positive sample images through the first sub-network model to obtain N second target feature maps includes:
[0192] Performing image correction on N positive sample images through the first sub-network model to obtain N third images to be detected. The N positive sample images and the N third images to be detected are in one-to-one correspondence;
[0193] Performing feature extraction and correction on N third images to be detected through the first sub-network model to obtain N second target feature maps. The N second target feature maps and the N third images to be detected are in one-to-one correspondence;
[0194] Among them, the feature extraction and correction of any one of the N third images to be detected includes:
[0195] Performing the operation of the first feature extraction and correction on any one of the third images to be detected through the first sub-network model to obtain the first third feature map;
[0196] The operation of performing the (r + 1)-th feature extraction and correction on the r-th third feature map is carried out through the first sub-network model to obtain the (r + 1)-th third feature map;
[0197] At least two operations of feature extraction and correction are carried out through the first sub-network model to obtain R third feature maps, where R is an integer greater than or equal to 2. The R third feature maps include the first third feature map and the (r + 1)-th third feature map;
[0198] The second target feature map corresponding to any third image to be detected is obtained through the first sub-network model according to the R third feature maps.
[0199] Specifically, after the first sub-network model is trained, the first sub-network model is used to re-infer N positive sample images, and they are roughly aligned through the image set rough alignment module to obtain N third images to be detected. The feature extraction network is used to perform feature extraction and correction on the N third images to be detected. The feature map output by each feature-level rough alignment module is the third feature map. For any third image to be detected, there are corresponding multiple third feature maps. Among them, the processing process of the feature extraction network can refer to the description of performing feature extraction and correction on the first image to be detected in step 302. For the R third feature maps of any third image to be detected, scale normalization and fusion are carried out in the manner of step 302 to obtain the second target feature map. In this way, N second target feature maps can be obtained. The features at the corresponding positions in the N second target feature maps are subjected to Gaussian fitting to construct a pixel-level target Gaussian distribution, so as to facilitate calculating the anomaly score at each position in the first target feature map based on the first target feature map and the target Gaussian distribution in the application stage, and finally obtaining the anomaly detection result.
[0200] It can be seen that in the embodiment of the present application, the first image to be detected is obtained by performing image correction on the original image; the first target feature map is obtained by performing feature extraction and correction on the first image to be detected; the anomaly detection result is obtained according to the first target feature map and the feature map of the positive sample image. In this way, image correction is performed on the original image to align the features at each position of the original image with the feature distribution of the normal image, which is beneficial to capturing fine-grained anomalies from each position subsequently, realizing pixel-level anomaly localization. In addition, the first target feature map is obtained by fusing feature maps of different layers, which can aggregate the anomalies at each position in the feature map, making the pixel-level anomalies more obvious. Pixel-level anomaly localization and more obvious pixel-level anomalies mean higher accuracy of anomaly localization. Based on the first target feature map with higher localization accuracy and the feature map of the normal sample image, it is beneficial to improve the accuracy of the anomaly detection result.
[0201] Based on the description of the above method embodiments, the embodiments of the present application further provide an image anomaly detection device. Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an image anomaly detection device provided by an embodiment of the present application. As Figure 9 shown, the device includes an acquisition unit 901 and a processing unit 902;
[0202] The acquisition unit 901 is configured to perform image correction on the original image to obtain a first image to be detected;
[0203] The processing unit 902 is configured to perform feature extraction and correction on the first image to be detected to obtain a first target feature map;
[0204] The processing unit 902 is further configured to obtain an anomaly detection result according to the first target feature map and the feature map of the positive sample image.
[0205] It can be seen that in Figure 9 the image anomaly detection device shown, by performing image correction on the original image, a first image to be detected is obtained; by performing feature extraction and correction on the first image to be detected, a first target feature map is obtained; and an anomaly detection result is obtained according to the first target feature map and the feature map of the positive sample image. In this way, image correction is performed on the original image to align the features at each position of the original image with the feature distribution of the normal image, which is beneficial to capturing fine-grained anomalies from each position later and realizing pixel-level anomaly localization. In addition, the first target feature map is obtained by fusing feature maps of different layers, which can aggregate the anomalies at each position in the feature map, making the pixel-level anomalies more obvious. Pixel-level anomaly localization and more obvious pixel-level anomalies mean higher accuracy of anomaly localization. Based on the first target feature map with higher localization accuracy and the feature map of the normal sample image, it is beneficial to improve the accuracy of the anomaly detection result.
[0206] In a possible implementation manner, in terms of obtaining an anomaly detection result according to the first target feature map and the feature map of the positive sample image, the processing unit 902 is specifically configured to:
[0207] Obtain the target Gaussian distribution corresponding to each position in the feature map of the positive sample image;
[0208] According to the features at each position in the first target feature map and the target Gaussian distribution, obtain the anomaly score at each position in the first target feature map;
[0209] Determine the anomaly score map composed of the anomaly scores at each position in the first target feature map as the anomaly detection result.
[0210] In a possible implementation manner, in terms of obtaining the anomaly score of each position in the first target feature map based on the features of each position in the first target feature map and the target Gaussian distribution, the processing unit 902 is specifically configured to:
[0211] Calculate the first distance between the feature of each position in the first target feature map and the target Gaussian distribution corresponding to this position;
[0212] Determine the first distance as the anomaly score of each position in the first target feature map.
[0213] In a possible implementation manner, in terms of performing image correction on the original image to obtain the first image to be detected, the acquisition unit 901 is specifically configured to:
[0214] Acquire the first transformation matrix;
[0215] Align the original image using the first transformation matrix to obtain the first image to be detected.
[0216] In a possible implementation manner, in terms of performing feature extraction and correction on the first image to be detected to obtain the first target feature map, the processing unit 902 is specifically configured to:
[0217] Perform feature extraction on the first image to be detected to obtain the feature map to be corrected;
[0218] Acquire the second transformation matrix;
[0219] Align the feature map to be corrected using the second transformation matrix to obtain the first feature map;
[0220] Obtain the first target feature map based on the first feature map.
[0221] In a possible implementation manner, in terms of performing feature extraction and correction on the first image to be detected to obtain the first target feature map, the processing unit 902 is specifically configured to:
[0222] Perform the operation of the first feature extraction and correction on the first image to be detected to obtain the first first feature map;
[0223] Perform the operation of the (r + 1)th feature extraction and correction on the rth first feature map to obtain the (r + 1)th first feature map, where r is an integer greater than or equal to 1;
[0224] Perform the operation of feature extraction and correction at least 2 times to obtain M first feature maps, where M is an integer greater than or equal to 2, and the M first feature maps include the first first feature map and the (r + 1)th first feature map;
[0225] Obtain the first target feature map according to the M first feature maps.
[0226] In a possible implementation, the first image to be detected and the first target feature map are obtained through a first sub-network model, and the first sub-network model is obtained by training the first sub-network of the neural network; the processing unit 902 is further configured to:
[0227] Perform image correction on N positive sample images in the sample image set through the first sub-network to obtain N second images to be detected, where the N positive sample images correspond one-to-one to the N second images to be detected, and N is an integer greater than 1;
[0228] Perform feature extraction and correction on the N second images to be detected through the first sub-network to obtain N second feature maps, and the N second images to be detected correspond one-to-one to the N second feature Figure 1 maps;
[0229] Determine the target loss according to the first loss determined from the N second images to be detected and the second loss determined from the N second feature maps;
[0230] Adjust the parameters of the neural network, iterate the positive sample image set to make the target loss converge, and obtain the first sub-network model.
[0231] In a possible implementation, in terms of determining the first loss, the processing unit 902 is specifically configured to:
[0232] Randomly select two second images to be detected from the N second images to be detected;
[0233] Calculate the second distance between the two second images to be detected and determine the second distance as the first loss.
[0234] In a possible implementation, the neural network further includes a second sub-network and a third sub-network. In terms of determining the second loss, the processing unit 902 is specifically configured to:
[0235] Encode any second feature map a in the N second feature maps through the second sub-network to obtain a first set of feature vectors;
[0236] Encode any second feature map b in the shuffled N second feature maps through the third sub-network to obtain a second set of feature vectors;
[0237] Determine the second loss according to the first set of feature vectors and the second set of feature vectors.
[0238] In a possible implementation, in terms of encoding any second feature map a in the N second feature maps through the second sub-network to obtain a first set of feature vectors, the processing unit 902 is specifically configured to:
[0239] For each first position in any second feature map a, encode and project the features at the first position through a second sub-network to obtain a first feature vector;
[0240] Form a first feature vector set from the first feature vectors corresponding to each first position;
[0241] In terms of encoding any second feature map b among the N shuffled second feature maps through a third sub-network to obtain a second feature vector set, the processing unit 902 is specifically configured to:
[0242] For each second position in any second feature map b, encode the features at the second position through a third sub-network to obtain a second feature vector;
[0243] Form a second feature vector set from the second feature vectors corresponding to each second position.
[0244] In a possible implementation manner, in terms of determining a second loss according to the first feature vector set and the second feature vector set, the processing unit 902 is specifically configured to:
[0245] Determine a target second feature map with the same sequence identifier as any second feature map a from the N shuffled second feature maps, and any second feature map b includes the target second feature map;
[0246] Obtain a target second feature vector set corresponding to the target second feature map, and the second feature vector set includes the target second feature vector set;
[0247] Determine target second feature vectors corresponding to the first feature vectors in the first feature vector set from the target second feature vector set;
[0248] Determine the second loss according to the first feature vector and the corresponding target second feature vector.
[0249] In a possible implementation manner, in terms of extracting and correcting features of N second images to be detected through a first sub-network to obtain N second feature maps, the processing unit 902 is specifically configured to:
[0250] For any one of the N second images to be detected, perform at least 2 operations of feature extraction and correction through the first sub-network to obtain a second feature map of any one of the N second images to be detected;
[0251] Form N second feature maps from the second feature maps of each of the N second images to be detected.
[0252] In a possible implementation manner, the processing unit 902 is further configured to:
[0253] Process N positive sample images through the first sub-network model to obtain N second target feature maps, and the N positive sample images and the N second target features Figure 1 are in one-to-one correspondence;
[0254] Perform Gaussian fitting on the features corresponding to the positions in the N second target feature maps to obtain the target Gaussian distribution of the positive sample images.
[0255] In a possible implementation manner, in terms of processing N positive sample images through the first sub-network model to obtain N second target feature maps, the processing unit 902 is specifically configured to:
[0256] Perform image correction on the N positive sample images through the first sub-network model to obtain N third images to be detected, and the N positive sample images and the N third images to be detected are in one-to-one correspondence;
[0257] Perform feature extraction and correction on the N third images to be detected through the first sub-network model to obtain N second target feature maps, and the N second target feature maps and the N third images to be detected are in one-to-one correspondence;
[0258] Among them, in terms of feature extraction and correction of any one of the N third images to be detected, the processing unit 902 is specifically configured to:
[0259] Perform the operation of the first feature extraction and correction on any one of the third images to be detected through the first sub-network model to obtain the first third feature map;
[0260] Perform the operation of the (r + 1)th feature extraction and correction on the rth third feature map through the first sub-network model to obtain the (r + 1)th third feature map;
[0261] Perform at least 2 operations of feature extraction and correction through the first sub-network model to obtain R third feature maps, where R is an integer greater than or equal to 2, and the R third feature maps include the first third feature map and the (r + 1)th third feature map;
[0262] Obtain the second target feature map corresponding to any one of the third images to be detected through the first sub-network model according to the R third feature maps.
[0263] According to an embodiment of the present application, Figure 9Each unit in the image anomaly detection device shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, based on the image anomaly detection device, other units may also be included. In practical applications, these functions can also be assisted and realized by other units, and can be realized through the cooperation of multiple units.
[0264] According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding method shown in Figure 3 or Figure 6 on a general computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct an image anomaly detection device as shown in Figure 9 and to implement the image anomaly detection method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, and be loaded into the above computing device through the computer-readable recording medium and run therein.
[0265] Based on the descriptions of the above method embodiments and device embodiments, the embodiments of this application also provide an electronic device. Please refer to Figure 10 , this electronic device at least includes a processor 1001, an input device 1002, an output device 1003, and a computer storage medium 1004. Among them, the processor 1001, the input device 1002, the output device 1003, and the computer storage medium 1004 within the electronic device can be connected through a bus or other means.
[0266] The computer storage medium 1004 can be stored in the memory of the electronic device. The computer storage medium 1004 is used to store a computer program. The computer program includes program instructions. The processor 1001 is used to execute the program instructions stored in the computer storage medium 1004. The processor 1001 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the electronic device, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to thereby implement the corresponding method process or corresponding function.
[0267] In one embodiment, the processor 1001 of the electronic device provided by the embodiments of this application can be used to perform a series of image anomaly detections:
[0268] Perform image correction on the original image to obtain the first image to be detected;
[0269] Perform feature extraction and correction on the first image to be detected to obtain the first target feature map;
[0270] Obtain the anomaly detection result based on the first target feature map and the feature map of the positive sample image.
[0271] It can be seen that in the Figure 10 shown electronic device, perform image correction on the original image to obtain the first image to be detected; perform feature extraction and correction on the first image to be detected to obtain the first target feature map; obtain the anomaly detection result based on the first target feature map and the feature map of the positive sample image. In this way, perform image correction on the original image to align the features at each position of the original image with the feature distribution of the normal image, which is beneficial to subsequent capturing of fine-grained anomalies from each position, realizing pixel-level anomaly localization. In addition, the first target feature map is obtained by fusing feature maps of different layers, which can aggregate the anomalies at each position in the feature map, making the pixel-level anomalies more obvious. Pixel-level anomaly localization and more obvious pixel-level anomalies mean higher accuracy of anomaly localization. Based on the first target feature map with higher localization accuracy and the feature map of the normal sample image, it is beneficial to improve the accuracy of the anomaly detection result.
[0272] In another embodiment, when the processor 1001 executes to obtain the anomaly detection result based on the first target feature map and the feature map of the positive sample image, it includes:
[0273] Obtain the target Gaussian distribution corresponding to each position in the feature map of the positive sample image;
[0274] Obtain the anomaly score of each position in the first target feature map according to the feature of each position in the first target feature map and the target Gaussian distribution;
[0275] Determine the anomaly score map composed of the anomaly scores of each position in the first target feature map as the anomaly detection result.
[0276] In another embodiment, when the processor 1001 executes to obtain the anomaly score of each position in the first target feature map according to the feature of each position in the first target feature map and the target Gaussian distribution, it includes:
[0277] Calculate the first distance between the feature of each position in the first target feature map and the target Gaussian distribution corresponding to this position;
[0278] Determine the first distance as the anomaly score of each position in the first target feature map.
[0279] In another embodiment, the processor 1001 performs image correction on the original image to obtain a first image to be detected, including:
[0280] Obtain a first transformation matrix;
[0281] Align the original image using the first transformation matrix to obtain a first image to be detected.
[0282] In another embodiment, the processor 1001 performs feature extraction and correction on the first image to be detected to obtain a first target feature map, including:
[0283] Perform feature extraction on the first image to be detected to obtain a feature map to be corrected;
[0284] Obtain a second transformation matrix;
[0285] Align the feature map to be corrected using the second transformation matrix to obtain a first feature map;
[0286] Obtain a first target feature map based on the first feature map.
[0287] In another embodiment, the processor 1001 performs feature extraction and correction on the first image to be detected to obtain a first target feature map, including:
[0288] Perform the operation of the first feature extraction and correction on the first image to be detected to obtain the first first feature map;
[0289] Perform the operation of the (r + 1)th feature extraction and correction on the rth first feature map to obtain the (r + 1)th first feature map, where r is an integer greater than or equal to 1;
[0290] Perform at least two operations of feature extraction and correction to obtain M first feature maps, where M is an integer greater than or equal to 2, and the M first feature maps include the first first feature map and the (r + 1)th first feature map;
[0291] Obtain a first target feature map according to the M first feature maps.
[0292] In another embodiment, the first image to be detected and the first target feature map are obtained through a first sub-network model, and the first sub-network model is obtained by training the first sub-network of the neural network;
[0293] The processor 1001 performs the training steps of the first sub-network model, including:
[0294] Perform image correction on N positive sample images in the sample image set through the first sub-network to obtain N second images to be detected, where the N positive sample images and the N second images to be detected are in one-to-one correspondence, and N is an integer greater than 1;
[0295] Extract and correct the features of N second images to be detected through the first sub-network, obtaining N second feature maps, with one-to-one correspondence between the N second images to be detected and the N second features Figure 1 One-to-one correspondence;
[0296] Determine the target loss based on the first loss determined from the N second images to be detected and the second loss determined from the N second feature maps;
[0297] Adjust the parameters of the neural network, iterate the positive sample image set to make the target loss converge, and obtain the first sub-network model.
[0298] In another embodiment, the processor 1001 executes the step of determining the first loss, including:
[0299] Randomly select two second images to be detected from the N second images to be detected;
[0300] Calculate the second distance between the two second images to be detected, and determine the second distance as the first loss.
[0301] In another embodiment, the neural network further includes a second sub-network and a third sub-network. The processor 1001 executes the step of determining the second loss, including:
[0302] Encode any second feature map a in the N second feature maps through the second sub-network to obtain a first set of feature vectors;
[0303] Encode any second feature map b in the shuffled N second feature maps through the third sub-network to obtain a second set of feature vectors;
[0304] Determine the second loss according to the first set of feature vectors and the second set of feature vectors.
[0305] In another embodiment, the processor 1001 executes encoding any second feature map a in the N second feature maps through the second sub-network to obtain a first set of feature vectors, including:
[0306] For each first position in any second feature map a, encode and project the feature at the first position through the second sub-network to obtain a first feature vector;
[0307] The first set of feature vectors is composed of the first feature vectors corresponding to each first position;
[0308] The processor 1001 executes encoding any second feature map b in the shuffled N second feature maps through the third sub-network to obtain a second set of feature vectors, including:
[0309] For each second position in any second feature map b, encode the feature at the second position through a third sub-network to obtain a second feature vector;
[0310] A second feature vector set is composed of second feature vectors corresponding to each second position.
[0311] In another embodiment, the processor 1001 executes to determine a second loss according to the first feature vector set and the second feature vector set, including:
[0312] Determine a target second feature map with the same sequence identifier as any second feature map a from the shuffled N second feature maps, and any second feature map b includes the target second feature map;
[0313] Obtain a target second feature vector set corresponding to the target second feature map, and the second feature vector set includes the target second feature vector set;
[0314] Determine a target second feature vector corresponding to the first feature vector in the first feature vector set from the target second feature vector set;
[0315] Determine the second loss according to the first feature vector and the corresponding target second feature vector.
[0316] In another embodiment, the processor 1001 executes to perform feature extraction and correction on N second images to be detected through a first sub-network to obtain N second feature maps, including:
[0317] For any one of the N second images to be detected, perform at least 2 operations of feature extraction and correction through the first sub-network to obtain a second feature map of any one of the N second images to be detected;
[0318] N second feature maps are composed of second feature maps of each of the N second images to be detected.
[0319] In another embodiment, after obtaining the first sub-network model, the processor 1001 is further configured to:
[0320] Process N positive sample images through the first sub-network model to obtain N second target feature maps, and the N positive sample images and the N second target features Figure 1 Are in one-to-one correspondence;
[0321] Perform Gaussian fitting on the features corresponding to the positions in the N second target feature maps to obtain the target Gaussian distribution of the positive sample images.
[0322] In another embodiment, the processor 1001 executes to process N positive sample images through the first sub-network model to obtain N second target feature maps, including:
[0323] The first sub-network model is used to perform image correction on N positive sample images to obtain N third images to be detected, and the N positive sample images correspond one-to-one with the N third images to be detected;
[0324] The first sub-network model is used to perform feature extraction and correction on the N third images to be detected to obtain N second target feature maps, and the N second target feature maps correspond one-to-one with the N third images to be detected;
[0325] Among them, the feature extraction and correction of any one of the N third images to be detected includes:
[0326] The first sub-network model is used to perform the operation of the first feature extraction and correction on any one of the third images to be detected to obtain the first third feature map;
[0327] The first sub-network model is used to perform the operation of the (r + 1)-th feature extraction and correction on the r-th third feature map to obtain the (r + 1)-th third feature map;
[0328] The first sub-network model is used to perform at least two operations of feature extraction and correction to obtain R third feature maps, where R is an integer greater than or equal to 2, and the R third feature maps include the first third feature map and the (r + 1)-th third feature map;
[0329] The first sub-network model is used to obtain the second target feature map corresponding to any one of the third images to be detected according to the R third feature maps.
[0330] Exemplarily, the electronic device may include but is not limited to a processor 1001, an input device 1002, an output device 1003, and a computer storage medium 1004. The input device 1002 may be a keyboard, a touch screen, etc., and the output device 1003 may be a speaker, a display, a radio frequency transmitter, etc. Those skilled in the art can understand that the schematic diagram is only an example of the electronic device, does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine some components, or different components.
[0331] It should be noted that since the processor 1001 of the electronic device implements the steps in the above image anomaly detection method when executing a computer program, the embodiments of the above image anomaly detection method are all applicable to the electronic device and can achieve the same or similar beneficial effects.
[0332] The embodiments of the present application further provide a computer storage medium (Memory). The computer storage medium is a memory device in an electronic device and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the terminal and, of course, the extended storage medium supported by the terminal. The computer storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor 1001 are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor 1001. In one embodiment, one or more instructions stored in the computer storage medium can be loaded and executed by the processor 1001 to implement the corresponding steps of the above-mentioned image anomaly detection method.
[0333] Exemplarily, the computer program of the computer storage medium includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0334] It should be noted that since the computer program of the computer storage medium implements the steps in the above-mentioned image anomaly detection method when executed by the processor, all embodiments of the above-mentioned image anomaly detection method are applicable to this computer storage medium and can achieve the same or similar beneficial effects.
[0335] The embodiments of the present application further provide a computer program product. Among them, the above computer program product includes a computer program, and the above computer program can operate to cause the computer to execute the steps in the above-mentioned image anomaly detection method. This computer program product can be a software installation package.
[0336] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image anomaly detection method, characterized in that, The method includes: Performing image correction on the original image to obtain a first image to be detected; Performing feature extraction and correction on the first image to be detected to obtain a first target feature map; Obtaining an anomaly detection result according to the first target feature map and the feature map of the positive sample image; The performing feature extraction and correction on the first image to be detected to obtain a first target feature map includes: Performing the operation of the first feature extraction and correction on the first image to be detected to obtain the first first feature map; Performing the operation of the (r + 1)-th feature extraction and correction on the r-th first feature map to obtain the (r + 1)-th first feature map, where r is an integer greater than or equal to 1; Performing at least two operations of feature extraction and correction to obtain M first feature maps, where M is an integer greater than or equal to 2, and the M first feature maps include the first first feature map and the (r + 1)-th first feature map; Obtaining the first target feature map according to the M first feature maps.
2. The method according to claim 1, wherein The obtaining an anomaly detection result according to the first target feature map and the feature map of the positive sample image includes: Obtaining the target Gaussian distribution corresponding to each position in the feature map of the positive sample image; Obtaining the anomaly score of each position in the first target feature map according to the feature of each position in the first target feature map and the target Gaussian distribution; Determining the anomaly score map composed of the anomaly scores of each position in the first target feature map as the anomaly detection result.
3. The method according to claim 2, wherein The obtaining the anomaly score of each position in the first target feature map according to the feature of each position in the first target feature map and the target Gaussian distribution includes: Calculating the first distance between the feature of each position in the first target feature map and the target Gaussian distribution corresponding to the position; Determining the first distance as the anomaly score of each position in the first target feature map.
4. The method according to any one of claims 1 to 3, characterized in that The performing image correction on the original image to obtain a first image to be detected includes: Obtaining a first transformation matrix; Aligning the original image by using the first transformation matrix to obtain the first image to be detected.
5. The method according to any one of claims 1 to 3, characterized in that, The performing feature extraction and correction on the first image to be detected to obtain a first target feature map includes: Performing feature extraction on the first image to be detected to obtain a feature map to be corrected; Obtaining a second transformation matrix; Aligning the feature map to be corrected by using the second transformation matrix to obtain a first feature map; Obtaining the first target feature map based on the first feature map.
6. The method according to claim 1, wherein The first image to be detected and the first target feature map are obtained through a first sub-network model, and the first sub-network model is obtained by training a first sub-network of a neural network; The first sub-network model is trained by the following steps: Performing image correction on N positive sample images in a sample image set through the first sub-network to obtain N second images to be detected, where the N positive sample images and the N second images to be detected are in one-to-one correspondence, and N is an integer greater than 1; Performing feature extraction and correction on the N second images to be detected through the first sub-network, obtaining N second feature maps, where the N second images to be detected correspond one-to-one with the N second feature maps; Determining a target loss according to a first loss determined from the N second images to be detected and a second loss determined from the N second feature maps; Adjusting parameters of the neural network and iterating the positive sample image set to make the target loss converge, obtaining the first sub-network model.
7. The method according to claim 6, wherein The determining step of the first loss includes: Randomly selecting two second images to be detected from the N second images to be detected; Calculating a second distance between the two second images to be detected and determining the second distance as the first loss.
8. The method according to claim 6, wherein The neural network further includes a second sub-network and a third sub-network. The determining step of the second loss includes: Encoding an arbitrary second feature map a among the N second feature maps through the second sub-network to obtain a first set of feature vectors; Encoding an arbitrary second feature map b among the shuffled N second feature maps through the third sub-network to obtain a second set of feature vectors; Determining the second loss according to the first set of feature vectors and the second set of feature vectors.
9. The method according to claim 8, wherein The encoding an arbitrary second feature map a among the N second feature maps through the second sub-network to obtain a first set of feature vectors includes: For each first position in the arbitrary second feature map a, encoding and projecting the feature at the first position through the second sub-network to obtain a first feature vector; Composing the first set of feature vectors from the first feature vectors corresponding to each first position; The encoding an arbitrary second feature map b among the shuffled N second feature maps through the third sub-network to obtain a second set of feature vectors includes: For each second position in the arbitrary second feature map b, encoding the feature at the second position through the third sub-network to obtain a second feature vector; Composing the second set of feature vectors from the second feature vectors corresponding to each second position.
10. The method according to claim 9, wherein The determining the second loss according to the first set of feature vectors and the second set of feature vectors includes: Determining a target second feature map with the same sequence identifier as the arbitrary second feature map a from the shuffled N second feature maps, where the arbitrary second feature map b includes the target second feature map; Obtaining a target second set of feature vectors corresponding to the target second feature map, where the second set of feature vectors includes the target second set of feature vectors; Determining a target second feature vector corresponding to the first feature vector in the first set of feature vectors from the target second set of feature vectors; Determining the second loss according to the first feature vector and the corresponding target second feature vector.
11. The method according to any one of claims 6 - 10, characterized in that, The performing feature extraction and correction on the N second images to be detected through the first sub-network to obtain N second feature maps includes: For any one of the N second images to be detected, at least two operations of feature extraction and correction are performed through the first sub-network to obtain a second feature map of any one of the second images to be detected; The N second feature maps are composed of the second feature maps of each of the N second images to be detected.
12. The method according to any one of claims 6-10, characterized in that After obtaining the first sub-network model, the method further includes: Processing the N positive sample images through the first sub-network model to obtain N second target feature maps, where the N positive sample images and the N second target feature maps correspond one by one; Performing Gaussian fitting on the features corresponding to the positions in the N second target feature maps to obtain the target Gaussian distribution of the positive sample images.
13. The method according to claim 12, characterized in that, The processing the N positive sample images through the first sub-network model to obtain N second target feature maps includes: Performing image correction on the N positive sample images through the first sub-network model to obtain N third images to be detected, where the N positive sample images and the N third images to be detected correspond one by one; Performing feature extraction and correction on the N third images to be detected through the first sub-network model to obtain the N second target feature maps, where the N second target feature maps and the N third images to be detected correspond one by one; Among them, the feature extraction and correction of any one of the N third images to be detected includes: Performing the first operation of feature extraction and correction on any one of the third images to be detected through the first sub-network model to obtain the first third feature map; Performing the (r + 1)-th operation of feature extraction and correction on the r-th third feature map through the first sub-network model to obtain the (r + 1)-th third feature map; Performing at least two operations of feature extraction and correction through the first sub-network model to obtain R third feature maps, where R is an integer greater than or equal to 2, and the R third feature maps include the first third feature map and the (r + 1)-th third feature map; Obtaining the second target feature map corresponding to any one of the third images to be detected through the first sub-network model according to the R third feature maps.
14. An image anomaly detection device, characterized in that, The device includes an acquisition unit and a processing unit; The acquisition unit is configured to perform image correction on the original image to obtain a first image to be detected; The processing unit is configured to perform feature extraction and correction on the first image to be detected to obtain a first target feature map; The processing unit is further configured to obtain an anomaly detection result according to the first target feature map and the feature map of the positive sample image; In terms of performing feature extraction and correction on the first image to be detected to obtain a first target feature map, the processing unit is specifically configured to: Performing the first operation of feature extraction and correction on the first image to be detected to obtain the first first feature map; Performing the (r + 1)-th operation of feature extraction and correction on the r-th first feature map to obtain the (r + 1)-th first feature map, where r is an integer greater than or equal to 1; Perform the operations of feature extraction and correction at least twice to obtain M first feature maps, where M is an integer greater than or equal to 2, and the M first feature maps include the first first feature map and the (r + 1)-th first feature map; Obtain the first target feature map according to the M first feature maps.
15. An electronic device, comprising an input device and an output device, characterized in that, It further includes a processor and a computer storage medium; The processor is adapted to implement one or more instructions; and, The computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the method according to any one of claims 1-13.
16. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the method according to any one of claims 1-13.
Citation Information
Patent Citations
Anomaly detection method based on passive-aggressive Gaussian online learning
CN107832716A
Printed matter defect detection method based on feature registration and gradient shape matching fusion
CN112508826A