Image anomaly detection method, system and equipment based on resolution scaling and medium
By introducing multi-scale resolution image object detection network and dynamic scale sampling technology in image anomaly detection, the shortcomings of existing methods in computing complexity and generalization capabilities are solved, and more efficient and flexible abnormality detection capabilities are achieved.
Patent Information
- Application Number
- CN202510348140.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
The existing anomaly image detection methods have shortcomings in terms of computational complexity and generalization capabilities, and are difficult to effectively apply in monitoring systems with high real-time requirements or equipment with limited computing resources. The image detection capabilities of the mixture of new anomalies and multiple anomaly factors are limited.
A multi-scale resolution image object detection network is introduced, and image multi-scale resolution scaling technology is used to achieve high similarity of normal sample detection results and high differences between abnormal sample prediction results through dynamic scale sampling and weighted fusion feature maps, thereby improving detection accuracy and generalization capabilities.
It significantly improves the accuracy and generalization ability of image abnormality detection, optimizes computing efficiency, reduces error detection rate, and supports a variety of hardware devices and application scenarios.
Smart Images

Figure CN120219848A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image detection, and particularly relates to an image anomaly detection method, system, device, medium and system based on resolution scaling. Background Art
[0002] High computational complexity is a significant problem faced by anomaly image detection in the field of object detection. Currently, many advanced anomaly image detection methods were initially developed in the classification field and then tried to be applied to the object detection field, but they all face the problem of huge consumption of computing resources and time. For example, multi-scale feature fusion detection models based on deep learning and anomaly score calculation based on complex probability models require a large amount of computing resources in the classification field. When migrated to the object detection field, due to the need to process objects at different positions and of different sizes in the image, the computational amount increases significantly.
[0003] Existing methods already require a large amount of computation when classifying samples in the classification field. When applied to the object detection field, due to the complex model structure, not only do they need to perform multiple iterative calculations on a large amount of image data, but also need to locate and judge multiple objects in the image, and their computational amount is several times that of the training of ordinary object detection models. This high computational requirement makes it difficult to apply this method in monitoring systems with high real-time requirements, or in embedded devices such as mobile devices and small intelligent cameras with limited computing resources, greatly limiting its popularization in actual scenarios.
[0004] Thirdly, insufficient generalization ability is another important defect of existing methods. Most anomaly image detection methods are designed for specific types of anomaly images, such as only for anomaly images generated in specific scenarios such as sudden light changes and object occlusion. When faced with new types of anomalies or images with a mixture of multiple anomaly factors, the performance of these methods will be greatly reduced. Some studies have pointed out that existing detection methods often overfit specific anomaly patterns and lack adaptability to unknown anomaly situations. For example, if only anomaly images caused by daytime light changes are considered during model training, when encountering anomaly images generated at night due to low light and accompanied by complex environments such as fog, the model will be difficult to accurately detect. This limitation makes it difficult for existing methods to cope with the increasingly diverse anomaly image scenarios, greatly limiting their reliability and effectiveness in practical applications.
[0005] In summary, there are still obvious deficiencies in existing anomaly detection methods in terms of computational complexity, generalization ability, etc. Summary of the Invention
[0006] The object of the present invention is to solve the problems of insufficient generalization ability and poor detection accuracy of existing anomaly detection technologies in the field of image classification. By introducing a multi-scale resolution image object detection network and using the image multi-scale resolution scaling technology, the present invention can achieve high similarity of the detection results of normal samples and high difference of the prediction results of abnormal samples, thereby significantly improving the detection accuracy and generalization ability of anomaly detection.
[0007] To solve the above technical problems, the present invention proposes an image anomaly detection method, system, device and medium based on resolution scaling.
[0008] In the first aspect, the present invention proposes an image anomaly detection method based on resolution scaling, the method comprising:
[0009] S1: Obtain an image to be detected, perform dynamic scale sampling on the image to be detected to obtain multi-scale images;
[0010] S2: Input the multi-scale feature maps into a trained multi-scale resolution image object detection model to obtain a weighted fusion feature map;
[0011] S3: Measure the distribution difference between the feature maps in a preset feature library and the weighted fusion feature map to obtain a feature distribution difference result, and judge the feature distribution difference result and a dynamic threshold to obtain a detection result.
[0012] Further, the training process of the multi-scale resolution image object detection network comprises:
[0013] Obtain a training sample set, perform dynamic scale sampling on the training samples to obtain multi-scale images;
[0014] Extract features from the multi-scale images to obtain multi-branch preliminary feature maps;
[0015] Calculate the weight of each branch according to the multi-scale images and the multi-branch preliminary feature maps;
[0016] According to the multi-branch preliminary feature maps and the weight of each branch, perform weighted fusion on the feature maps of each branch to obtain a weighted fusion feature map and store it in the feature library;
[0017] Use the object detection head to predict the weighted fusion feature map to obtain a prediction result, calculate the total loss between the prediction result and the actual label of the training sample, and backpropagate through this loss to optimize the network parameters.
[0018] Further, the trained multi-scale resolution image object detection model includes: a trained multi-branch feature extraction network, a trained weight prediction network WPN, and a trained feature fusion network layer; the multi-scale resolution image object detection network includes a multi-branch feature extraction network, a weight prediction network WPN, a feature fusion network layer, and an object prediction head.
[0019] In a second aspect, the present invention proposes an image anomaly detection system, which includes:
[0020] An input module, configured to collect the image to be detected;
[0021] A multi-scale transformation module, configured to perform dynamic scale sampling on the image to be detected to obtain multi-scale images;
[0022] A feature extraction module, configured to extract features from the multi-scale images to obtain multi-scale preliminary feature maps;
[0023] A multi-branch weight calculation module, which is internally provided with a pre-trained weight prediction network WPN, configured to calculate the weights of each branch;
[0024] A feature fusion module, configured to perform feature fusion based on the multi-scale preliminary feature maps and the weights of each branch to obtain a weighted fusion feature map;
[0025] A feature library, configured to store the feature maps of normal image samples or abnormal image samples;
[0026] An anomaly detection module, configured to judge the weighted fusion feature map against the feature maps of normal image samples in the feature library to obtain an anomaly detection result;
[0027] An output module, configured to output the anomaly detection result.
[0028] In a third aspect, the present invention proposes an electronic device, including:
[0029] A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the method for detecting image anomalies in object detection based on resolution scaling as proposed in the first aspect.
[0030] In a fourth aspect, the present invention proposes a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for detecting image anomalies based on multi-scale resolution scaling as proposed in the first aspect.
[0031] Advantages of the present invention: By introducing a multi-scale resolution image object detection network and a dynamic scale sampling technique, the present invention significantly improves the accuracy and generalization ability of image anomaly detection. Specifically, the multi-scale feature extraction and weighted fusion mechanism in the multi-scale resolution image object detection network can capture multi-level detailed information in the image, magnify the feature differences between abnormal samples and normal samples, thereby improving the detection accuracy; the dynamic scale sampling and feature distribution difference measurement enhance the adaptability of the model to diverse data and solve the problem of insufficient generalization ability of traditional methods. In addition, the present invention optimizes the computing efficiency, reduces the false detection rate, and supports multiple hardware devices and application scenarios through parallel computing, end-to-end training, and dynamic threshold judgment, and has important practical value and promotion potential. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of an existing typical abnormal picture example;
[0033] Figure 2 It is a flowchart of the steps of the embodiment of the present invention;
[0034] Figure 3 It is an overall flowchart of the embodiment of the present invention;
[0035] Figure 4 It is a flowchart of the network training steps of the multi-scale resolution image object detection network in the embodiment of the present invention;
[0036] Figure 5 It is a schematic diagram of the process of feature distribution difference measurement in the embodiment of the present invention;
[0037] Figure 6 It is a schematic diagram of the system structure of the image anomaly detection system in the embodiment of the present invention. Detailed Embodiments
[0038] The terms "first", "second", "third", "fourth", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, and this is only a way of distinguishing objects with the same attributes when describing the embodiments of this application.
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] Figure 1It is an example diagram of typical existing abnormal images.
[0041] Refer to Figure 1 As shown, in various application scenarios, the effects of existing common abnormal images include: blurring, smudging, thermal noise, quantization noise, compression artifacts, Gaussian noise, salt-and-pepper noise, motion blur, low-light noise, shadow interference, raindrop effect, and fogging effect, etc. The application scenarios of these abnormal images cover various situations such as natural environments (such as rain and fog, shadows), sensor problems (such as noise, quantization errors), and distortions during the image processing process (such as compression artifacts, motion blur). Various application scenarios include: autonomous driving environment, rain and fog environment, security monitoring scenario, industrial quality inspection scenario, etc.
[0042] To solve the existing technology and improve the detection accuracy of image anomaly detection, the embodiment of the present invention proposes an object detection image anomaly detection method based on resolution scaling.
[0043] Figure 2 It is the step flowchart of the embodiment of the present invention.
[0044] Figure 3 It is the overall flowchart of the embodiment of the present invention. Figure 3Among them, the overall process of the embodiments of the present invention includes a training stage and a detection stage. The training stage includes: obtaining training samples, dynamically scale-sampling the training samples to obtain multi-scale images, inputting the multi-scale images into a feature extraction network (n branch networks) to respectively obtain multi-scale preliminary feature maps (F1, F2, F3, …, Fn); inputting the multi-scale images and the multi-scale preliminary feature maps (F1, F2, F3, …, Fn) into a weight prediction network (Weight Prediction Network, WPN) to calculate the weights (W1, W2, W3, …, Wn) of each branch network; inputting the multi-scale preliminary feature maps (F1, F2, F3, …, Fn) and the weights (*W1, *W2, *W3, …, *Wn) output by each branch network into a feature fusion network layer for feature weighted fusion to obtain a fused feature map; storing the fused feature map in a feature library; using an object detection head to perform anomaly detection on the fused feature map, calculating the total loss, and backpropagating according to the total loss to optimize network parameters. The prediction stage includes: obtaining an image to be detected, dynamically scale-sampling the image to be detected to obtain multi-scale images, inputting the multi-scale images into a feature extraction network (n branch networks) to respectively obtain multi-scale preliminary feature maps (F1, F2, F3, …, Fn); inputting the multi-scale images and the multi-scale preliminary feature maps (F1, F2, F3, …, Fn) into a weight prediction network (Weight Prediction Network, WPN) to calculate the weights (W1, W2, W3, …, Wn) of each branch network; inputting the multi-scale preliminary feature maps (F1, F2, F3, …, Fn) and the weights (*W1, *W2, *W3, …, *Wn) output by each branch network into a feature fusion network layer for feature weighted fusion to obtain a weighted fused feature map; using a difference metric module to perform multi-modal distribution difference metric on the weighted fused feature map and the feature maps in a preset feature library to obtain a feature difference metric result; judging the feature distribution difference result with a dynamic threshold to obtain a detection result.
[0045] Referring to Figure 2 、 Figure 3 As shown, the method for anomaly detection of target detection images based on resolution scaling includes:
[0046] S1: Obtain an image to be detected, and perform dynamic scale sampling on the image to be detected to obtain multi-scale images.
[0047] Before performing dynamic scale sampling, perform feature analysis on the targets in each image. Calculate the size (area) of the targets, the proportion of the targets in the image, and the complexity of the targets (such as the shape, texture, etc. of the targets). Dynamically determine the scaling ratio of the image according to the target feature analysis results.
[0048] Use EfficientNet to extract features from the image to be detected and predict the scale distribution and complexity score of the target. Dynamically determine the scaling ratio of the image to be detected according to the scale distribution and complexity score of the target to obtain multi-scale images.
[0049] Exemplary:
[0050] 1. For small-object-dominated images: If the scale distribution S shows that the image contains a large number of small objects (e.g., s1 > 0.8), increase the high-resolution scaling ratio (e.g., 1x, 2x, 3x, 4x).
[0051] 2. For large-object-dominated images: If the scale distribution S shows that the image is dominated by large objects (e.g., s k > 0.8), increase the low-resolution scaling ratio (e.g., 0.1x, 0.25x, 0.5x, 0.7x).
[0052] 3. For mixed-object images: If the scale distribution shows that the image contains both small and large objects, adopt multiple sets of scaling ratios (e.g., 0.5x, 1x, 2x, 3x).
[0053] It should be noted that EfficientNet is an efficient convolutional neural network (CNN) architecture. This model is based on a compound scaling method that simultaneously considers three dimensions: depth (number of network layers), width (number of filters per layer), and resolution (input image size) to optimize network performance. EfficientNet mainly includes the MBConv module and the compound scaling module.
[0054] S2: Input the multi-scale feature maps into the trained multi-scale resolution image object detection network to obtain a weighted fusion feature map.
[0055] Refer to Figure 3 As shown, the multi-scale resolution image object detection network includes a multi-branch feature extraction network, a weight prediction network WPN, a feature fusion network layer, and an object prediction head. After the multi-scale resolution image object detection network is trained, the network parameters are fixed; in the prediction stage, the trained multi-scale resolution image object detection network includes: the trained multi-branch feature extraction network, the trained weight prediction network WPN, and the trained feature fusion network layer; while the object prediction head is not used in the prediction stage.
[0056] Figure 4 This is the flowchart of the network training steps of the multi-scale resolution image object detection network in the embodiments of the present invention.
[0057] In a preferred embodiment, refer to Figure 3 、 Figure 4As shown in the figure, the training process of the multi-scale resolution image target detection network includes:
[0058] S201: Obtain a training sample set, perform dynamic scale sampling on the training samples to obtain multi-scale images.
[0059] Widely collect a data set containing normal images and various abnormal images (such as blurred, damaged, fogged, etc.), and accurately annotate the collected images. For normal images, mark them as the normal category (i.e., the normal label, such as label = 1); for abnormal images (i.e., abnormal labels, such as label = 0), clearly mark the abnormal type (such as blurred, damaged, fogged, etc.), and perform bounding box annotation on the abnormal area. The collected image data comes from public image databases, actual business scenarios (such as product images on industrial production lines, security monitoring images, etc.). Ensure that the data set has diversity, covering images with different scenarios, lighting conditions, shooting angles and resolutions to enhance the generalization ability of the model. Take the annotated image data set as the training sample set of the embodiment of the present invention.
[0060] Perform dynamic scale sampling on the training samples. The dynamic scale sampling method is the same as that for the image to be detected in S1 above and will not be elaborated here.
[0061] S202: Input the multi-scale images into the multi-branch feature extraction network respectively to obtain multi-branch preliminary feature maps.
[0062] The specific number of branches n of the multi-branch feature extraction network is determined according to the number of branches of the multi-scale feature maps. The feature extraction network can be a convolutional neural network layer, such as EfficientNet.
[0063] In the training stage, the feature extraction network Efficient Net loads the pre-trained weights, removes the classification head, retains the feature extraction part, and adds a task head: (1) Add a fully connected layer to the scale distribution prediction head to output the scale distribution of the target. Use the Softmax function to normalize the output into a probability distribution. (2) Add a fully connected layer to output the complexity score of the target, and use the Sigmoid function to limit the output between 0 and 1.
[0064] In the training stage, the loss function adopted by the feature extraction network Efficient Net is:
[0065] L total =L scale +λL complexity
[0066] In the formula, L total represents the total loss of the feature extraction network Efficient Net, L scaleDenotes the scale distribution loss, and calculates the error of scale distribution prediction using the Cross-Entropy Loss; L complexity Denotes the complexity score loss, and calculates the error of complexity score prediction using the Mean Squared Error (MSE); λ represents the weight coefficient of the complexity score loss. For example, λ = 0.5.
[0067] For example, the specific number of branches n of the multi-branch feature extraction network is n = 4; feature maps (F1, F2, F3, F4) at different resolutions are extracted; these feature maps contain information about the target at different scales, but the importance of the features may vary depending on the input image.
[0068] S203: Input the multi-scale image and the preliminary feature maps of multiple branches into the Weight Prediction Network WPN to calculate the weight of each branch.
[0069] The input of the Weight Prediction Network (WPN) WPN includes the multi-scale image and the preliminary feature maps (F1, F2, F3, …, Fn) of each branch. In order to integrate this information from different sources, the input needs to be appropriately processed. For example, downsample the original image or perform feature extraction to match its dimension with that of the feature map; perform operations such as global average pooling on the feature map to reduce the computational amount and extract the global information of the features.
[0070] WPN can adopt the structure of a Multi-Layer Perceptron (MLP), and its specific structure can be designed according to the actual situation. WPN can include multiple fully connected layers, with activation functions (such as ReLU) interspersed in the middle to introduce non-linearity.
[0071] For example, if the input dimension of WPN is 4 and the output dimension is 4 (corresponding to the weights of 4 branches, i.e., n = 4), then the forward propagation process of WPN can be expressed as:
[0072] h1 = ReLU(w1x + b1)
[0073] h2 = ReLU(w1x + b1)
[0074] …
[0075] w = Softmax(w n h n -1 + b n )
[0076] where x is the input vector (containing image and feature map information), W i and b iThey are the weight matrix and bias vector of the i-th layer respectively. w = [w1, w2, w3, w4] is the output weight vector and satisfies The non-negativity and normalization of the weights are ensured by the Softmax function.
[0077] The weight prediction network WPN outputs the weights of each branch (*W1, *W2, *W3, …, *Wn). The weight of each branch reflects the contribution degree of the features of this branch to the final fused features under the current input image. For example, if the input image contains a large number of small targets, the features of the low-resolution branch may be relatively unimportant, and WPN will calculate a lower weight; while the high-resolution branch can capture the detailed information of small targets, and its weight will be relatively high.
[0078] S204: Input the weights of each branch and the preliminary feature maps of multiple branches into the feature fusion network layer, perform weighted fusion on the feature maps of each branch to obtain a weighted fusion feature map, and store it in the feature library.
[0079] Specifically, the preliminary feature maps (F1, F2, F3, …, Fn) of each branch are weighted and fused with the weights (*W1, *W2, *W3, …, *Wn) of each branch to obtain a weighted fusion feature map.
[0080] The weighted fusion feature maps obtained according to the training sample set are stored in the feature library to form a preset feature library. Since there are both normal images and abnormal images in the training sample set, there are both normal feature maps and abnormal feature maps and their labels in the preset feature library.
[0081] S205: Input the weighted fusion feature map into the target detection head to obtain a prediction result, calculate the total loss between the prediction result and the actual label of the training sample, and backpropagate through this loss to optimize the network parameters.
[0082] Exemplarily, the target detection head can be the detection head of Faster R-CNN, YOLO or SSD. Each detection head includes a classification head and a regression head. The classification head is used to predict the target class score in the training sample, and the regression head is used to predict the offset of the bounding box in the training sample.
[0083] In the training stage, the total loss adopted by the multi-scale resolution image target detection network includes classification loss, regression loss and scale consistency loss, and its specific calculation formula is:
[0084] L = L class + λ reg L reg + λ scale L scale
[0085] Wherein, \(L\) represents the total loss of the multi-scale resolution image target detection network, \(L\) class represents the classification loss, \(L\) reg represents the regression loss, \(\lambda\) scale represents the weight coefficient of the regression loss, \(L\) scale represents the scale consistency loss, \(\lambda\) scale represents the weight coefficient of the scale consistency loss, which is used to balance the influence of different types of losses.
[0086] The calculation formula of the classification loss is as follows:
[0087]
[0088] Wherein, \(y\) represents the true class label of the training sample, represents the predicted class probability distribution of the training sample.
[0089] If the target detection head locates to the abnormal area of the training sample, the regression loss function (such as smooth L1 loss) is used to calculate the error between the predicted position coordinates and the true position coordinates of the training sample.
[0090] The calculation formula of the regression loss is as follows:
[0091]
[0092] Wherein, \((x, y, w, h)\) represents the true position coordinates of the training sample, \(x\) represents the abscissa of the center point of the true bounding box, \(y\) represents the ordinate of the center point of the true bounding box, \(w\) represents the width of the true bounding box, and \(h\) represents the height of the true bounding box, represents the predicted position coordinates of the training sample, represents the abscissa of the center point of the predicted bounding box, represents the ordinate of the center point of the predicted bounding box represents the width of the bounding box, represents the height of the center point of the bounding box.
[0093] If the training sample is a normal image, there is no need to calculate the regression loss.
[0094] In order to ensure the feature consistency between different scale branches, the scale consistency loss is introduced. The cosine similarity or mean square error is used to calculate the difference between the features output by different scale branches.
[0095] The calculation formula of the scale consistency loss is as follows:
[0096] \(L\) scale = MSE(F1, F2) or
[0097] \(L\) scale= 1 - Cosine_similarity(F1,F2)
[0098] In the formula, F1 and F2 respectively represent the feature maps of two different scale branches.
[0099] The scale consistency losses between all pairs of different scale branches are weighted and summed to obtain the total scale consistency loss.
[0100] S3: Measure the distribution difference between the feature map in the preset feature library and the weighted fusion feature map to obtain a feature distribution difference result, and judge the feature distribution difference result against a dynamic threshold to obtain a detection result.
[0101] The preset feature library is a set of feature maps extracted from the training sample set during the training process of the multi-scale resolution image target detection network.
[0102] To improve the accuracy of feature map distribution difference measurement, the embodiments of the present invention organically combine traditional distance measurement, statistical test methods, divergence measurement, and dynamic threshold adjustment, and propose a multi-modal distribution difference measurement algorithm.
[0103] Figure 5 This is a schematic diagram of the process of measuring the feature distribution difference in the embodiments of the present invention.
[0104] Refer to Figure 5 As shown, the specific process of measuring the feature distribution difference includes:
[0105] S301: Calculate the Euclidean distance between the weighted fusion feature map and the normal sample feature map in the preset feature library in sequence to obtain a first feature distribution difference result and judge it against a dynamic threshold to obtain a first detection result; if the first detection result is an abnormal image, output the detection result as an abnormal image.
[0106] Use the Euclidean distance for preliminary screening to quickly evaluate the absolute difference between the new sample (i.e., the weighted fusion feature map of the image to be detected) and the historical normal sample (i.e., the normal sample feature map in the preset feature library) in the feature space.
[0107] Taking the Euclidean distance as an example, it is one of the most commonly used distance measurement methods, and its calculation formula is for two n-dimensional vectors x = (x1, x2,..., x n ) and y = (y1, y2,..., y n ), the Euclidean distance This measurement method intuitively reflects the absolute distance between two feature points in the n-dimensional space and performs well in dealing with some simple data distributions and feature relationships.
[0108] After calculating the Euclidean distance, a first feature distribution difference result is obtained, which is judged against a dynamic threshold to obtain a first detection result. The first detection result includes a normal image or an abnormal image.
[0109] To overcome the limitations of traditional distance metrics, a statistical test method can be used to quantify the deviation degree of the feature distribution of a new sample feature (i.e., the weighted fusion feature map of the image to be detected) from that of historical normal samples (i.e., the normal sample feature maps in the preset feature library). Statistical test is a method of inferring population characteristics based on sample data. By comparing the statistical characteristics of the distribution of new sample features with the distribution of historical normal samples, it is possible to more accurately determine whether the new sample is abnormal.
[0110] S302: If the first detection result is a normal image, calculate the chi-square statistic between the weighted fusion feature map and the normal sample feature maps in the preset feature library to obtain a second feature distribution difference result, and judge it against the dynamic threshold to obtain a second detection result; if the second detection result is an abnormal image, output the detection result as an abnormal image.
[0111] The chi-square test is a commonly used statistical test method for comparing the difference between observed frequencies and expected frequencies. Its process includes:
[0112] In feature distribution comparison, divide the weighted fusion feature map into several intervals, and count the number of samples of the new sample (i.e., the weighted fusion feature map) in each interval as the observed frequency; calculate the expected frequency of each interval according to the distribution of historical normal samples (i.e., the normal sample feature maps in the preset feature library).
[0113] Calculate the chi-square statistic, and its calculation formula is:
[0114]
[0115] where x 2 represents the chi-square statistic, O i represents the observed frequency of the i-th interval in the weighted fusion feature map, and E i represents the expected frequency of the i-th interval of the normal sample feature map in the preset feature library.
[0116] According to the chi-square distribution table and a preset significance level (such as α = 0.01), judge whether there is a significant difference between the new sample feature distribution and the historical normal sample distribution. If the chi-square statistic is greater than the critical value, reject the null hypothesis and consider that the new sample feature distribution has a significant deviation from the historical normal sample distribution.
[0117] After the chi-square test, a second feature distribution difference result is obtained, which is judged against the dynamic threshold to obtain a second detection result. The second detection result includes a normal image or an abnormal image.
[0118] S303: If the second detection result is a normal image, calculate the divergence between the weighted fusion feature map and the normal sample feature map in the preset feature library to obtain a third feature distribution difference result, and compare it with a dynamic threshold to obtain a third detection result, where the third detection result is an abnormal image or a normal image.
[0119] Specifically, regard the features of the weighted fusion feature map as a probability distribution, and calculate the KL divergence or JS divergence between the feature distribution of the weighted fusion feature map and the historical normal distribution (i.e., the feature distribution of the normal sample feature map in the preset feature library). Through the KL divergence or JS divergence, not only can single-point differences be captured, but also the overall shift of the feature distribution can be discovered (for example, the feature variance of abnormal samples significantly increases at certain scales).
[0120] Through steps S301 - S303, the feature distribution differences between the weighted fusion feature map and the normal sample feature map in the preset feature library are calculated in a progressive manner, which can not only effectively improve the accuracy of image anomaly detection, but also take into account the computational efficiency and improve the efficiency of anomaly detection.
[0121] In image anomaly detection, traditional methods usually use a fixed threshold to determine whether a sample is anomalous. This method is simple and intuitive. By setting a fixed threshold, when a certain feature value of a sample exceeds this threshold, the sample is considered anomalous. However, in practical applications, the distribution of data often changes over time, and this phenomenon is called data drift. To avoid the inadaptability of the fixed threshold to data drift, the anomaly threshold can be dynamically determined by combining the confidence interval. The confidence interval refers to the range within which the population parameter may take values at a certain confidence level. Among them, the 3σ principle is a commonly used method for determining the confidence interval based on the normal distribution. In the normal distribution, about 99.7% of the data falls within the range of the mean ± 3 standard deviations. The mean and standard deviation of each scale feature can be calculated according to the feature distribution of historical normal samples, and then the 3σ principle is used to determine the anomaly threshold. Specifically, for each scale of feature, calculate its mean and standard deviation in historical normal samples. When a new sample is input, calculate the deviation degree of the feature value of this sample at this scale from the mean. If the deviation degree exceeds the threshold, it is considered that the feature of this sample at this scale is anomalous. As new samples are continuously input, the statistical information (such as the mean and standard deviation) of historical normal samples can be updated regularly, so as to dynamically adjust the anomaly threshold. In this way, even if the data distribution drifts, the anomaly threshold can adapt to this change in a timely manner, ensuring the accuracy and stability of anomaly detection. In addition, the range of the confidence interval can be adjusted according to different application scenarios and risk preferences, such as using the 2σ principle or the 4σ principle, etc. When using the 2σ principle, the confidence interval is ± 2 times the standard deviation, and about 95.4% of the data falls within this interval. At this time, the sensitivity of anomaly detection is relatively high, but the false detection rate may also increase accordingly; when using the 4σ principle, the confidence interval is ± 4 times the standard deviation, and about 99.994% of the data falls within this interval. At this time, the false detection rate of anomaly detection is relatively low, but some anomalous samples may be missed. Therefore, it is necessary to select an appropriate confidence interval and dynamic threshold adjustment strategy according to the specific situation.
[0122] Based on the same inventive concept, an embodiment of the present invention proposes an image anomaly detection system, which has the same or similar technical features as the above-mentioned image anomaly detection method based on multi-scale resolution scaling; for the same or similar technical features, they will not be elaborated below.
[0123] Figure 6 It is a schematic structural diagram of the image anomaly detection system in the embodiment of the present invention.
[0124] Refer to Figure 6 As shown, the image anomaly detection system includes:
[0125] An input module, used to collect the image to be detected;
[0126] A multi-scale transformation module, which is used to perform dynamic scale sampling on the image to be detected to obtain multi-scale images;
[0127] A feature extraction module, which is used to extract features from the multi-scale images to obtain multi-scale preliminary feature maps;
[0128] A multi-branch weight calculation module, which is internally provided with a pre-trained weight prediction network WPN and is used to calculate the weights of each branch;
[0129] A feature fusion module, which is used to perform feature fusion based on the multi-scale preliminary feature maps and the weights of each branch to obtain a weighted fusion feature map;
[0130] A feature library, which is used to store the feature maps of normal image samples or abnormal image samples;
[0131] An anomaly detection module, which is used to judge the weighted fusion feature map and the feature maps of normal image samples in the feature library to obtain an anomaly detection result;
[0132] An output module, which is used to output the anomaly detection result.
[0133] An embodiment of the present invention provides an electronic device, including:
[0134] A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned image anomaly detection method based on resolution scaling is implemented.
[0135] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-mentioned image anomaly detection method based on multi-scale resolution scaling is implemented.
[0136] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, and the storage medium can include: ROM, RAM, magnetic disk, optical disk, etc.
[0137] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image anomaly detection method based on resolution scaling, characterized in that: include: Acquire an image to be detected, perform dynamic scale sampling on the image to be detected, and obtain a multi-scale image; Input the multi-scale feature map into the trained multi-scale resolution image object detection network to obtain a weighted fusion feature map; The distribution difference between the feature map in the preset feature library and the weighted fusion feature map is measured to obtain the feature distribution difference result, and the feature distribution difference result is judged with the dynamic threshold to obtain the detection result.
2. The image anomaly detection method based on resolution scaling according to claim 1, characterized in that: Efficient Net is used to extract features from the image to be detected and predict the scale distribution and complexity score of the target. The scaling ratio of the image to be detected is dynamically determined according to the scale distribution and complexity score of the target to obtain a multi-scale image.
3. The image anomaly detection method based on resolution scaling according to claim 1, characterized in that: The trained multi-scale resolution image target detection network includes: a trained multi-branch feature extraction network, a trained weight prediction network WPN and a trained feature fusion network layer; the multi-scale resolution image target detection network includes a multi-branch feature extraction network, a weight prediction network WPN, a feature fusion network layer and a target prediction head.
4. The image anomaly detection method based on resolution scaling according to claim 1, characterized in that: The training process of the multi-scale resolution image object detection network includes: Obtain a training sample set, perform dynamic scale sampling on the training samples, and obtain a multi-scale image; Extract features from multi-scale images and obtain preliminary feature maps of multiple branches; Calculating the weight of each branch according to the multi-scale image and the preliminary feature map of the multiple branches; According to the preliminary feature maps of multiple branches and the weight of each branch, the feature maps of each branch are weighted fused to obtain a weighted fused feature map and store it in the feature library; The target detection head is used to predict the weighted fusion feature map to obtain the prediction result. The total loss between the prediction result and the actual label of the training sample is calculated. The loss is back-propagated to optimize the network parameters.
5. The image anomaly detection method based on resolution scaling according to claim 2, characterized in that: The total loss used by the Efficient Net during network training is: THE total =L scale +λL complexity Where, L total Represents the total loss of the feature extraction network Efficient Net, L scale represents the scale distribution loss, L complexity represents the complexity score loss, and λ represents the weight coefficient of the complexity score loss.
6. The image anomaly detection method based on resolution scaling according to claim 1, characterized in that: The loss function used by the multi-scale resolution image target detection network during training is: L=L class +λ reg L reg +λ scale L scale Where L represents the total loss of the multi-scale resolution image target detection network, L class represents the classification loss, L reg represents the regression loss, λ scale Represents the weight coefficient of regression loss, L scale represents the scale consistency loss, λ scale Represents the weight coefficient of scale consistency loss.
7. The image anomaly detection method based on resolution scaling according to claim 1, characterized in that: The distribution difference between the feature map in the preset feature library and the weighted fusion feature map is measured. The specific distribution difference measurement process includes: The Euclidean distance between the weighted fusion feature map and the normal sample feature map in the preset feature library is calculated in sequence to obtain a first feature distribution difference result and compare it with the dynamic threshold to obtain a first detection result; if the first detection result is an abnormal image, the output detection result is an abnormal image; If the first detection result is a normal image, the chi-square statistic between the weighted fusion feature map and the normal sample feature map in the preset feature library is calculated to obtain the second feature distribution difference result and compare it with the dynamic threshold to obtain the second detection result; if the second detection result is an abnormal image, the detection result is output as an abnormal image; If the second detection result is a normal image, the divergence between the weighted fusion feature map and the normal sample feature map in the preset feature library is calculated to obtain the third feature distribution difference result and compare it with the dynamic threshold to obtain the third detection result, which is an abnormal image or a normal image.
8. An image anomaly detection system, characterized in that: The system includes: An input module, used for collecting images to be detected; A multi-scale transformation module is used to perform dynamic scale sampling on the image to be detected to obtain a multi-scale image; The feature extraction module is used to extract features from multi-scale images and obtain preliminary multi-scale feature maps; A multi-branch weight calculation module, in which a pre-trained weight prediction network WPN is set to calculate the weight of each branch; The feature fusion module is used to fuse features based on the multi-scale preliminary feature map and the weight of each branch to obtain a weighted fusion feature map; A feature library, used to store feature maps of normal image samples or abnormal image samples; The anomaly detection module is used to judge the weighted fusion feature map and the normal image sample feature map in the feature library to obtain the anomaly detection result; The output module is used to output anomaly detection results.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the image anomaly detection method based on resolution scaling as claimed in any one of claims 1 to 7 is implemented.
10. A computer readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image anomaly detection method based on multi-scale resolution scaling as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Anomaly detection method and system based on RGB-D fusion and dual-time flow feature learning
CN121074851A
Abnormality detection method and system based on RGB-D fusion and double time flow feature learning
CN121074851B
Intelligent equipment state monitoring method and device, electronic equipment and storage medium
CN121191094A
An intelligent monitoring method and device for device state, an electronic device, and a storage medium
CN121191094B