Cross-domain small-sample anomaly detection method, device, electronic device and storage medium

By dividing abnormal detection into three categories: low-level, intermediate-level, and high-level semantic anomalies, and based on the comprehensive comparison of image sub-blocks and text features, the problem of difficulty in expanding across fields in the existing technology is solved, cross-domain small sample abnormality detection is realized, and detection performance and efficiency are improved.

CN119887762BActive Publication Date: 2025-07-08INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510363946.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing anomaly detection technology is highly specialized and customized, and it is difficult to expand and generalize across fields. The training data demands are large, which limits the standardization research of anomaly detection technology.

Method used

By dividing exception detection into three categories: low-level semantic exceptions, intermediate-level semantic exceptions, and high-level semantic exceptions, and based on the comprehensive comparison of image sub-block features, sub-component features and text features, cross-domain small sample anomaly detection is realized to reduce dependence on training data.

Benefits of technology

Cross-domain anomaly detection methods are implemented, detection performance is improved, knowledge gaps in different fields are all alleviated, and different anomaly types can be detected in a small number of samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887762B_ABST
    Figure CN119887762B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of anomaly detection, and provides a cross-domain few-shot anomaly detection method, device, electronic device and storage medium. The method includes: extracting the text features of the image content knowledge corresponding to the normal images; determining the test image reconstruction result of the test image based on the matching difference between the image sub-block features of the test image and the image sub-block features of the normal images, and determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result; determining the intermediate-level semantic anomaly detection result in the test image based on the first sub-component features of the test image and the second sub-component features of the normal images; determining the high-level semantic anomaly detection result based on the text features and the overall image features of the test image, and determining the target anomaly detection result based on the low-level, intermediate-level and high-level semantic anomaly detection results. This method does not require a large number of samples for training, thereby reducing the cost of anomaly detection training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly detection, and particularly to a cross-domain small-sample anomaly detection method, device, electronic device, and storage medium. Background Art

[0002] Anomaly detection aims to use computer vision technology to detect whether visible anomalies exist in a test sample by taking pictures or scanning images. It is a practical task involving multiple fields, including defect detection in industrial production, logical anomaly detection in product assembly processes, lesion detection in the medical field, and abnormal vehicle detection in traffic scenarios. Existing anomaly detection technologies usually design dedicated detection frameworks and methods for each field. For example, the method of image sub-block feature matching is commonly used in industrial defect detection, the method based on image component segmentation is commonly used in logical anomaly detection, and the method based on image reconstruction is commonly used in medical and traffic anomaly detection.

[0003] However, the existing anomaly detection technologies mainly have the following deficiencies: (1) Existing anomaly detection technologies need to design dedicated detection methods for different data characteristics and anomaly types in each field, and the anomaly detection methods designed for one field are difficult to apply to the data of another field. (2) Within the same field, existing anomaly detection methods usually need to train dedicated models for each type of object to be detected, and the trained models can only be used for anomaly detection of this type of item, with poor generalization performance. (3) Existing anomaly detection technologies need to use a large number of samples for training, with a large demand for training data volume and a cumbersome training process.

[0004] In summary, the existing anomaly detection technologies are highly specialized and customized, making it difficult to expand and generalize, which limits the standardized research of anomaly detection technologies. Summary of the Invention

[0005] The present invention provides a cross-domain small-sample anomaly detection method, device, electronic device, and storage medium to solve the defect that the existing anomaly detection technologies in the prior art are highly specialized and customized, making it difficult to expand and generalize, which limits the standardized research of anomaly detection technologies.

[0006] The present invention provides a cross-domain small-sample anomaly detection method, including the following steps:

[0007] Respectively obtain the image content knowledge of normal images and test images, and extract the text features of the image content knowledge corresponding to the normal images; based on the matching difference between the first image sub-block features of the test images and the second image sub-block features of the normal images, determine the test image reconstruction result of the test images, and based on the reconstruction difference corresponding to the test image reconstruction result, determine the low-level semantic anomaly detection result in the test images;

[0008] Determine the intermediate semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image;

[0009] Determine the high-level semantic anomaly detection result in the test image based on the text feature and the overall image feature of the test image, and determine the target anomaly detection result of the test image based on the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result, and the high-level semantic anomaly detection result.

[0010] According to a cross-domain few-shot anomaly detection method provided by the present invention, determining the test image reconstruction result based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image includes:

[0011] Sort the matching differences between the first image sub-block feature of the test image and the second image sub-block feature of the normal image from smallest to largest to obtain the target matching difference with the smallest matching difference;

[0012] Determine the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature;

[0013] Replace the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result.

[0014] According to a cross-domain few-shot anomaly detection method provided by the present invention, determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result includes:

[0015] Determine the reconstruction difference between each image sub-block feature in the test image reconstruction result and each image sub-block feature in the test image;

[0016] When the reconstruction difference is greater than the first threshold, determine that the image sub-block feature in the test image is a low-level semantic anomaly.

[0017] According to a cross-domain few-shot anomaly detection method provided by the present invention, the first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image;

[0018] The second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image.

[0019] According to a cross - domain small - sample anomaly detection method provided by the present invention, the image content knowledge includes all objects in the image, detection frames of all the objects, and segmentation masks of the objects in the detection frames;

[0020] The determining step of the first image sub - block feature includes:

[0021] Multiply the segmentation mask in the image content knowledge of the test image pixel - by - pixel with the test image to obtain a first target sub - block;

[0022] Use a feature extractor to extract features from the first target sub - block to obtain the first image sub - block feature;

[0023] The determining step of the second image sub - block feature includes:

[0024] Multiply the segmentation mask in the image content knowledge of the normal image pixel - by - pixel with the normal image to obtain a second target sub - block;

[0025] Use the feature extractor to extract features from the second target sub - block to obtain the second image sub - block feature.

[0026] According to a cross - domain small - sample anomaly detection method provided by the present invention, determining the intermediate - level semantic anomaly detection result in the test image based on the first sub - component feature of the test image and the second sub - component feature of the normal image includes:

[0027] Determine the sub - component anomaly score between the first sub - component feature and the second sub - component feature;

[0028] When the sub - component anomaly score is greater than a second threshold, determine that the sub - component feature in the test image is an intermediate - level semantic anomaly.

[0029] According to a cross - domain small - sample anomaly detection method provided by the present invention, determining the high - level semantic anomaly detection result in the test image based on the text feature and the overall image feature of the test image includes:

[0030] Determine the overall anomaly score between the text feature of the normal image and the overall image feature of the test image;

[0031] When the overall anomaly score is greater than a third threshold, determine that the overall image feature of the test image is a high - level semantic anomaly.

[0032] The present invention also provides a cross - domain small - sample anomaly detection device, including the following units:

[0033] An acquisition unit, configured to acquire the image content knowledge of a normal image and a test image respectively, and extract the text features of the image content knowledge corresponding to the normal image;

[0034] A first determination unit, configured to determine the test image reconstruction result of the test image based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, and determine the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result;

[0035] A second determination unit, configured to determine the intermediate semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image;

[0036] A third determination unit, configured to determine the high-level semantic anomaly detection result in the test image based on the text features and the overall image feature of the test image, and determine the target anomaly detection result of the test image based on the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result, and the high-level semantic anomaly detection result.

[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the cross-domain small-sample anomaly detection method as described in any one of the above is implemented.

[0038] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the cross-domain small-sample anomaly detection method as described in any one of the above is implemented.

[0039] The present invention also provides a computer program product, including a computer program, where when the computer program is executed by a processor, the cross-domain small-sample anomaly detection method as described in any one of the above is implemented.

[0040] The cross - domain small - sample anomaly detection method, device, electronic device and storage medium provided by the present invention: (1) classifies different types of anomalies in different domains into three categories: low - level semantic anomalies, medium - level semantic anomalies, and high - level semantic anomalies, and detects these three categories of anomalies, thus realizing a cross - domain anomaly detection method; (2) conducts a detailed content analysis of normal images and test images, which can reduce the knowledge gap between different domains, provide image content knowledge for subsequent anomaly detection, and improve the cross - domain anomaly detection performance; (3) comprehensively uses detection methods based on feature comparison and reconstruction. Only a small number of normal samples need to be provided during the test to achieve anomaly detection of test images, without the need for training on proprietary data, solving the problem of large training requirements of existing methods; (4) by comprehensively comparing features at different semantic levels, as well as text features and overall image features, this method can detect different types of anomalies that may exist in test images, realizing cross - domain small - sample anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is one of the flow diagrams of the cross - domain small - sample anomaly detection method provided by the present invention.

[0043] Figure 2 is the flow diagram for determining the detection result of low - level semantic anomalies in the test image provided by the present invention.

[0044] Figure 3 is the second flow diagram of the cross - domain small - sample anomaly detection method provided by the present invention.

[0045] Figure 4 is the third flow diagram of the cross - domain small - sample anomaly detection method provided by the present invention.

[0046] Figure 5 is the structural diagram of the cross - domain small - sample anomaly detection device provided by the present invention.

[0047] Figure 6 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0049] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same category.

[0050] Figure 1 is one of the flow schematic diagrams of the cross-domain small-sample anomaly detection method provided by the present invention. As Figure 1 shown, the method includes step 110, step 120, step 130, and step 140.

[0051] Step 110, respectively obtain the image content knowledge of the normal image and the test image, and extract the text features of the image content knowledge corresponding to the normal image.

[0052] Specifically, the image content knowledge of the normal image and the test image can be obtained. For example, a vision-based model can be used to perform content recognition, object detection, and semantic segmentation on the image content in the normal image and the test image respectively to obtain the image content knowledge in the normal image and the image content knowledge in the test image.

[0053] Here, the normal image refers to an expected or standard image, usually used for training the model to help the model learn the features under normal conditions. The test image is an image used to verify the performance of the model, usually not participating in the training, and may contain normal or abnormal situations.

[0054] Among them, the image content knowledge includes all the objects in the image, the detection frames of all the objects, and the segmentation masks of the objects in the detection frames. Here, all the objects are, for example, wooden boards, nails, fruits, etc., and the embodiments of the present invention do not make specific limitations in this regard.

[0055] Then, extract the text features of the image content knowledge corresponding to the normal image. Specifically, a text description corresponding to the image content knowledge of the normal image can be written, and a text encoder can be used to extract the text features of the text description corresponding to the normal image.

[0056] Among them, the text encoder can be a BERT (Bidirectional Encoder Representations from Transformers) model, an RNN (Recurrent Neural Network), a CNN (Convolutional Neural Network), etc. The embodiments of the present invention do not make specific limitations thereon.

[0057] Step 120: Determine the test image reconstruction result of the test image based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, and determine the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result.

[0058] Specifically, Figure 2 is a schematic flow chart of determining the low-level semantic anomaly detection result in the test image provided by the present invention. As Figure 2 shown, the test image reconstruction result of the test image can be determined based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image. For example, bipartite graph matching can be performed on the first image sub-block feature of the test image and the second image sub-block feature of the normal image.

[0059] Here, an image encoder can be used to extract features from the normal image to obtain the second image sub-block feature, and an image encoder can be used to extract features from the test image to obtain the first image sub-block feature.

[0060] Among them, the cosine distance or the Manhattan distance can be used to calculate the matching difference. The embodiments of the present invention do not make specific limitations thereon.

[0061] Then, determine the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result. Among them, the reconstruction difference corresponding to the test image reconstruction result can be determined based on the difference between the test image and the test image reconstruction result.

[0062] Among them, the difference between the test image and the test image reconstruction result can be determined using the cosine distance or the Manhattan distance. The embodiments of the present invention do not make specific limitations thereon.

[0063] Here, the low-level semantic anomaly detection result is used to reflect the low-level semantic anomaly situation of the test image. The low-level semantic anomaly refers to extremely local anomalies occurring in a certain image sub-block, such as scratches and damages on a wooden board. The embodiments of the present invention do not make specific limitations thereon.

[0064] Step 130: Determine the intermediate semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image.

[0065] Specifically, the intermediate semantic anomaly detection result in the test image can be determined based on the first sub-component feature of the test image and the second sub-component feature of the normal image.

[0066] Here, the intermediate semantic anomaly detection result is used to reflect the intermediate semantic anomaly situation of the test image. Intermediate semantic anomaly refers to an anomaly occurring at the level of a certain sub-component in the target. For example, if a mechanical component consists of two screws and two nuts, and it is detected that a certain component has only one screw or one nut, then the intermediate semantic anomaly detection result is an intermediate semantic anomaly.

[0067] Step 140: Determine the high-level semantic anomaly detection result in the test image based on the text feature and the overall image feature of the test image, and determine the target anomaly detection result of the test image based on the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result, and the high-level semantic anomaly detection result.

[0068] Specifically, the high-level semantic anomaly detection result in the test image can be determined based on the text feature and the overall image feature of the test image. Here, the high-level semantic anomaly detection result is used to reflect the high-level semantic anomaly situation of the test image. High-level semantic anomaly refers to an overall anomaly of the test image. For example, during the detection of a wooden board, an iron plate appears.

[0069] The high-level semantic anomaly detection result can be determined based on the overall anomaly score between the text feature and the overall image feature of the test image. The calculation of the overall anomaly score can use cosine distance, or Manhattan distance, etc. The embodiments of the present invention do not make specific limitations on this.

[0070] After obtaining the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result, and the high-level semantic anomaly detection result, the target anomaly detection score can be determined based on the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result, and the high-level semantic anomaly detection result, and the target anomaly detection result of the test image can be determined based on the target anomaly detection score.

[0071] Here, the formula for the target anomaly detection score is as follows:

[0072]

[0073] Among them, represents the target anomaly detection score, represents the low-level semantic anomaly detection result, represents the intermediate semantic anomaly detection result, Indicates the advanced semantic anomaly detection result, which is an adjustable weighted hyperparameter.

[0074] It should be noted that when the target anomaly detection score is greater than the preset threshold, the target anomaly detection result is determined to be that the test image is abnormal.

[0075] The method provided by the embodiments of the present invention: (1) classifies different types of anomalies in different fields into three categories: low-level semantic anomalies, intermediate-level semantic anomalies, and advanced semantic anomalies, and detects these three categories of anomalies, realizing a cross-domain anomaly detection method; (2) conducts a detailed content analysis of normal images and test images, which can reduce the knowledge gap between different fields, provide image content knowledge for subsequent anomaly detection, and improve the cross-domain anomaly detection performance; (3) comprehensively uses detection methods based on feature comparison and reconstruction, and only a small number of normal samples need to be provided during the test to achieve anomaly detection of test images, without the need to train on proprietary data, solving the problem of large training requirements of existing methods; (4) by comprehensively comparing features at different semantic levels, as well as text features and overall image features, this method can detect different types of anomalies that may exist in test images, realizing cross-domain small-sample anomaly detection.

[0076] Based on the above embodiments, determining the test image reconstruction result of the test image based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image in step 120 includes:

[0077] Step 121, sorting the matching differences between the first image sub-block feature of the test image and the second image sub-block feature of the normal image from smallest to largest to obtain the target matching difference with the smallest matching difference;

[0078] Step 122, determining the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature;

[0079] Step 123, replacing the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result.

[0080] Specifically, the matching differences between the first image sub-block feature of the test image and the second image sub-block feature of the normal image are sorted from smallest to largest to obtain the target matching difference with the smallest matching difference.

[0081] Here, the matching difference can be calculated using the cosine distance, and the formula for the matching difference is as follows:

[0082]

[0083] Among them, represents the matching difference, represents the first image sub-block feature, represents the second image sub-block feature.

[0084] Then, determine the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature.

[0085] Finally, replace the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result, thereby realizing the reconstruction of the test image. The reconstruction formula is as follows:

[0086]

[0087] Among them, represents the test image reconstruction result, represents the second target image sub-block feature, represents the first target image sub-block feature.

[0088] Based on the above embodiments, determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result includes:

[0089] Step 210, determine the reconstruction difference between each image sub-block feature in the test image reconstruction result and each image sub-block feature in the test image;

[0090] Step 220, when the reconstruction difference is greater than the first threshold, determine that the image sub-block feature in the test image is a low-level semantic anomaly.

[0091] Specifically, determine the reconstruction difference between each image sub-block feature in the test image reconstruction result and each image sub-block feature in the test image, and judge whether each sub-block of the test image is abnormal according to the size of the reconstruction difference.

[0092] Among them, the formula for the reconstruction difference is as follows:

[0093]

[0094] Among them, represents the reconstruction difference, represents the test image reconstruction result, represents the test image.

[0095] In the case where the reconstruction difference is greater than the first threshold, it is determined that the image sub-block feature in the test image is a low-level semantic anomaly.

[0096] Based on the above embodiments, the first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image;

[0097] The second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image.

[0098] Specifically, the first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image, and the formula is as follows:

[0099] ,

[0100] Wherein, represents the first sub-component feature, is the number of image sub-blocks included in each sub-component i , represents the first image sub-block feature of the test image, represents the image content knowledge of the test image.

[0101] Here, the second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image, and the formula is as follows:

[0102]

[0103] Wherein, represents the second sub-component feature, is the number of image sub-blocks included in each sub-component i , represents the second image sub-block feature of the normal image, represents the image content knowledge of the test image.

[0104] Based on the above embodiments, the image content knowledge includes all objects in the image, detection frames of all the objects, and segmentation masks of the objects in the detection frames;

[0105] The determination steps of the first image sub-block feature include:

[0106] Multiply the segmentation mask in the image content knowledge of the test image pixel by pixel with the test image to obtain a first target sub-block;

[0107] Use a feature extractor to extract features from the first target sub-block to obtain the first image sub-block feature;

[0108] The steps for determining the second image sub-block features include:

[0109] Multiply the segmentation mask in the image content knowledge of the normal image pixel by pixel with the normal image to obtain a second target sub-block;

[0110] Use the feature extractor to extract features from the second target sub-block to obtain the second image sub-block features.

[0111] Specifically, the image content knowledge includes all objects in the image, the detection bounding boxes of all objects, and the segmentation masks of the objects in the detection bounding boxes.

[0112] Specifically, a basic visual object recognition model can be used to recognize all objects included in the image, and then a basic visual object detection model is used to detect the recognized objects to obtain the detection bounding boxes of all objects. Finally, a basic visual semantic segmentation model is used to segment all objects in all detection bounding boxes to obtain the segmentation mask of each object.

[0113] Among them, the basic visual object recognition model can be the Recognize Anything model, the basic visual object detection model can be the Grounding DINO model, the basic visual semantic segmentation model can be the Segment Anything model, etc. The embodiments of the present invention do not make specific limitations in this regard.

[0114] Here, the steps for determining the first image sub-block features include:

[0115] Multiply the segmentation mask in the image content knowledge of the test image pixel by pixel with the test image to obtain a first target sub-block. Here, the segmentation mask represents the target area in the form of a binary image, where the area with a mask value of 1 represents the target object, and the area with a mask value of 0 represents the background.

[0116] Then use the feature extractor to extract features from the first target sub-block to obtain the first image sub-block features. The feature extractor can be a CNN model or other feature extractors, etc.

[0117] Here, the steps for determining the second image sub-block features include:

[0118] Multiply the segmentation mask in the image content knowledge of the normal image pixel by pixel with the normal image to obtain a second target sub-block, and then use the feature extractor to extract features from the second target sub-block to obtain the second image sub-block features. The first image sub-block features and the second image sub-block features refer to the features extracted from the local area (i.e., sub-block) of the image and are used to describe the visual information of this area.

[0119] Based on the above embodiments, step 130 includes:

[0120] Step 131, determine the sub-component anomaly score between the first sub-component feature and the second sub-component feature;

[0121] Step 132, when the sub-component anomaly score is greater than the second threshold, determine that the sub-component feature in the test image is a medium-level semantic anomaly.

[0122] Specifically, to determine the sub-component anomaly score between the first sub-component feature and the second sub-component feature, for example, the nearest neighbor comparison can be performed on the first sub-component feature and the second sub-component feature to determine whether there is an anomaly in each sub-component of the test image. The formula is as follows:

[0123]

[0124] where, represents the sub-component anomaly score, represents the first sub-component feature, represents the second sub-component feature, represents the cosine distance between the first sub-component feature and the second sub-component feature.

[0125] When the sub-component anomaly score is greater than the second threshold, determine that the sub-component feature in the test image is a medium-level semantic anomaly.

[0126] Based on the above embodiments, in step 140, determining the high-level semantic anomaly detection result in the test image based on the text feature and the overall image feature of the test image includes:

[0127] Step 141, determine the overall anomaly score between the text feature of the normal image and the overall image feature of the test image;

[0128] Step 142, when the overall anomaly score is greater than the third threshold, determine that the overall image feature of the test image is a high-level semantic anomaly.

[0129] Specifically, the overall anomaly score between the text feature of the normal image and the overall image feature of the test image can be determined to judge whether the overall content in the test image has changed, and accordingly, the high-level semantic anomaly at the overall level of the test image is detected. The formula is as follows:

[0130]

[0131] where, represents the overall anomaly score, represents the overall image feature, represents the text feature of the normal image, represents the cosine distance between the overall image feature and the text feature.

[0132] Based on any of the above embodiments, Figure 3 is the second schematic flowchart of the cross-domain small-sample anomaly detection method provided by the present invention, Figure 4 is the third schematic flowchart of the cross-domain small-sample anomaly detection method provided by the present invention. As Figure 3 , Figure 4 shown, the method includes:

[0133] In the first step, use a visual base model to perform content recognition, object detection, and semantic segmentation on normal images and test images respectively to obtain image content knowledge, and use a text encoder to extract the text features of the image content knowledge corresponding to the normal images.

[0134] In the second step, use an image encoder to extract the feature vectors of the test images and normal images, including the feature vectors of the overall image and the feature vectors of each image sub-block.

[0135] In the third step, perform bipartite graph matching on the first image sub-block feature of the test image and the second image sub-block feature of the normal image, sort the matching differences from smallest to largest to obtain the target matching difference with the smallest matching difference, then determine the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature. Finally, replace the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result.

[0136] In the fourth step, determine the reconstruction differences between the image sub-block features in the test image reconstruction result and the image sub-block features in the test image. When the reconstruction difference is greater than the first threshold, determine that the image sub-block feature in the test image is a low-level semantic anomaly.

[0137] In the fifth step, perform nearest neighbor comparison on the first sub-component feature of the test image and the second sub-component feature of the normal image to determine the intermediate semantic anomaly detection result of the test image. The first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image, and the second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image.

[0138] In the sixth step, determine the high-level semantic anomaly detection result of the test image based on the text features and the overall image feature of the test image, and determine the target anomaly detection result based on the low-level semantic anomaly detection result, intermediate semantic anomaly detection result, and high-level semantic anomaly detection result.

[0139] The cross - domain small - sample anomaly detection device provided by the present invention will be described below. The cross - domain small - sample anomaly detection device described below can be mutually corresponding and referred to the cross - domain small - sample anomaly detection method described above.

[0140] Based on any of the above - mentioned embodiments, the present invention provides a cross - domain small - sample anomaly detection device. Figure 5 It is a schematic structural diagram of the cross - domain small - sample anomaly detection device provided by the present invention. As Figure 5 shown, the device includes:

[0141] An acquisition unit 510, configured to respectively acquire the image content knowledge of the normal image and the test image, and extract the text features of the image content knowledge corresponding to the normal image.

[0142] A first determination unit 520, configured to determine the test - image reconstruction result of the test image based on the matching difference between the first image sub - block feature of the test image and the second image sub - block feature of the normal image, and determine the low - level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test - image reconstruction result.

[0143] A second determination unit 530, configured to determine the intermediate - level semantic anomaly detection result in the test image based on the first sub - component feature of the test image and the second sub - component feature of the normal image.

[0144] A third determination unit 540, configured to determine the high - level semantic anomaly detection result in the test image based on the text features and the overall image feature of the test image, and determine the target anomaly detection result of the test image based on the low - level semantic anomaly detection result, the intermediate - level semantic anomaly detection result, and the high - level semantic anomaly detection result.

[0145] The device provided by the embodiment of the present invention: (1) classifies different types of anomalies in different domains into three categories: low - level semantic anomalies, intermediate - level semantic anomalies, and high - level semantic anomalies, and detects these three categories of anomalies, realizing a cross - domain anomaly detection method; (2) conducts a detailed content analysis of the normal image and the test image, which can reduce the knowledge gap between different domains, provide image content knowledge for subsequent anomaly detection, and improve the cross - domain anomaly detection performance; (3) comprehensively uses a detection method based on feature comparison and a detection method based on reconstruction, and only needs to provide a small number of normal samples during the test to realize the anomaly detection of the test image, without the need for training on proprietary data, solving the problem of large training requirements of the existing methods; (4) by comprehensively comparing features at different semantic levels, as well as text features and the overall image feature, this method can detect different types of anomalies that may exist in the test image, realizing cross - domain small - sample anomaly detection.

[0146] Based on any of the above embodiments, the first determining unit 520 is specifically configured to:

[0147] Sort the matching differences between the first image sub-block features of the test image and the second image sub-block features of the normal image from smallest to largest to obtain the target matching difference with the smallest matching difference;

[0148] Determine the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature;

[0149] Replace the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result.

[0150] Based on any of the above embodiments, the first determining unit 520 is specifically configured to:

[0151] Determine the reconstruction difference between each image sub-block feature in the test image reconstruction result and each image sub-block feature in the test image;

[0152] When the reconstruction difference is greater than the first threshold, determine that the image sub-block feature in the test image is a low-level semantic anomaly.

[0153] Based on any of the above embodiments, the first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image;

[0154] The second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image.

[0155] Based on any of the above embodiments, the image content knowledge includes all targets in the image, detection frames of all the targets, and segmentation masks of the targets in the detection frames;

[0156] It further includes a first image sub-block feature determining unit, and the first image sub-block feature determining unit is specifically configured to:

[0157] Multiply the segmentation mask in the image content knowledge of the test image pixel by pixel with the test image to obtain a first target sub-block;

[0158] Use a feature extractor to extract features from the first target sub-block to obtain the first image sub-block feature;

[0159] It further includes a second image sub-block feature determining unit, and the second image sub-block feature determining unit is specifically configured to:

[0160] Multiply the segmentation mask in the image content knowledge of the normal image pixel by pixel with the normal image to obtain a second target sub-block;

[0161] Use the feature extractor to perform feature extraction on the second target sub-block to obtain the second image sub-block feature.

[0162] Based on any of the above embodiments, the second determination unit 530 is specifically configured to:

[0163] Determine the sub-component anomaly score between the first sub-component feature and the second sub-component feature;

[0164] In the case where the sub-component anomaly score is greater than a second threshold, determine that the sub-component feature in the test image is a medium-level semantic anomaly.

[0165] Based on any of the above embodiments, the third determination unit 540 is specifically configured to:

[0166] Determine the overall anomaly score between the text feature of the normal image and the overall image feature of the test image;

[0167] In the case where the overall anomaly score is greater than a third threshold, determine that the overall image feature of the test image is a high-level semantic anomaly.

[0168] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention, as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the cross-domain small-sample anomaly detection method, which includes: respectively obtaining the image content knowledge of the normal image and the test image, and extracting the text features of the image content knowledge corresponding to the normal image; based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, determining the test image reconstruction result of the test image, and based on the reconstruction difference corresponding to the test image reconstruction result, determining the low-level semantic anomaly detection result in the test image; based on the first sub-component feature of the test image and the second sub-component feature of the normal image, determining the intermediate-level semantic anomaly detection result in the test image; based on the text features and the overall image feature of the test image, determining the high-level semantic anomaly detection result in the test image, and based on the low-level semantic anomaly detection result, the intermediate-level semantic anomaly detection result, and the high-level semantic anomaly detection result, determining the target anomaly detection result of the test image.

[0169] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0170] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the cross-domain small-sample anomaly detection method provided by each of the above methods. The method includes: respectively obtaining the image content knowledge of a normal image and a test image, and extracting the text features of the image content knowledge corresponding to the normal image; determining the test image reconstruction result of the test image based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, and determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result; determining the intermediate-level semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image; determining the high-level semantic anomaly detection result in the test image based on the text features and the overall image feature of the test image, and determining the target anomaly detection result of the test image based on the low-level semantic anomaly detection result, the intermediate-level semantic anomaly detection result, and the high-level semantic anomaly detection result.

[0171] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the cross-domain small-sample anomaly detection method provided by each of the above methods. The method includes: respectively obtaining the image content knowledge of a normal image and a test image, and extracting the text features of the image content knowledge corresponding to the normal image; determining the test image reconstruction result of the test image based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, and determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result; determining the intermediate-level semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image; determining the high-level semantic anomaly detection result in the test image based on the text features and the overall image feature of the test image, and determining the target anomaly detection result of the test image based on the low-level semantic anomaly detection result, the intermediate-level semantic anomaly detection result, and the high-level semantic anomaly detection result.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross - domain small - sample anomaly detection method, characterized in that, Including: Respectively obtain the image content knowledge of the normal image and the test image, and extract the text features of the image content knowledge corresponding to the normal image; The image content knowledge includes all the targets in the image, the detection frames of all the targets, and the segmentation masks of the targets in the detection frames; Based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, determine the test image reconstruction result of the test image, and based on the reconstruction difference corresponding to the test image reconstruction result, determine the low-level semantic anomaly detection result in the test image; Based on the first sub-component feature of the test image and the second sub-component feature of the normal image, determine the intermediate semantic anomaly detection result in the test image; the first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image; the second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image; Based on the text feature and the overall image feature of the test image, determine the high-level semantic anomaly detection result in the test image, and based on the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result and the high-level semantic anomaly detection result, determine the target anomaly detection result of the test image; The determining the test image reconstruction result of the test image based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image includes: Sort the matching differences between the first image sub-block feature of the test image and the second image sub-block feature of the normal image from smallest to largest to obtain the target matching difference with the smallest matching difference; Determine the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature; Replace the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result.

2. The cross-domain small-sample anomaly detection method according to claim 1, wherein The determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result includes: Determine the reconstruction difference between each image sub-block feature in the test image reconstruction result and each image sub-block feature in the test image; When the reconstruction difference is greater than the first threshold, determine that the image sub-block feature in the test image is a low-level semantic anomaly.

3. The cross-domain small-sample anomaly detection method according to any one of claims 1 to 2, characterized in that The image content knowledge includes all the targets in the image, the detection frames of all the targets, and the segmentation masks of the targets in the detection frames; The determining step of the first image sub-block feature includes: Multiply the segmentation mask in the image content knowledge of the test image pixel by pixel with the test image to obtain a first target sub-block; Use a feature extractor to extract features from the first target sub-block to obtain the first image sub-block feature; The determining step of the second image sub-block feature includes: Multiply the segmentation mask in the image content knowledge of the normal image pixel by pixel with the normal image to obtain a second target sub-block; Use the feature extractor to extract features from the second target sub-block to obtain the second image sub-block feature.

4. The cross-domain small-sample anomaly detection method according to any one of claims 1 to 2, characterized in that The determining the intermediate semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image includes: Determine the sub-component anomaly score between the first sub-component feature and the second sub-component feature; When the sub-component anomaly score is greater than a second threshold, determine that the sub-component feature in the test image is an intermediate semantic anomaly.

5. The cross-domain small-sample anomaly detection method according to any one of claims 1 to 2, characterized in that The determining the high-level semantic anomaly detection result in the test image based on the text feature and the overall image feature of the test image includes: Determine the overall anomaly score between the text feature of the normal image and the overall image feature of the test image; When the overall anomaly score is greater than a third threshold, determine that the overall image feature of the test image is a high-level semantic anomaly.

6. A cross - domain small - sample anomaly detection device, characterized in that, Includes: An acquisition unit for respectively acquiring the image content knowledge of a normal image and a test image, and extracting the text feature of the image content knowledge corresponding to the normal image; The image content knowledge includes all targets in the image, the detection frames of all the targets, and the segmentation masks of the targets in the detection frames; A first determination unit for determining the test image reconstruction result based on the matching difference between the first image sub-block feature of the test image and the second image sub-block feature of the normal image, and determining the low-level semantic anomaly detection result in the test image based on the reconstruction difference corresponding to the test image reconstruction result; A second determination unit for determining the intermediate semantic anomaly detection result in the test image based on the first sub-component feature of the test image and the second sub-component feature of the normal image; the first sub-component feature is determined based on the image content knowledge of the test image and the first image sub-block feature of the test image; the second sub-component feature is determined based on the image content knowledge of the normal image and the second image sub-block feature of the normal image; A third determination unit for determining the high-level semantic anomaly detection result in the test image based on the text feature and the overall image feature of the test image, and determining the target anomaly detection result of the test image based on the low-level semantic anomaly detection result, the intermediate semantic anomaly detection result, and the high-level semantic anomaly detection result; The first determination unit is specifically configured to: Sort the matching differences between the first image sub-block feature of the test image and the second image sub-block feature of the normal image from smallest to largest to obtain the target matching difference with the smallest matching difference; Determine the first target image sub-block feature in the normal image corresponding to the target matching difference, and the second target image sub-block feature in the test image corresponding to the first target image sub-block feature; Replace the second target image sub-block feature in the test image with the first target image sub-block feature to obtain the test image reconstruction result.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the cross-domain small-sample anomaly detection method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the cross-domain small-sample anomaly detection method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device, equipment, storage medium and computer program product

    CN117710301A

  • Zero sample anomaly detection method based on multi-scale feature aggregation and semantic guidance

    CN119180780A