Expandable category guide anomaly detection method and device for multiple categories of targets

By using a lightweight neural network classifier and a multi-model collaborative architecture, combined with multi-scale enhanced preprocessing and dynamic routing mechanisms, the problems of insufficient model generalization and poor scalability in multi-class object detection are solved, achieving high-precision and low-latency anomaly detection.

CN120931579APending Publication Date: 2025-11-11苏州旗开得电子科技有限公司

Patent Information

Application Number
CN202511031015.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as insufficient model generalization, high training and inference costs, lack of dynamic routing strategies, and poor scalability in anomaly detection of multiple target types, making it difficult to balance detection accuracy and efficiency.

Method used

A lightweight neural network classifier is used for category determination. Combined with a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion comparison anomaly model, adaptive anomaly detection is achieved through multi-scale enhanced preprocessing and a dynamic soft routing mechanism.

Benefits of technology

It improves detection accuracy and efficiency, supports flexible model expansion and deployment, and adapts to the multi-target detection needs of different industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931579A_ABST
    Figure CN120931579A_ABST
Patent Text Reader

Abstract

The invention discloses an expandable category guide anomaly detection method and device for multiple categories of targets, and relates to the field of computer vision. The method comprises the following steps: carrying out abnormal region guided adaptive enhancement preprocessing on an industrial image to obtain a target image; judging the category of the target image according to a pre-constructed lightweight neural network classifier; activating at least one anomaly detection model according to the category of the target image; and performing anomaly prediction on the target image by using the activated anomaly detection model, and fusing with the confidence corresponding to the anomaly detection model to realize adaptive anomaly discrimination of the industrial image. According to the method, the sample category is quickly judged through the lightweight classifier, the pre-screening and path guidance of the anomaly detection model are realized, the category guidance weight is generated by using a confidence coefficient mechanism, the subsequent model fusion strategy is endowed with higher adaptability, and the structural clarity, the model selection accuracy and the overall calculation efficiency of the system are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to an scalable category-guided anomaly detection method and apparatus for multiple target types. Background Technology

[0002] In modern industrial inspection, electronics manufacturing, automotive electronics, and automated assembly, target objects often exhibit a wide variety of categories and forms. Anomaly detection, as a key technology for identifying defects, faults, or non-compliance with standards, typically relies on image- or sensor data-based models to locate and identify abnormal areas.

[0003] However, traditional anomaly detection methods mostly use a unified model or general process to process all types of targets, making it difficult to balance detection accuracy and model generalization ability. On the one hand, different types of targets differ significantly in morphology, texture, structure, and defect distribution, making it difficult for general models to accurately adapt and easily leading to missed or false detections. On the other hand, in order to adapt to multi-class detection scenarios, it is necessary to build ultra-large models or perform global training to improve generalization ability, resulting in high training costs and low inference efficiency, making it difficult to meet the needs of actual deployment. Summary of the Invention

[0004] Therefore, it is necessary to provide an scalable category-guided anomaly detection method and device for multiple types of targets to address the above-mentioned technical problems and improve detection accuracy and efficiency.

[0005] Firstly, this application provides an extensible category-guided anomaly detection method for multiple target types. The method includes:

[0006] Acquire industrial images, perform adaptive enhancement preprocessing guided by anomaly regions on the industrial images, and acquire target images;

[0007] The target image is classified based on a pre-built lightweight neural network classifier.

[0008] Based on the category of the target image, at least one anomaly detection model is activated from a pre-constructed subset of structurally heterogeneous anomaly detection models; wherein, the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion contrast anomaly model;

[0009] Anomaly prediction of the target image is performed using the activated anomaly detection model, and the prediction is fused with the confidence scores of the anomaly detection model to achieve adaptive anomaly discrimination of industrial images.

[0010] In one embodiment, performing anomaly region-guided adaptive enhancement preprocessing on an industrial image to obtain a target image includes:

[0011] Industrial images are resampled at multiple scales to obtain multi-scale information, and the multi-scale information is then fused to obtain an enhanced image.

[0012] A salient target detection model is used to identify candidate regions of anomalies in the enhanced image. CLAHE is then applied to these candidate regions for local enhancement to obtain the target image.

[0013] In one embodiment, a lightweight neural network classifier includes an input layer, a feature extraction layer, a fully connected layer, and an output layer;

[0014] Based on a pre-built lightweight neural network classifier, the category of the target image is determined, including:

[0015] The target image is input into a lightweight neural network classifier through an input layer; feature vectors are extracted through a feature extraction layer; the feature vectors are integrated and mapped through a fully connected layer to obtain the probability distribution of each category; and the category of the target image is output through an output layer based on the probability distribution.

[0016] In one embodiment, anomaly prediction of a target image using a structure-aware reconstruction anomaly detector includes:

[0017] The target image is reconstructed using an encoder and a decoder to obtain the reconstructed image;

[0018] Construct a multidimensional reconstruction loss function that includes pixel loss and gradient domain loss;

[0019] Based on the target image and the reconstructed image, anomaly heatmaps or masks are generated by combining structure-aware reconstruction errors.

[0020] In one embodiment, using an improved pseudo-defect classifier to predict anomalies in a target image includes:

[0021] Define a pseudo-transform set including different image enhancement methods, use the elements of the pseudo-transform set to generate pseudo samples based on the target image, and construct a pseudo sample set;

[0022] A parallel multi-branch discriminant structure is constructed, which includes parameter sub-branches corresponding to different elements of the pseudo-transform set; wherein each sub-branch supports hot-plugging.

[0023] The classification results are obtained by using a parallel multi-branch discriminant structure.

[0024] In one embodiment, using a semantic fusion contrast anomaly model to predict anomalies in a target image includes:

[0025] The global vector of the target image is extracted using an encoder;

[0026] Construct natural language category-guided text prompts and obtain the embedding vectors of the prompts;

[0027] The target image is segmented to obtain several local attention maps. A guide map is generated by fusing the local attention maps with the embedding vectors. An anomaly heatmap or a mask is generated based on the semantic difference between the guide map and the embedding vectors.

[0028] The semantic fusion comparison anomaly model also introduces a learnable cue generator corresponding to different anomaly categories, which is concatenated with the guidance text cue to update the guidance text cue.

[0029] Secondly, this application also provides an expandable category-guided anomaly detection device for multiple target types. The device includes:

[0030] The preprocessing module is used to acquire industrial images, perform adaptive enhancement preprocessing guided by abnormal regions on the industrial images, and acquire target images.

[0031] The coarse classification module is used to determine the category of the target image based on a pre-built lightweight neural network classifier;

[0032] The dynamic scheduling and routing module is used to activate at least one anomaly detection model from a pre-built subset of structurally heterogeneous anomaly detection models based on the category of the target image; wherein the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion contrast anomaly model.

[0033] The output module is used to perform anomaly prediction on the target image using the activated anomaly detection model, and to fuse the prediction with the confidence level of the anomaly detection model to achieve adaptive anomaly discrimination of industrial images.

[0034] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the aforementioned scalable category-guided anomaly detection method for multiple target types.

[0035] Fourthly, this application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps in the aforementioned extensible category-guided anomaly detection method for multiple target types.

[0036] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps in the aforementioned extensible category-guided anomaly detection method for multiple target types.

[0037] The aforementioned scalable category-guided anomaly detection method and apparatus for multiple target types includes: acquiring an industrial image; performing adaptive enhancement preprocessing on the industrial image to guide anomaly regions; acquiring a target image; determining the category of the target image based on a pre-built lightweight neural network classifier; activating at least one anomaly detection model from a pre-built subset of structurally heterogeneous anomaly detection models based on the category of the target image; wherein the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion comparison anomaly model; using the activated anomaly detection models to predict anomalies in the target image, and fusing the predictions with the corresponding confidence scores of the anomaly detection models to achieve adaptive anomaly discrimination of industrial images. The aforementioned method and apparatus rapidly determine sample categories using a lightweight classifier, achieving pre-screening and path guidance for anomaly detection models. The use of a confidence mechanism to generate category-guided weights gives subsequent model fusion strategies greater adaptability, effectively improving the system's structural clarity, model selection accuracy, and overall computational efficiency. Attached Figure Description

[0038] Figure 1 This is a flowchart illustrating an scalable category-guided anomaly detection method for multiple target types in one embodiment.

[0039] Figure 2 This is a structural block diagram of an scalable category-guided anomaly detection device for multiple target types in one embodiment. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0041] Some research has attempted to use deep learning models for anomaly detection in industrial images, such as reconstruction models based on self-supervised learning and contrastive methods using representation learning. While these methods perform well on specific datasets, they still face the following technical challenges in complex scenarios involving multiple classes:

[0042] (1) Insufficient model generalization: The general anomaly detection model cannot adapt to the differences in texture, boundary and defect performance between different types of targets, resulting in reduced accuracy;

[0043] (2) High training and inference costs: In order to take into account generalization ability in multi-class environments, large models or multi-branch structures are often required, but there is a lack of efficient scheduling mechanisms.

[0044] (3) Lack of dynamic routing strategy: Existing methods generally adopt static model process, which cannot intelligently select the optimal detection model according to the input category;

[0045] (4) Poor scalability: When adding new categories or defect types, it is often necessary to reconstruct or retrain the model, which increases the complexity of system maintenance and updates.

[0046] Therefore, how to introduce a target category guidance mechanism and build a flexible and efficient multi-model anomaly detection system has become one of the key technical problems that urgently need to be solved.

[0047] To address the aforementioned problems, embodiments of this application provide an extensible category-guided anomaly detection method for multiple target types, such as... Figure 1 As shown, it includes the following steps:

[0048] Step 102: Acquire industrial images, perform adaptive enhancement preprocessing guided by abnormal regions on the industrial images, and acquire target images.

[0049] Industrial images are acquired by using cameras to capture images from scenarios such as modern industrial inspection, electronic manufacturing, automotive electronics, and automated assembly.

[0050] To enhance the system's adaptability to complex backgrounds and multi-morphological targets, a multi-scale enhancement preprocessing step is introduced.

[0051] For the input industrial image I, define the set of multi-scale resampling ratios:

[0052] S = {s1, s2, s3, ..., s} n}s i ∈(0,1]

[0053] The enhanced image is obtained by resampling and fusing the images.

[0054]

[0055] Where n is the number of scales; s i ∈(0,1] represents the i-th scaling factor; Resize(I,s) i ) Scale the image by s i The scaling function; α i Learnable weighting coefficients for scaled images; I multi Multi-scale image fusion results.

[0056] Next, an anomalous region-guided adaptive CLAHE (Contrast Limited Adaptive Histogram Equalization) image enhancement method is adopted.

[0057] The aforementioned adaptive CLAHE image enhancement guided by anomaly regions uses a salient object detection model to first identify suspected anomaly regions, i.e., candidate anomaly regions, and then applies CLAHE only locally to these regions. The specific local processing operation involves first using a shallow convolutional neural network to predict a mask R(x,y)∈[0,1] for possible anomaly regions, then applying local CLAHE enhancement to regions where R(x,y)>0.5, while leaving other regions unchanged. The formula for local CLAHE enhancement is expressed as:

[0058]

[0059] Where R(x,y) is the heatmap of the suspected abnormal area predicted in the preliminary analysis; τ is the enhancement threshold.

[0060] The shallow convolutional neural network described above consists of an input layer, hidden layers, and an output layer. There can be one or more hidden layers. In a shallow neural network, information flows from the input layer to the output layer; each node (neuron) receives input from the previous layer and influences the output of the next layer. Compared to deep convolutional neural networks, its characteristic is that it has relatively fewer network layers.

[0061] Furthermore, to facilitate subsequent image processing, the target image needs to undergo channel normalization to obtain the final target image I', expressed by the formula:

[0062]

[0063] Step 102 integrates learnable weighted multi-scale image construction and performs unified preprocessing before general classification and anomaly detection to enhance edge features and local anomaly responses.

[0064] Step 104: Determine the category of the target image based on the pre-built lightweight neural network classifier.

[0065] A lightweight neural network classifier is used to determine the sub-category of the target image, such as determining the major category (QFN, SOIC, etc.) of the target image of the current AOI device, which is then used for subsequent model distribution.

[0066] In one embodiment, a lightweight neural network classifier includes an input layer, a feature extraction layer, a fully connected layer, and an output layer;

[0067] Based on a pre-built lightweight neural network classifier, the category of the target image is determined, including:

[0068] Step 1041, input image I through the input layer ' Input a lightweight neural network classifier.

[0069] Step 1042: Extract feature vectors through the feature extraction layer. Input the standardized target image I described above. ' Feature vectors are extracted through a feature extraction layer:

[0070] z = f cls (I ' )

[0071] Among them, f cls Classification backbone network; feature vectors extracted from z-images.

[0072] Step 1043: Integrate and map the feature vectors through a fully connected layer to obtain the probability distribution p for each category, expressed as:

[0073] p = Softmax(Wz + b) ∈ R k

[0074] Where W is the classification network weight, b is the bias, and k is the number of classes.

[0075] Step 1044: Output the category of the target image based on the probability distribution through the output layer.

[0076] The main category number can be predicted based on the probability distribution:

[0077] c * =argmax(p)

[0078] Among them, c * The index of the target category with the highest probability is predicted.

[0079] Construct confidence weight vectors for subsequent anomaly fusion

[0080]

[0081] Where, θ k The fusion weight for category k is ∈ a very small positive number to prevent the denominator from being zero. This weighting method highlights the high-confidence category and suppresses the influence of the low-confidence fuzzy category.

[0082] Step 104 uses a squared normalized confidence factor to improve the class sensitivity of subsequent anomaly maps.

[0083] Step 106: Based on the category of the target image, activate at least one anomaly detection model from a pre-constructed subset of structurally heterogeneous anomaly detection models; wherein the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion contrast anomaly model.

[0084] The subset of anomaly detection models is represented as follows:

[0085]

[0086] in, This represents the Kth anomaly detection model.

[0087] The selected anomaly detection model F select Represented as:

[0088]

[0089] in, Indicates the cth * Anomaly detection model for classes.

[0090] In one embodiment, a Top-N soft routing strategy is used to activate the anomaly detection models. This involves scoring the anomaly detection models and selecting the N highest-scoring models for activation. The scoring criterion can be the model's matching degree. The formula for selecting and activating the anomaly detection models is expressed as:

[0091]

[0092] Where τ is the threshold, used to filter categories; C active F is the set of candidate categories that are activated. select This is the selected subset of anomaly detection models.

[0093] In one embodiment, the selection of the anomaly detection model is determined based on the matching degree between the model and the target image anomaly recognition task. Step 104, after determining the category of the target image using a lightweight neural network classifier, queries a matching table to determine the anomaly detection model corresponding to that category, and then activates the corresponding anomaly detection model. The matching relationship between categories and anomaly detection models in the matching table is manually specified. For example, if the target image is identified as belonging to the QFN category, which includes five subcategories: QFN_a, QFN_b, QFN_c, QFN_d, and QFN_e, categories QFN_a and QFN_b in the matching table both correspond to anomaly detection model 1_1, and categories QFN_c, QFN_d, and QFN_e both correspond to anomaly detection model 1_2. The corresponding anomaly detection model can be determined by looking up the table and then activated.

[0094] In other embodiments, to further improve the flexibility and adaptability of anomaly detection model matching, probability distributions for each category are extracted from a lightweight neural network classifier. Categories with probabilities greater than a preset threshold are selected, and anomaly detection models under multiple categories are aggregated. The aggregated anomaly detection models are then evaluated for matching degree with the target image, primarily by calculating the similarity between the training data of the anomaly detection models and the target image. The higher the similarity, the higher the matching degree between the anomaly detection model and the target image. A Top-N algorithm is then used to select N anomaly detection models with high matching degrees to activate the target image.

[0095] Step 106 decouples model calls from categories, supports hot-swappable structures, and enhances inference robustness through a soft routing mechanism.

[0096] Step 108: The activated anomaly detection model is used to predict anomalies in the target image, and the predictions are fused with the confidence scores of the anomaly detection model to achieve adaptive anomaly discrimination for industrial images.

[0097] The anomaly map is fused based on the activated anomaly detection model and its confidence level to generate the final output mask. Standardized score map:

[0098]

[0099] Among them, S k (x,y) is the original anomaly score map output by the k-th model; The value is normalized to [0,1] after being standardized by the Sigmoid function. The outlier maps from different models are then fused using the following formula:

[0100]

[0101] Where, ω k The corresponding confidence weighting coefficient (see step 104); C active Activated model set; S final The final fused anomaly score map. Based on the anomaly score map, a thresholded output mask is created, using the following formula:

[0102]

[0103] Step 108 uses a confidence-driven multi-model fusion mechanism to achieve adaptive anomaly detection for different categories.

[0104] This embodiment introduces a weighted multi-scale fusion strategy, which can highlight multi-level structural features (edges, textures, brightness differences) and enhance the response to small-scale anomalies. Combined with CLAHE and normalization enhancement, it enhances the robustness of images under strong and weak lighting and complex backgrounds, supports subsequent models to capture anomalies in a fine-grained manner, and improves the overall detection accuracy.

[0105] By using a lightweight classifier to quickly determine the sample category, the system can perform pre-screening and path guidance for anomaly detection models. The confidence level squared normalization mechanism is used to generate category guidance weights, which gives the subsequent model fusion strategy a stronger adaptability and effectively improves the system's structural clarity, model selection accuracy and overall computational efficiency.

[0106] The system adopts a Top-N dynamic soft routing strategy, which activates the optimal combination of detection models based on category confidence. It supports "hot-swappable registration" of models, which can expand the target detection capability without modifying the backbone structure, thereby improving the modular deployment capability and maintainability of the system in real industrial scenarios.

[0107] A confidence-squared weighted fusion mechanism is proposed, which combines classification-guided distribution for dynamic anomaly scoring. It supports multi-model output fusion by category perception, effectively avoiding conflicts between different models. The resulting anomaly heatmap and mask have higher spatial resolution, positioning accuracy and interpretability, which is beneficial for subsequent alarms, defect tracing and human-computer interaction.

[0108] In one embodiment, anomaly prediction of a target image using a structure-aware reconstruction anomaly detector includes:

[0109] The target image is reconstructed using an encoder and a decoder to obtain the reconstructed image;

[0110] Construct a multidimensional reconstruction loss function that includes pixel loss and gradient domain loss;

[0111] Based on the target image and the reconstructed image, anomaly heatmaps or masks are generated by combining structure-aware reconstruction errors.

[0112] Specifically, encoder E is used to extract deep semantic embedding representations of an image; decoder D reconstructs the original image from the embedding representations as much as possible.

[0113]

[0114] Where I' is the target image after standardization in step 102; This represents the reconstructed image, which is expected to be consistent with the input on normal samples. H and W are the height and width of the target image, respectively.

[0115] Construct a multidimensional reconstruction loss function, assuming pixel error:

[0116]

[0117] Introducing gradient domain loss:

[0118]

[0119] The total loss is:

[0120] L total =α·L pixel +β·L grad

[0121] Where α and β are the weighting coefficients for each loss.

[0122] The final anomaly heatmap is generated by combining structure-aware reconstruction errors:

[0123]

[0124] The structure-aware reconstruction anomaly detector supports per-pixel anomaly scoring, which can be used for anomaly region heatmap visualization or direct mask generation.

[0125] In one embodiment, using an improved pseudo-defect classifier to predict anomalies in a target image includes:

[0126] Define a pseudo-transform set including different image enhancement methods, use the elements of the pseudo-transform set to generate pseudo samples based on the target image, and construct a pseudo sample set;

[0127] A parallel multi-branch discriminant structure is constructed, which includes parameter sub-branches corresponding to different elements of the pseudo-transform set; wherein each sub-branch supports hot-plugging.

[0128] The classification results are obtained by using a parallel multi-branch discriminant structure.

[0129] Specifically, the improved pseudo-defect classifier adopts the CutPaste++ model, which combines "multi-transformation joint pseudo-defect generation + multi-branch enhancement". The pseudo-transformation set is defined as:

[0130] T = {T1, T2, T3, T4}

[0131] Among them, T i This indicates different enhancement methods.

[0132] Using pseudo-transform sets to generate pseudo-samples and assemble pseudo-sample sets:

[0133]

[0134] Design a parallel multi-branch discriminant structure, where each pseudo-transform corresponds to a set of dedicated parameter sub-branches. i For the i-th subnetwork, the corresponding transformation T i The final classification output will be:

[0135]

[0136] The improved pseudo-defect classifier can avoid the problem of feature interference when different types of pseudo-defects are mixed during training. It supports adding new pseudo-defect branches for specific target categories without reconstructing the backbone network, achieving the effect of structural decoupling and hot-plugging of sub-branches.

[0137] In one embodiment, using a semantic fusion contrast anomaly model to predict anomalies in a target image includes:

[0138] The global vector of the target image is extracted using an encoder;

[0139] Construct natural language category-guided text prompts and obtain the embedding vectors of the prompts;

[0140] The target image is segmented to obtain several local attention maps. A guide map is generated by fusing the local attention maps with the embedding vectors. An anomaly heatmap or a mask is generated based on the semantic difference between the guide map and the embedding vectors.

[0141] The semantic fusion comparison anomaly model also introduces a learnable cue generator corresponding to different anomaly categories, which is concatenated with the guidance text cue to update the guidance text cue.

[0142] Specifically, first, the image encoder extracts the global vector of the target image:

[0143] v i = CLIP.ImageEncoder(I)

[0144] Constructing natural language category-guided text prompts:

[0145] promt = {normal, abnormal}

[0146] Calculate the embedding vectors separately:

[0147] v nor ,v abn = CLIP.TextEncoder(promt)

[0148] Then, the semantic difference is calculated using a constant score function:

[0149] S clip =cos(v I ,v abn )-cos(v I ,v nor )

[0150] Where cos is the cosine similarity function. clip The higher the value, the more the target image leans towards the semantics of "abnormal" rather than "normal".

[0151] To achieve spatial saliency alignment, a guidance map f is generated by fusing the image attention map and the embedding vector of the guidance text cue when calculating the semantic difference. The set of guidance maps is represented as follows:

[0152] F I =|f1,f2,…,f N |

[0153] For each space token f i Calculate the similarity between the vector and the prompt vector:

[0154]

[0155] Obtain the anomaly attention map:

[0156]

[0157] Finally, for different target image categories c, a learnable cue generator is introduced:

[0158] p c =MLP(Embed(c))

[0159] Concatenate with prompt to generate a new text vector:

[0160]

[0161] It can achieve category-level Prompt fine-tuning to adapt to different target semantic backgrounds.

[0162] The semantic fusion comparison anomaly model introduces an attention-guided semantic region mapping mechanism to construct a spatially relevant anomaly response map, and finally generates an anomaly heatmap and outputs a mask.

[0163] This invention provides an scalable category-guided anomaly detection method and apparatus for multiple target types. This scheme introduces a category-aware mechanism into the detection process, combined with a multi-model collaborative architecture, to automatically guide the target to the corresponding dedicated anomaly detection model based on the coarse classification result of the input target, achieving high-precision, low-latency anomaly identification. The advantages of this invention are:

[0164] Category guidance mechanism: Use a lightweight classification module to identify the prior category of the input target to avoid the weakening of capabilities caused by a uniform model;

[0165] Multi-model branching structure: Customize an independent anomaly detection sub-model for each category of target, significantly improving detection accuracy;

[0166] Highly scalable: Supports dynamic expansion of model branches after adding new categories, without requiring overall retraining;

[0167] Flexible deployment: It can be combined with edge computing or cloud services for distributed model scheduling and management, adapting to different industrial scenarios.

[0168] This invention aims to improve the accuracy, efficiency, and maintainability of anomaly detection in various target environments, providing highly reliable solutions for fields such as intelligent manufacturing and automated quality inspection.

[0169] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0170] Based on the same inventive concept, this application also provides an expandable category-guided anomaly detection device for implementing the above-described expandable category-guided anomaly detection method for multiple target types. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the expandable category-guided anomaly detection device for multiple target types provided below can be found in the limitations of the expandable category-guided anomaly detection method for multiple target types described above, and will not be repeated here.

[0171] In one embodiment, such as Figure 2 As shown, a scalable category-guided anomaly detection device for multiple target types is provided, comprising:

[0172] The preprocessing module is used to acquire industrial images, perform adaptive enhancement preprocessing guided by abnormal regions on the industrial images, and acquire target images.

[0173] The coarse classification module is used to determine the category of the target image based on a pre-built lightweight neural network classifier;

[0174] The dynamic scheduling and routing module is used to activate at least one anomaly detection model from a pre-built subset of structurally heterogeneous anomaly detection models based on the category of the target image; wherein the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion contrast anomaly model.

[0175] The output module is used to predict anomalies in the target image using the activated anomaly detection model, and then fuses the predictions with the corresponding confidence scores of the anomaly detection model to achieve adaptive anomaly detection for industrial images. Each module in the aforementioned scalable category-guided anomaly detection device for multiple target types can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0176] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in all of the above method embodiments.

[0177] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in all of the above method embodiments.

[0178] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in all of the above method embodiments.

[0179] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0180] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0181] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0182] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An extensible category-guided anomaly detection method for multiple target types, characterized in that, The method includes: Acquire an industrial image, perform adaptive enhancement preprocessing guided by anomaly region on the industrial image, and acquire a target image; The category of the target image is determined based on a pre-built lightweight neural network classifier; Based on the category of the target image, at least one anomaly detection model is activated from a pre-constructed subset of structurally heterogeneous anomaly detection models; wherein, the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion contrast anomaly model; The activated anomaly detection model is used to predict anomalies in the target image, and the predictions are fused with the confidence scores of the anomaly detection model to achieve adaptive anomaly discrimination of the industrial image.

2. The method according to claim 1, characterized in that, The adaptive enhancement preprocessing guided by anomaly regions in the industrial image to obtain the target image includes: The industrial image is resampled at multiple scales to obtain multi-scale information, and the multi-scale information is fused to obtain an enhanced image. A salient target detection model is used to determine the abnormal candidate regions in the enhanced image, and CLAHE is used to locally enhance the abnormal candidate regions to obtain the target image.

3. The method according to claim 1, characterized in that, The lightweight neural network classifier includes an input layer, a feature extraction layer, a fully connected layer, and an output layer; The step of determining the category of the target image based on a pre-built lightweight neural network classifier includes: The target image is input into the lightweight neural network classifier through the input layer; feature vectors are extracted through the feature extraction layer; the feature vectors are integrated and mapped through the fully connected layer to obtain the probability distribution of each category; and the category of the target image is output through the output layer based on the probability distribution.

4. The method according to claim 1, characterized in that, The anomaly prediction of the target image using the structure-aware reconstruction anomaly detector includes: The target image is reconstructed using an encoder and a decoder to obtain a reconstructed image; Construct a multidimensional reconstruction loss function that includes pixel loss and gradient domain loss; Based on the target image and the reconstructed image, an abnormal heatmap or a mask is generated by combining structure-aware reconstruction errors.

5. The method according to claim 1, characterized in that, Anomaly prediction of the target image using the improved pseudo-defect classifier includes: Define a pseudo-transform set including different image enhancement methods, use the elements of the pseudo-transform set to generate pseudo samples based on the target image, and construct a pseudo sample set; A parallel multi-branch discriminant structure is constructed, wherein the multi-branch discriminant structure includes parameter sub-branches corresponding to different elements of the pseudo-transform set; wherein each sub-branch supports hot-plugging; The classification results are obtained using the parallel multi-branch discriminant structure.

6. The method according to claim 1, characterized in that, Using the semantic fusion contrast anomaly model to predict anomalies in the target image includes: The global vector of the target image is extracted using an encoder; Construct natural language category guidance text prompts and obtain the embedding vectors of the guidance text prompts; The target image is segmented to obtain several local attention maps. A guide map is generated by fusing the local attention maps with the embedding vector. An abnormal heatmap or a mask is generated based on the semantic difference between the guide map and the embedding vector. The semantic fusion comparison anomaly model also introduces a learnable prompt generator corresponding to different anomaly categories, which is concatenated with the guidance text prompt to update the guidance text prompt.

7. A scalable category-guided anomaly detection device for multiple target types, characterized in that, The device includes: The preprocessing module is used to acquire industrial images, perform adaptive enhancement preprocessing guided by abnormal regions on the industrial images, and acquire target images. A coarse classification module is used to determine the category of the target image based on a pre-built lightweight neural network classifier; A dynamic scheduling and routing module is used to activate at least one anomaly detection model from a pre-constructed subset of structurally heterogeneous anomaly detection models according to the category of the target image; wherein the anomaly detection model includes a structure-aware reconstruction anomaly detector, an improved pseudo-defect classifier, and a semantic fusion comparison anomaly model; The output module is used to perform anomaly prediction on the target image using the activated anomaly detection model, and to fuse the prediction with the confidence level of the anomaly detection model to achieve adaptive anomaly discrimination of the industrial image.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Single-stage object detection method based on lightweight image pyramid network

    CN110245655A

  • Two-stage image restoration method based on texture structure perception

    CN112801914A

  • Method and device for determining object category, electronic equipment and storage medium

    CN114529768A

  • Business anomaly detection method and device

    CN115766507A

  • Industrial scene-oriented multi-strategy image abnormal sample generation method

    CN116091987A

Cited By

  • Multi-class anomaly detection method based on memory guidance and class decoupling

    CN121600331A

  • A multi-class anomaly detection method based on memory guidance and category decoupling

    CN121600331B