Image anomaly detection method and device for industrial scene and electronic equipment

By constructing an image anomaly detection model for teacher networks, student networks and encoder networks in industrial scenarios, the problem of insufficient adaptability and real-time performance of image anomaly detection in industrial scenarios is solved, and higher detection accuracy and efficiency are achieved.

CN120070335APending Publication Date: 2025-05-30SHENZHEN LINGYUN VISION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063428.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has weak adaptability to image abnormality detection in industrial scenarios, and it is difficult to deal with diverse abnormal patterns, and has high calculation costs and insufficient real-time performance.

Method used

An image anomaly detection method for industrial scenarios is proposed. By constructing an image anomaly detection model including teacher network, student network and encoder network, the collaborative work of these networks is used to detect abnormalities on the detected images.

Benefits of technology

It improves the accuracy and efficiency of image abnormality detection in industrial scenarios, and can more accurately detect abnormal areas in the image to be detected, which is suitable for complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070335A_ABST
    Figure CN120070335A_ABST
Patent Text Reader

Abstract

The invention discloses an image anomaly detection method and device for an industrial scene and electronic equipment, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a to-be-detected image in an industrial scene; the to-be-detected image is input into an image anomaly detection model, an anomaly thermodynamic diagram of the to-be-detected image is determined, and the image anomaly detection model is obtained based on normal image sample training; and determining an anomaly detection result of the to-be-detected image according to the anomaly thermodynamic diagram of the to-be-detected image. The image anomaly detection method realizes anomaly detection of the to-be-detected image, is applied to an industrial scene, and improves the accuracy and efficiency of image anomaly detection in the industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to an image anomaly detection method, device, and electronic device for industrial scenarios. Background Art

[0002] With the rapid development of industrial automation and intelligent manufacturing, image detection technology is increasingly widely used in industrial scenarios. This technology monitors the status of equipment, products, and production processes through image detection, achieving real-time, non-contact quality control and fault detection. However, the application of existing image anomaly detection technologies in actual industrial scenarios still faces many challenges. Specifically, when traditional image processing technologies perform anomaly detection, in the face of complex industrial environments, their adaptability is weak, it is difficult to handle diverse anomaly forms, and it is not easy to expand functions; for image detection methods based on statistics, although anomaly detection is performed by setting thresholds, a single statistical method cannot comprehensively cover all anomaly types, and when the data volume is large, the calculation cost is high, making it difficult to ensure real-time performance; common image anomaly detection methods based on machine learning also have certain application limitations due to factors such as the scarcity of anomaly samples.

[0003] Therefore, how to provide an image anomaly detection solution that can adapt to complex industrial scenarios and has higher detection accuracy and real-time performance is an urgent problem to be solved in this field. Summary of the Invention

[0004] This application aims to solve at least one of the technical problems existing in the related technologies. For this purpose, this application proposes an image anomaly detection method, device, and electronic device for industrial scenarios, which realizes anomaly detection of the image to be detected, is applied to industrial scenarios, and improves the accuracy and efficiency of image anomaly detection in industrial scenarios.

[0005] In a first aspect, this application provides an image anomaly detection method for industrial scenarios, and the method includes:

[0006] Obtain an image to be detected in an industrial scenario;

[0007] Input the image to be detected into an image anomaly detection model to determine an anomaly heat map of the image to be detected, and the image anomaly detection model is trained based on normal image samples;

[0008] Determine an anomaly detection result of the image to be detected according to the anomaly heat map of the image to be detected.

[0009] In the above technical solution, a to-be-detected image in an industrial scenario is obtained, and the to-be-detected image is input into an image anomaly detection model trained based on normal image samples to obtain an anomaly heat map of the to-be-detected image. According to the anomaly heat map of the to-be-detected image, the anomaly detection result of the to-be-detected image can be determined, realizing the anomaly detection of the to-be-detected image, which is applied to the industrial scenario and improves the accuracy and efficiency of image anomaly detection in the industrial scenario.

[0010] According to an embodiment of the present application, the image anomaly detection model is trained through the following steps:

[0011] Construct a training sample set, where the training sample set includes normal image samples;

[0012] Construct an image anomaly detection network, where the image anomaly detection network includes: a teacher network, a student network, and an encoder network. The number of channels of the student network is twice that of the teacher network, and the output of the teacher network has the same scale as the output of the encoder network;

[0013] Input the normal image sample into the teacher network to obtain the teacher features of the normal image sample;

[0014] Input the normal image sample into the student network to obtain a first student feature and a second student feature of the normal image sample. The first student feature is the output of the first half of the channels of the student network, and the second student feature is the output of the second half of the channels of the student network;

[0015] Input the normal image sample into the encoder network to obtain the encoder features of the normal image sample;

[0016] Based on the teacher features and the first student feature, calculate a first loss. Based on the second student feature and the encoder feature, calculate a second loss. Based on the teacher feature and the encoder feature, calculate a third loss;

[0017] Freeze the parameters of the teacher network, optimize the parameters of the student network and the encoder network based on the first loss, the second loss, and the third loss. When the training end condition is met, retain the parameters of the student network and the encoder network to obtain the image anomaly detection model.

[0018] In the above technical solution, a training sample set including normal image samples is constructed, and an image anomaly detection network including a teacher network, a student network, and an encoder network is constructed. Among them, the number of channels of the student network is twice that of the teacher network, and the output scales of the teacher network and the encoder network are the same. The normal image samples are input into the teacher network to obtain the teacher features of the normal image samples. The normal image samples are input into the student network to obtain the first student features of the normal image samples output by the first half of the channels of the student network and the second student features output by the second half of the channels of the student network. The normal image samples are input into the encoder network to obtain the encoder features of the normal image samples. Based on the teacher features and the first student features, a first loss is calculated. Based on the second student features and the encoder features, a second loss is calculated. Based on the teacher features and the encoder features, a third loss is calculated. The parameters of the teacher network are frozen, and the parameters of the student network and the encoder network are optimized based on the first loss, the second loss, and the third loss. When the training end condition is met, the parameters of the student network and the encoder network are retained to obtain the image anomaly detection model, realizing the effective training of the image anomaly detection model. Through the collaborative work of the teacher network, the student network, and the encoder network, the trained image anomaly detection model can more accurately detect image anomalies to be detected, improving the accuracy of image anomaly detection in industrial scenarios.

[0019] According to an embodiment of the present application, the teacher network and the student network are obtained by loading a pre-trained lightweight feature extraction network, and the pre-trained lightweight feature extraction network is obtained through the following steps:

[0020] Construct a lightweight feature extraction network;

[0021] Transfer the knowledge of the deep convolutional neural network to the lightweight feature extraction network through knowledge distillation technology to obtain a pre-trained lightweight feature extraction network.

[0022] In the above technical solution, a lightweight feature extraction network is constructed, and the knowledge of the deep convolutional neural network is transferred to the lightweight feature extraction network through knowledge distillation technology to obtain a pre-trained lightweight feature extraction network. The teacher network and the student network are obtained by loading the pre-trained lightweight feature extraction network, reducing the training time of the teacher network and the student network, thereby improving the training efficiency of the image anomaly detection model. The lightweight feature extraction network reduces the number of parameters and computational overhead of the image anomaly detection model in the process of image anomaly detection. Applied to industrial scenarios, it improves the efficiency of image anomaly detection in industrial scenarios.

[0023] According to an embodiment of the present application, the image anomaly detection model includes: a teacher network, a student network, and an encoder network,

[0024] The teacher network is used to extract features from the image to be detected, obtaining the teacher features of the image to be detected;

[0025] The student network is used to extract features from the image to be detected, obtaining the first student features and the second student features of the image to be detected. The first student features are the output of the first half of the channels of the student network, and the second student features are the output of the second half of the channels of the student network;

[0026] The encoder network is used to reconstruct the image to be detected, obtaining the encoder features of the image to be detected.

[0027] In the above technical solution, the image anomaly detection model includes a teacher network, a student network, and an encoder network. The teacher network is used to extract features from the image to be detected, obtaining the teacher features of the image to be detected. The student network is used to extract features from the image to be detected, obtaining the first student features and the second student features of the image to be detected. The first student features are the output of the first half of the channels of the student network, and the second student features are the output of the second half of the channels of the student network. The encoder network is used to reconstruct the image to be detected, obtaining the encoder features of the image to be detected. The teacher features, encoder features, first student features, and second student features of the image to be detected can be used to determine the anomaly heat map of the image to be detected, thereby realizing the anomaly detection of the image to be detected.

[0028] According to an embodiment of the present application, the inputting the image to be detected into the image anomaly detection model to determine the anomaly heat map of the image to be detected includes:

[0029] Inputting the image to be detected into the teacher network, obtaining the teacher features of the image to be detected;

[0030] Inputting the image to be detected into the student network, obtaining the first student features and the second student features of the image to be detected;

[0031] Inputting the image to be detected into the encoder network, obtaining the encoder features of the image to be detected;

[0032] Determining the anomaly heat map of the image to be detected according to the teacher features, the first student features, the second student features, and the encoder features.

[0033] In the above technical solution, the image to be detected is input into the teacher network to obtain the teacher features of the image to be detected. The image to be detected is input into the student network to obtain the first student features and the second student features of the image to be detected. The image to be detected is input into the encoder network to obtain the encoder features of the image to be detected. According to the teacher features, the first student features, the second student features, and the encoder features, an anomaly heatmap of the image to be detected is determined. Furthermore, the anomaly detection result of the image to be detected can be determined based on the anomaly heatmap of the image to be detected. By combining the teacher network, the student network, and the encoder network, the anomaly detection of the image to be detected is realized, which is applied to the industrial scenario and improves the accuracy of image anomaly detection in the industrial scenario.

[0034] According to an embodiment of the present application, the determining the anomaly heatmap of the image to be detected according to the teacher features, the first student features, the second student features, and the encoder features includes:

[0035] Determining a structural anomaly heatmap of the image to be detected according to the teacher features of the image to be detected and the first student features of the image to be detected;

[0036] Determining a logical anomaly heatmap of the image to be detected according to the second student features of the image to be detected and the encoder features of the image to be detected;

[0037] Determining the anomaly heatmap of the image to be detected according to the structural anomaly heatmap of the image to be detected and the logical anomaly heatmap of the image to be detected.

[0038] In the above technical solution, a structural anomaly heatmap of the image to be detected is determined according to the teacher features of the image to be detected and the first student features of the image to be detected; a logical anomaly heatmap of the image to be detected is determined according to the second student features of the image to be detected and the encoder features of the image to be detected; the anomaly heatmap of the image to be detected is determined according to the structural anomaly heatmap of the image to be detected and the logical anomaly heatmap of the image to be detected. By using the difference in the ability of the teacher network and the student network to extract anomaly regions, the structural anomaly is determined, thereby determining the structural anomaly heatmap, and by using the difference in the reconstruction results of the encoder network and the student network for the image to be detected, the logical anomaly is determined, thereby determining the logical anomaly heatmap. Applied to the industrial scenario, it improves the accuracy and robustness of image anomaly detection in the industrial scenario.

[0039] According to an embodiment of the present application, the determining the anomaly detection result of the image to be detected according to the anomaly heatmap of the image to be detected includes:

[0040] Based on the anomaly heatmap of the image to be detected, determining the anomaly score of the image to be detected;

[0041] Compare the anomaly score of the image to be detected with a preset threshold to determine whether there is an anomaly in the image to be detected;

[0042] And / or, based on the anomaly heat map of the image to be detected, determine the binary map of the image to be detected, where the binary map is used to display the location of the anomaly area in the image to be detected.

[0043] In the above technical solution, based on the anomaly heat map of the image to be detected, determine the anomaly score of the image to be detected, compare the anomaly score of the image to be detected with a preset threshold to determine whether there is an anomaly in the image to be detected, and / or, based on the anomaly heat map of the image to be detected, determine the binary map of the image to be detected, where the binary map is used to display the location of the anomaly area in the image to be detected, and the anomaly detection result of the image to be detected is determined, thereby realizing the anomaly detection of the image to be detected.

[0044] In a second aspect, the present application provides an image anomaly detection device for an industrial scenario, and the device includes:

[0045] An image acquisition unit for acquiring an image to be detected in an industrial scenario;

[0046] An anomaly detection unit for inputting the image to be detected into an image anomaly detection model to determine the anomaly heat map of the image to be detected, where the image anomaly detection model is trained based on normal image samples;

[0047] An anomaly determination unit for determining the anomaly detection result of the image to be detected according to the anomaly heat map of the image to be detected.

[0048] In the above technical solution, the image anomaly detection device for an industrial scenario acquires an image to be detected in an industrial scenario, inputs the image to be detected into an image anomaly detection model trained based on normal image samples to obtain the anomaly heat map of the image to be detected, and according to the anomaly heat map of the image to be detected, the anomaly detection result of the image to be detected can be determined, realizing the anomaly detection of the image to be detected, and being applied to an industrial scenario, improving the accuracy and efficiency of image anomaly detection in an industrial scenario.

[0049] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the image anomaly detection method for an industrial scenario as described in the first aspect above.

[0050] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the image anomaly detection method for an industrial scenario as described in the first aspect above.

[0051] Fifth aspect, the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the image anomaly detection method for industrial scenarios as described in the first aspect.

[0052] Sixth aspect, the present application provides a computer program product, including a computer program, which when executed by a processor, implements the image anomaly detection method for industrial scenarios as described in the first aspect above.

[0053] The additional aspects and advantages of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application. Description of the Drawings

[0054] The above and / or additional aspects and advantages of the present application will become apparent and be easily understood from the description of the embodiments in conjunction with the following drawings, where:

[0055] Figure 1 is a schematic flowchart of the image anomaly detection method for industrial scenarios provided by some embodiments of the present application;

[0056] Figure 2 is one of the schematic flowcharts of training the image anomaly detection model provided by some embodiments of the present application;

[0057] Figure 3 is another schematic flowchart of training the image anomaly detection model provided by some embodiments of the present application;

[0058] Figure 4 is a schematic structural diagram of the teacher network provided by some embodiments of the present application;

[0059] Figure 5 is a schematic flowchart of determining the anomaly detection result provided by some embodiments of the present application;

[0060] Figure 6 is a schematic structural diagram of the image anomaly detection device for industrial scenarios provided by some embodiments of the present application;

[0061] Figure 7 is a schematic structural diagram of the electronic device provided by some embodiments of the present application.

[0062] Description of the Reference Numerals:

[0063] 600: Image anomaly detection device for industrial scenarios; 601: Image acquisition unit; 602: Anomaly detection unit; 603: Anomaly determination unit; 700: Electronic device; 701: Processor; 702: Memory. Detailed Embodiments

[0064] The technical solutions in the embodiments of the present application will be clearly described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0065] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0066] With the rapid development of industrial automation and intelligent manufacturing, image detection technology is increasingly widely used in industrial scenarios. By detecting images to monitor the status of equipment, products, and production processes, real-time and non-contact quality control and fault detection can be achieved. However, the application of existing image anomaly detection technology in actual industrial scenarios still faces many challenges. The following are several current common technical solutions and their limitations:

[0067] (1) Anomaly detection based on traditional image processing: Traditional image processing techniques rely on algorithms such as edge detection, texture analysis, and color feature extraction. These methods detect abnormal parts in images by manually designing features. For example, surface defects are detected using information such as the gradient and gray histogram of the image. However, this rule-based detection method has poor adaptability to complex industrial environments, cannot cope with diverse anomaly forms, especially when facing noise interference and environmental changes, the false detection rate is relatively high, and the scalability is limited.

[0068] (2) Statistical image detection methods: Some anomaly detection techniques analyze abnormal situations in images based on the statistical characteristics of image data, such as mean, variance, etc. Such methods can judge the normality or abnormality of images by setting thresholds. However, anomalies in industrial scenarios usually have diversity, and a single statistical method often cannot cover all anomaly types, and when the data volume is large, the calculation cost is high and the real-time performance is insufficient.

[0069] (3) Image anomaly detection based on machine learning: With the rise of machine learning, many researchers have tried to improve the performance of image anomaly detection through supervised learning or unsupervised learning methods. Supervised learning methods such as convolutional neural networks (CNNs) rely on a large amount of labeled normal and abnormal image data for training. However, labeling abnormal data often requires a large amount of human cost, and in industrial scenarios, abnormal samples are usually scarce, which limits the popularization and application of these methods. Unsupervised learning methods, such as autoencoders or generative adversarial networks (GANs), detect abnormal data that deviates from the normal pattern by learning the feature distribution of normal images, which alleviates the problem of dependence on labeled data to a certain extent. However, when faced with complex environments, lighting changes, and diverse abnormal patterns, their detection effects are still not robust enough.

[0070] To solve the above technical problems, the embodiments of the present application provide an image anomaly detection method, device, and electronic device for industrial scenarios. The following will be described in detail through specific embodiments and their application scenarios with reference to the accompanying drawings.

[0071] The image anomaly detection method for industrial scenarios provided by the embodiments of the present application may be executed by an electronic device or a functional module or entity in the electronic device that can implement the image anomaly detection method for industrial scenarios. The electronic devices mentioned in the embodiments of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices, etc. The following takes the electronic device as the execution subject to illustrate the image anomaly detection method for industrial scenarios provided by the embodiments of the present application.

[0072] Figure 1 is a schematic flowchart of the image anomaly detection method for industrial scenarios provided by some embodiments of the present application. As Figure 1 shown, the image anomaly detection method for industrial scenarios includes: step 110, step 120, and step 130.

[0073] Step 110: Obtain the image to be detected in the industrial scenario.

[0074] Step 120: Input the image to be detected into the image anomaly detection model to determine the anomaly heat map of the image to be detected. The image anomaly detection model is trained based on normal image samples.

[0075] It can be understood that in an industrial scenario, abnormal situations are not common, and the cost of annotating abnormal image samples is relatively high. Therefore, it is difficult to collect a large number of abnormal image samples under normal production conditions. The image anomaly detection model in the embodiments of the present application is trained based on normal image samples, that is, it does not rely on abnormal image samples during the training process, and only uses abnormal image samples in the stage of testing the image anomaly detection model to evaluate the model's ability to detect image anomalies, reducing the cost of model training and improving the applicability and flexibility of image anomaly detection in actual industrial scenarios using the image anomaly detection model.

[0076] The trained image anomaly detection model can effectively detect the abnormal regions in the image to be detected that do not match the normal image samples. The image to be detected is input into the trained image anomaly detection model to determine the abnormal heat map of the image to be detected, and then the abnormal detection result of the image to be detected is determined based on the abnormal heat map of the image to be detected, reducing the labor cost and time cost and improving the accuracy and efficiency of abnormal detection of the image to be detected.

[0077] Step 130: Determine the abnormal detection result of the image to be detected according to the abnormal heat map of the image to be detected.

[0078] It can be understood that the abnormal heat map is a visual representation of the abnormalities in the image to be detected, and the degree of abnormality of each region in the image to be detected is represented by the color intensity. Optionally, a threshold is set, and when the degree of abnormality of a certain region in the abnormal heat map exceeds the threshold, it is determined that there is an abnormality in that region. Optionally, the abnormal detection result of the image to be detected includes information such as the location and size of the abnormal region in the image to be detected.

[0079] In the above technical solution, the image to be detected in the industrial scenario is obtained, the image to be detected is input into the image anomaly detection model trained based on normal image samples to obtain the abnormal heat map of the image to be detected, and the abnormal detection result of the image to be detected can be determined according to the abnormal heat map of the image to be detected, realizing the abnormal detection of the image to be detected, applying it to the industrial scenario, and improving the accuracy and efficiency of image anomaly detection in the industrial scenario.

[0080] Figure 2 It is one of the flow diagrams for training an image anomaly detection model provided by some embodiments of the present application. As Figure 2 shown, in some embodiments of the present application, the image anomaly detection model is trained through the following steps:

[0081] Step 210: Construct a training sample set, where the training sample set includes normal image samples.

[0082] Step 220: Construct an image anomaly detection network, which includes a teacher network, a student network, and an encoder network. The number of channels of the student network is twice that of the teacher network, and the output of the teacher network has the same scale as the output of the encoder network.

[0083] Figure 3 It is the second schematic diagram of the process for training an image anomaly detection model provided by some embodiments of the present application. As Figure 3 shown, the image anomaly detection network includes a teacher network, a student network, and an encoder network.

[0084] It can be understood that in some embodiments of the present application, both the teacher network and the student network are obtained by loading a pre-trained lightweight feature extraction network, and the number of channels of the student network is twice that of the teacher network. After division, they are respectively used to learn the feature representation of the teacher network and the feature representation of the encoder network.

[0085] The encoder network is an independent network that has not been pre-trained, but is optimized together with the student network during the training process. During the training process of the image anomaly detection model, the encoder network attempts to reconstruct the input normal image samples to output a feature representation with the same scale as the teacher network.

[0086] The number of channels refers to the number of feature maps output by the convolutional layer. The number of channels of the student network being twice that of the teacher network means that the student network will have twice the number of feature maps of the teacher network at the corresponding layer. For example, if a certain convolutional layer of the teacher network outputs 64 feature maps, then the corresponding convolutional layer of the student network will output 128 feature maps. The channels of the student network are divided to obtain the first half channels and the second half channels of the student network.

[0087] The output scale refers to the size of the feature map output by the last layer of the network. When the output of the teacher network has the same scale as the output of the encoder network, it means that they have the same dimension in the feature map generated by the last layer (for example, both are 64x64 pixels).

[0088] Step 230: Input the normal image samples into the teacher network to obtain the teacher features of the normal image samples.

[0089] Step 240: Input the normal image samples into the student network to obtain the first student features and the second student features of the normal image samples. The first student features are output by the first half channels of the student network, and the second student features are output by the second half channels of the student network.

[0090] Step 250: Input the normal image samples into the encoder network to obtain the encoder features of the normal image samples.

[0091] It can be understood that during the training process, the encoder network is used to reconstruct normal image samples to obtain the encoder features of the normal image samples.

[0092] Step 260: Calculate a first loss based on the teacher features and the first student features, calculate a second loss based on the second student features and the encoder features, and calculate a third loss based on the teacher features and the encoder features.

[0093] It can be understood that based on the teacher features and the first student features, a first loss is calculated. The first loss is used to measure the difference between the first student features output by the first half channels of the student network and the teacher features output by the teacher network, and can be used to guide the first half channels of the student network to learn the feature representation of the teacher network.

[0094] Based on the second student features and the encoder features, a second loss is calculated. The second loss is used to measure the difference between the second student features output by the second half channels of the student network and the encoder features output by the encoder network, and can be used to guide the second half channels of the student network to learn the feature representation of the encoder network.

[0095] Based on the teacher features and the encoder features, a third loss is calculated. The third loss is used to measure the difference between the teacher features output by the teacher network and the encoder features output by the encoder network, which helps to improve the consistency between the feature representation of the encoder network and the feature representation of the teacher network.

[0096] It can be understood that in this technical solution, the loss function includes the first loss, the second loss, and the third loss, and no additional losses such as penalty term loss (the second norm of the results of the student network on the dataset) are introduced, avoiding problems such as increased storage space that may be caused by additional losses, making the image anomaly detection model more applicable to resource-constrained industrial devices, and improving the applicability of the image anomaly detection method in industrial scenarios.

[0097] Step 270: Freeze the parameters of the teacher network, optimize the parameters of the student network and the encoder network based on the first loss, the second loss, and the third loss, and when the training end condition is met, retain the parameters of the student network and the encoder network to obtain the image anomaly detection model.

[0098] It can be understood that the parameters of the teacher network are frozen during training, which can avoid overfitting during the training process using only normal image samples. By freezing the parameters, the teacher network can retain the feature representation learned from normal image samples and also has the ability to extract features from abnormal regions if there are abnormal regions in the image.

[0099] Different from the teacher network, the parameters of the student network can be freely updated during training. Its training objective is to overfit the features of normal image samples as much as possible by training on normal image samples. This training method enables the student network to perform excellently on normal image samples. However, if there are abnormal regions in the image, its feature extraction ability is limited, and it is difficult to effectively extract the features of the abnormal regions.

[0100] The differential training mechanism of the teacher network and the student network provides a basis for anomaly detection. For example, when an image to be detected with anomalies is input into the image anomaly detection model, because the parameters of the teacher network are frozen, it still has the ability to extract the features of the abnormal regions. While the student network, due to overfitting the features of normal image samples, is difficult to extract the features of the abnormal regions. By comparing the differences between the teacher features output by the teacher network and the first student features output by the first half of the channels of the student network, the structural anomalies in the image to be detected can be effectively detected. Structural anomalies usually refer to local features or attributes in the image to be detected that are different from normal image samples, manifested as cracks, dents, holes, contaminations, damages or other local area defects that usually do not exist in normal image samples.

[0101] During training, the encoder network attempts to reconstruct the input normal image samples and optimize based on the third loss. The encoder network can learn the features of normal image samples and adapt to situations such as rotation, color differences, and brightness changes. However, since abnormal image samples are not used during the training process, the encoder network cannot reconstruct the features of abnormal image samples. In addition, when there are texture features in the image samples, there will also be large errors in the reconstruction of the encoder network. During this process, align the second student features output by the second half of the channels of the student network with the encoder features output by the encoder network, so that the student network can learn the reconstruction features (i.e., encoder features) of a blurred and less background - focused normal image sample. When an image to be detected with anomalies is input into the image anomaly detection model, by comparing the differences between the encoder features output by the encoder network and the second student features output by the second half of the channels of the student network in the trained image anomaly detection model, the logical anomalies in the abnormal image samples can be effectively detected. Logical anomalies refer to anomalies in the image to be detected that violate the potential logical constraints of the content in normal image samples. For example, the number, arrangement, or composition of components does not meet expectations, or there are minor deviations in details or color anomalies.

[0102] It can be understood that through the above training steps, differential learning of the student network is achieved. That is, the first half of the channels of the student network learn the feature representation of the teacher network, and the second half of the channels of the student network learn the feature representation of the encoder network. Furthermore, structural anomalies can be determined based on the teacher features output by the teacher network and the first student features output by the first half of the channels of the student network, and logical anomalies can be determined based on the encoder features output by the encoder network and the second student features output by the second half of the channels of the student network. This realizes the differentiation between logical anomalies and structural anomalies while performing anomaly detection, improving the accuracy of anomaly detection for images, and is particularly suitable for detecting complex anomalies in industrial images.

[0103] In the above technical solution, a training sample set including normal image samples is constructed, and an image anomaly detection network including a teacher network, a student network, and an encoder network is constructed. Among them, the number of channels of the student network is twice that of the teacher network, and the output scales of the teacher network and the encoder network are the same. The normal image samples are input into the teacher network to obtain the teacher features of the normal image samples. The normal image samples are input into the student network to obtain the first student features of the normal image samples output by the first half of the channels of the student network and the second student features output by the second half of the channels of the student network. The normal image samples are input into the encoder network to obtain the encoder features of the normal image samples. Based on the teacher features and the first student features, a first loss is calculated. Based on the second student features and the encoder features, a second loss is calculated. Based on the teacher features and the encoder features, a third loss is calculated. The parameters of the teacher network are frozen, and the parameters of the student network and the encoder network are optimized based on the first loss, the second loss, and the third loss. When the training end condition is met, the parameters of the student network and the encoder network are retained to obtain the image anomaly detection model, realizing the effective training of the image anomaly detection model. Through the collaborative work of the teacher network, the student network, and the encoder network, the trained image anomaly detection model can more accurately detect image anomalies to be detected, improving the accuracy of image anomaly detection in industrial scenarios.

[0104] In some embodiments of the present application, the teacher network and the student network are obtained by loading a pre-trained lightweight feature extraction network, and the pre-trained lightweight feature extraction network is obtained through the following steps:

[0105] Construct a lightweight feature extraction network;

[0106] Transfer the knowledge of the deep convolutional neural network to the lightweight feature extraction network through knowledge distillation technology to obtain the pre-trained lightweight feature extraction network.

[0107] It can be understood that, in order to reduce the complexity of the model to adapt to the computing resource limitations of industrial scenarios, a lightweight feature extraction network is constructed. Through structural optimization, this network can reduce the number of parameters and computational overhead while ensuring the accuracy of feature extraction, enabling it to operate efficiently in embedded devices or edge computing environments. To incorporate stronger feature learning capabilities into the lightweight feature extraction network, knowledge distillation technology can be used to transfer the knowledge of a deep convolutional neural network to this lightweight feature extraction network.

[0108] Optionally, the deep convolutional neural network is a WideResNet101 network. Transferring the knowledge of the WideResNet101 network to the lightweight feature extraction network through knowledge distillation technology specifically includes:

[0109] (1) Select a subset from the ImageNet dataset as the distillation training data;

[0110] (2) Distill the feature knowledge of the WideResNet101 network into the lightweight feature extraction network by minimizing the mean squared error (MSE Loss) between the outputs of the lightweight feature extraction network and the WideResNet101 network.

[0111] The above distillation process enables the lightweight feature extraction network to obtain a feature extraction ability similar to that of the WideResNet101 network, and has the advantage of being lightweight, that is, while reducing the consumption of computing resources, it can still efficiently process feature extraction tasks.

[0112] Figure 4 It is a schematic diagram of the structure of the teacher network provided by some embodiments of the present application. As Figure 4 shown, in some embodiments, the teacher network is obtained by loading the lightweight feature extraction network. The teacher network receives a normal image sample with a size of 256x256 and 3 channels as input and outputs a feature map with a size of 64x64 and 192 channels (i.e., the teacher feature of the normal image sample). The teacher network consists of four convolutional layers and two max-pooling layers. Among them, Kx represents that the kernel size of the current layer is x, Sx represents that the stride of the current layer is x, and Dx represents that after being processed by the current layer, the number of channels of the output feature is x. That is, the teacher network includes:

[0113] The first convolutional layer (Conv1), the convolutional kernel size of the first convolutional layer is 4, the stride is 1, and after being processed by the first convolutional layer, the number of output feature channels is 64;

[0114] The second convolutional layer (Conv2), the convolutional kernel size of the second convolutional layer is 4, the stride is 1, and after being processed by the second convolutional layer, the number of output feature channels is 64;

[0115] The third convolutional layer (Conv3) has a convolutional kernel size of 3 and a stride of 1. After being processed by the third convolutional layer, the number of output feature channels is 128;

[0116] The fourth convolutional layer (Conv4) has a convolutional kernel size of 4 and a stride of 1. After being processed by the fourth convolutional layer, the number of output feature channels is 192;

[0117] The first max pooling layer (MaxPool1) has a pooling kernel size of 2 and a stride of 2;

[0118] The second max pooling layer (MaxPool2) has a pooling kernel size of 2 and a stride of 2.

[0119] Double the number of channels of the teacher network to obtain the corresponding student network.

[0120] In the embodiments of the present application, the lightweight feature extraction network has a relatively shallow depth and adopts a convolution design without padding, which can prevent the features Figure Four around the convolution output from appearing abnormally.

[0121] The teacher network and the student network are obtained by loading the pre-trained lightweight feature extraction network, which reduces the training time of the teacher network and the student network, thereby improving the training efficiency of the image anomaly detection model, and also reducing the number of parameters and computational overhead of the image anomaly detection model during the image anomaly detection process, and improving the efficiency of image anomaly detection.

[0122] In the above technical solution, a lightweight feature extraction network is constructed, and the knowledge of the deep convolutional neural network is transferred to the lightweight feature extraction network through the knowledge distillation technology to obtain a pre-trained lightweight feature extraction network. The teacher network and the student network are obtained by loading the pre-trained lightweight feature extraction network, which reduces the training time of the teacher network and the student network, thereby improving the training efficiency of the image anomaly detection model. The lightweight feature extraction network reduces the number of parameters and computational overhead of the image anomaly detection model during the image anomaly detection process, and is applied to the industrial scenario, improving the efficiency of image anomaly detection in the industrial scenario.

[0123] In some embodiments of the present application, the image anomaly detection model includes: a teacher network, a student network, and an encoder network,

[0124] The teacher network is used to extract features from the image to be detected to obtain the teacher features of the image to be detected;

[0125] The student network is used to extract features from the image to be detected, obtaining a first student feature and a second student feature of the image to be detected. The first student feature is the output of the first half of the channels of the student network, and the second student feature is the output of the second half of the channels of the student network;

[0126] The encoder network is used to reconstruct the image to be detected, obtaining an encoder feature of the image to be detected.

[0127] It can be understood that the channels of the student network are divided into the first half and the second half. The channels in the first half of the student network are used to learn the feature representation of the teacher network, and the channels in the second half of the student network are used to learn the feature representation of the encoder network, respectively outputting the first student feature and the second student feature of the image to be detected.

[0128] The encoder network is used to reconstruct the image to be detected to obtain the encoder feature of the image to be detected. This process is usually based on auto encoding of the image, that is, the encoder network first encodes the image to be detected into a low-dimensional latent space representation, and then reconstructs the original image from this representation so that the reconstructed image is as similar as possible to the original image.

[0129] In the above technical solution, the image anomaly detection model includes a teacher network, a student network, and an encoder network. The teacher network is used to extract features from the image to be detected, obtaining a teacher feature of the image to be detected. The student network is used to extract features from the image to be detected, obtaining a first student feature and a second student feature of the image to be detected. The first student feature is the output of the first half of the channels of the student network, and the second student feature is the output of the second half of the channels of the student network. The encoder network is used to reconstruct the image to be detected, obtaining an encoder feature of the image to be detected. The teacher feature, encoder feature, first student feature, and second student feature of the image to be detected can be used to determine the anomaly heat map of the image to be detected, thereby realizing the anomaly detection of the image to be detected.

[0130] In some embodiments of the present application, the inputting the image to be detected into the image anomaly detection model and determining the anomaly heat map of the image to be detected includes:

[0131] Inputting the image to be detected into the teacher network to obtain a teacher feature of the image to be detected;

[0132] Inputting the image to be detected into the student network to obtain a first student feature and a second student feature of the image to be detected;

[0133] Inputting the image to be detected into the encoder network to obtain an encoder feature of the image to be detected;

[0134] Determine the abnormal heat map of the image to be detected according to the teacher features, the first student features, the second student features, and the encoder features.

[0135] In some embodiments, the determining the abnormal heat map of the image to be detected according to the teacher features, the first student features, the second student features, and the encoder features includes:

[0136] Determine the structural abnormal heat map of the image to be detected according to the teacher features of the image to be detected and the first student features of the image to be detected;

[0137] Determine the logical abnormal heat map of the image to be detected according to the second student features of the image to be detected and the encoder features of the image to be detected;

[0138] Determine the abnormal heat map of the image to be detected according to the structural abnormal heat map of the image to be detected and the logical abnormal heat map of the image to be detected.

[0139] Specifically, perform per-pixel difference calculation on the teacher features of the image to be detected and the first student features of the image to be detected to obtain the teacher features of the image to be detected and the first student features of the image to be detected.

[0140] Perform per-pixel difference calculation on the second student features of the image to be detected and the encoder features of the image to be detected to obtain the logical abnormal heat map of the image to be detected.

[0141] It can be understood that the differential training mechanism of the teacher network and the student network provides a basis for anomaly detection. If there is an abnormal area in the image to be detected, the teacher network, due to the freezing of its parameters, still has the ability to extract the features of the abnormal area, while the student network, due to overfitting the features of normal image samples, is difficult to extract the feature representation of the abnormal area. By comparing the differences between the teacher features output by the teacher network and the first student features output by the first half of the channels of the student network in the trained image anomaly detection model, the structural anomalies in the image to be detected can be effectively detected.

[0142] It can be understood that since abnormal image samples are not used in the training process, if there is an abnormal area in the image to be detected, the encoder network cannot reconstruct the features of the abnormal area in the image to be detected. By comparing the differences between the encoder features output by the encoder network and the second student features output by the second half of the channels of the student network in the trained image anomaly detection model, the logical anomalies in the image to be detected can be effectively detected.

[0143] The structural abnormal heat map is a visual representation of the structural anomalies in the image to be detected, and the degree of structural anomalies in each area of the image to be detected is represented by the color intensity.

[0144] The logical anomaly heat map is a visual representation of logical anomalies in the image to be detected, and the degree of logical anomalies in each region of the image to be detected is represented by color intensity.

[0145] In some embodiments, determining the anomaly heat map of the image to be detected according to the structural anomaly heat map of the image to be detected and the logical anomaly heat map of the image to be detected includes:

[0146] Overlaying the structural anomaly heat map of the image to be detected and the logical anomaly heat map of the image to be detected to obtain the anomaly heat map of the image to be detected.

[0147] Optionally, overlay the structural anomaly heat map and the logical anomaly heat map according to a preset weight. For each pixel point in the anomaly heat map, obtain the pixel values of the corresponding pixel points in the structural anomaly heat map and the logical anomaly heat map, multiply them by the same preset weight (for example, 0.1) respectively, and then add them to calculate the pixel value of this pixel point in the anomaly heat map.

[0148] In the above technical solution, determining the structural anomaly heat map according to the teacher feature of the image to be detected and the first student feature of the image to be detected; determining the logical anomaly heat map according to the second student feature of the image to be detected and the encoder feature of the image to be detected; determining the anomaly heat map of the image to be detected according to the structural anomaly heat map of the image to be detected and the logical anomaly heat map of the image to be detected, realizes determining the structural anomaly by using the difference in the ability of the teacher network and the student network to extract abnormal regions, thereby determining the structural anomaly heat map, and determining the logical anomaly by the difference between the encoder network and the student network in the reconstruction result of the image to be detected, thereby determining the logical anomaly heat map, which is applied to industrial scenarios and improves the accuracy and robustness of image anomaly detection in industrial scenarios.

[0149] In the above technical solution, inputting the image to be detected into the teacher network to obtain the teacher feature of the image to be detected, inputting the image to be detected into the student network to obtain the first student feature and the second student feature of the image to be detected, inputting the image to be detected into the encoder network to obtain the encoder feature of the image to be detected, determining the anomaly heat map of the image to be detected according to the teacher feature, the first student feature, the second student feature and the encoder feature, and further determining the anomaly detection result of the image to be detected according to the anomaly heat map of the image to be detected. Combining the teacher network, the student network and the encoder network realizes the anomaly detection of the image to be detected, which is applied to industrial scenarios and improves the accuracy of image anomaly detection in industrial scenarios.

[0150] In some embodiments of the present application, the determining the anomaly detection result of the image to be detected according to the anomaly heat map of the image to be detected includes:

[0151] Determine the anomaly score of the image to be detected based on the anomaly heat map of the image to be detected;

[0152] Compare the anomaly score of the image to be detected with a preset threshold to determine whether there is an anomaly in the image to be detected;

[0153] And / or, determine the binary map of the image to be detected based on the anomaly heat map of the image to be detected, and the binary map is used to display the location of the anomaly area in the image to be detected.

[0154] Figure 5 It is a schematic flow chart of determining the anomaly detection result provided by some embodiments of the present application. As Figure 5 shown, based on the anomaly heat map of the image to be detected, determine the anomaly score of the image to be detected. The anomaly score can be the average value, maximum value or weighted sum of all pixel values in the anomaly heat map, etc. The larger the anomaly score, the greater the possibility that there is an anomaly in the image to be detected.

[0155] Compare the anomaly score with a preset threshold. If the anomaly score is greater than or equal to the preset threshold, it is determined that there is an anomaly (NG) in the image to be detected; if the anomaly score is less than the preset threshold, it is determined that there is no anomaly (OK) in the image to be detected.

[0156] Based on the anomaly heat map of the image to be detected, a binary map of the image to be detected can be further generated to more intuitively display the anomaly area in the image to be detected. The binary map is obtained by comparing the pixel values on the heat map with a threshold. The area where the pixel value is higher than the threshold is marked as an anomaly (usually represented by 1), and the area where the pixel value is lower than the threshold is marked as normal (usually represented by 0).

[0157] In the above technical solution, based on the anomaly heat map of the image to be detected, determine the anomaly score of the image to be detected, compare the anomaly score of the image to be detected with a preset threshold to determine whether there is an anomaly in the image to be detected, and / or, based on the anomaly heat map of the image to be detected, determine the binary map of the image to be detected, and the binary map is used to display the location of the anomaly area in the image to be detected, so as to determine the anomaly detection result of the image to be detected, thereby realizing the anomaly detection of the image to be detected.

[0158] For the image anomaly detection method for industrial scenarios provided by the embodiments of the present application, the execution subject can be an image anomaly detection device for industrial scenarios. In the embodiments of the present application, taking the image anomaly detection device for industrial scenarios executing the image anomaly detection method for industrial scenarios as an example, the image anomaly detection device for industrial scenarios provided by the embodiments of the present application is described.

[0159] Figure 6It is a schematic structural diagram of an image anomaly detection device for industrial scenarios provided by some embodiments of the present application.

[0160] As Figure 6 shown, the image anomaly detection device 600 for industrial scenarios includes:

[0161] An image acquisition unit 601, configured to acquire an image to be detected in an industrial scenario;

[0162] An anomaly detection unit 602, configured to input the image to be detected into an image anomaly detection model to determine an anomaly heat map of the image to be detected, where the image anomaly detection model is trained based on normal image samples;

[0163] An anomaly determination unit 603, configured to determine an anomaly detection result of the image to be detected according to the anomaly heat map of the image to be detected.

[0164] Optionally, the image anomaly detection model is trained through the following steps:

[0165] Construct a training sample set, where the training sample set includes normal image samples;

[0166] Construct an image anomaly detection network, where the image anomaly detection network includes: a teacher network, a student network, and an encoder network. The number of channels of the student network is twice that of the teacher network, and the output of the teacher network has the same scale as the output of the encoder network;

[0167] Input the normal image samples into the teacher network to obtain teacher features of the normal image samples;

[0168] Input the normal image samples into the student network to obtain first student features and second student features of the normal image samples. The first student features are the output of the first half of the channels of the student network, and the second student features are the output of the second half of the channels of the student network;

[0169] Input the normal image samples into the encoder network to obtain encoder features of the normal image samples;

[0170] Calculate a first loss based on the teacher features and the first student features, calculate a second loss based on the second student features and the encoder features, and calculate a third loss based on the teacher features and the encoder features;

[0171] Freeze the parameters of the teacher network, optimize the parameters of the student network and the encoder network based on the first loss, second loss, and third loss. When the training end condition is met, retain the parameters of the student network and the encoder network to obtain the image anomaly detection model.

[0172] Optionally, the teacher network and the student network are obtained by loading a pre-trained lightweight feature extraction network, and the pre-trained lightweight feature extraction network is obtained through the following steps:

[0173] Construct a lightweight feature extraction network;

[0174] Transfer the knowledge of the deep convolutional neural network to the lightweight feature extraction network through knowledge distillation technology to obtain the pre-trained lightweight feature extraction network.

[0175] Optionally, the image anomaly detection model includes: a teacher network, a student network, and an encoder network.

[0176] The teacher network is used to extract features from the image to be detected to obtain the teacher features of the image to be detected.

[0177] The student network is used to extract features from the image to be detected to obtain the first student features and the second student features of the image to be detected. The first student features are the output of the first half of the channels of the student network, and the second student features are the output of the second half of the channels of the student network.

[0178] The encoder network is used to reconstruct the image to be detected to obtain the encoder features of the image to be detected.

[0179] Optionally, the anomaly detection unit 602 is used to:

[0180] Input the image to be detected into the teacher network to obtain the teacher features of the image to be detected;

[0181] Input the image to be detected into the student network to obtain the first student features and the second student features of the image to be detected;

[0182] Input the image to be detected into the encoder network to obtain the encoder features of the image to be detected;

[0183] Determine the anomaly heat map of the image to be detected according to the teacher features, the first student features, the second student features, and the encoder features.

[0184] Optionally, the determining the anomaly heat map of the image to be detected according to the teacher features, the first student features, the second student features, and the encoder features includes:

[0185] Determine the structural anomaly heat map of the image to be detected according to the teacher features of the image to be detected and the first student features of the image to be detected;

[0186] Determine the logical anomaly heat map of the image to be detected according to the second student feature of the image to be detected and the encoder feature of the image to be detected;

[0187] Determine the anomaly heat map of the image to be detected according to the structural anomaly heat map of the image to be detected and the logical anomaly heat map of the image to be detected.

[0188] Optionally, the anomaly determination unit 603 is configured to:

[0189] Determine the anomaly score of the image to be detected based on the anomaly heat map of the image to be detected;

[0190] Compare the anomaly score of the image to be detected with a preset threshold to determine whether there is an anomaly in the image to be detected;

[0191] And / or, determine the binary map of the image to be detected based on the anomaly heat map of the image to be detected, and the binary map is used to display the position of the anomaly area in the image to be detected.

[0192] In the above technical solution, the image anomaly detection device for industrial scenarios obtains the image to be detected in the industrial scenario, inputs the image to be detected into the image anomaly detection model trained based on normal image samples, obtains the anomaly heat map of the image to be detected, and according to the anomaly heat map of the image to be detected, the anomaly detection result of the image to be detected can be determined, realizing the anomaly detection of the image to be detected, applied to industrial scenarios, and improving the accuracy and efficiency of image anomaly detection in industrial scenarios.

[0193] The image anomaly detection device for industrial scenarios in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0194] The image anomaly detection device for industrial scenarios in the embodiments of the present application can be a device with an operating system. The operating system can be the Microsoft (Windows) operating system, the Android operating system, the IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0195] The image anomaly detection device for industrial scenarios provided in the embodiments of the present application can implement each process implemented by the embodiments of the image anomaly detection method for industrial scenarios. To avoid repetition, it will not be elaborated here.

[0196] In some embodiments, as Figure 7 shown, the embodiments of the present application further provide an electronic device 700, including a processor 701, a memory 702, and a computer program stored on the memory 702 and executable on the processor 701. When the program is executed by the processor 701, it implements each process of the above-mentioned embodiments of the image anomaly detection method for industrial scenarios and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0197] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0198] The embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the image anomaly detection method for industrial scenarios and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0199] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.

[0200] The embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the above-mentioned image anomaly detection method for industrial scenarios.

[0201] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.

[0202] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the image anomaly detection method for industrial scenarios and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0203] It should be understood that the chip mentioned in the embodiment of the present application can also be called a system-level chip, system chip, chip system or system-on-chip, etc.

[0204] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0206] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

[0207] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0208] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present application. The scope of the present application is defined by the claims and their equivalents.

Claims

1. An image anomaly detection method for industrial scenes, characterized in that: include: Obtain images to be detected in industrial scenarios; Inputting the image to be detected into an image anomaly detection model to determine an abnormal heat map of the image to be detected, wherein the image anomaly detection model is trained based on normal image samples; An abnormality detection result of the image to be detected is determined according to the abnormality heat map of the image to be detected.

2. The image anomaly detection method for industrial scenes according to claim 1, characterized in that: The image anomaly detection model is trained by the following steps: Constructing a training sample set, wherein the training sample set includes normal image samples; Constructing an image anomaly detection network, the image anomaly detection network comprising: a teacher network, a student network and an encoder network, the number of channels of the student network is twice the number of channels of the teacher network, and the output of the teacher network has the same scale as the output of the encoder network; Inputting the normal image sample into the teacher network to obtain the teacher features of the normal image sample; Input the normal image sample into the student network to obtain a first student feature and a second student feature of the normal image sample, wherein the first student feature is the first half channel output of the student network, and the second student feature is the second half channel output of the student network; Inputting the normal image sample into the encoder network to obtain encoder features of the normal image sample; A first loss is calculated based on the teacher feature and the first student feature, a second loss is calculated based on the second student feature and the encoder feature, and a third loss is calculated based on the teacher feature and the encoder feature; Freeze the parameters of the teacher network, optimize the parameters of the student network and the encoder network based on the first loss, the second loss and the third loss, and retain the parameters of the student network and the encoder network when the training end condition is met to obtain the image anomaly detection model.

3. The image anomaly detection method for industrial scenes according to claim 2, characterized in that: The teacher network and the student network are obtained by loading a pre-trained lightweight feature extraction network, and the pre-trained lightweight feature extraction network is obtained by the following steps: Build a lightweight feature extraction network; The knowledge of the deep convolutional neural network is transferred to the lightweight feature extraction network through the knowledge distillation technology to obtain a pre-trained lightweight feature extraction network.

4. The image anomaly detection method for industrial scenes according to claim 1, characterized in that: The image anomaly detection model includes: a teacher network, a student network and an encoder network. The teacher network is used to extract features of the image to be detected to obtain teacher features of the image to be detected; The student network is used to extract features of the image to be detected, and obtain a first student feature and a second student feature of the image to be detected, wherein the first student feature is the first half channel output of the student network, and the second student feature is the second half channel output of the student network; The encoder network is used to reconstruct the image to be detected to obtain encoder features of the image to be detected.

5. The image anomaly detection method for industrial scenes according to claim 4, characterized in that: The step of inputting the image to be detected into an image anomaly detection model to determine an abnormality heat map of the image to be detected includes: Inputting the image to be detected into the teacher network to obtain the teacher features of the image to be detected; Inputting the image to be detected into the student network to obtain a first student feature and a second student feature of the image to be detected; Inputting the image to be detected into the encoder network to obtain encoder features of the image to be detected; An abnormal heat map of the image to be detected is determined according to the teacher feature, the first student feature, the second student feature and the encoder feature.

6. The image anomaly detection method for industrial scenes according to claim 5, characterized in that: The step of determining the abnormal heat map of the image to be detected according to the teacher feature, the first student feature, the second student feature and the encoder feature includes: Determining a structural anomaly thermodynamic map of the image to be detected according to the teacher feature of the image to be detected and the first student feature of the image to be detected; Determining a logic anomaly heat map of the image to be detected according to the second student feature of the image to be detected and the encoder feature of the image to be detected; The abnormality thermogram of the image to be detected is determined according to the structural abnormality thermogram of the image to be detected and the logical abnormality thermogram of the image to be detected.

7. The image anomaly detection method for industrial scenes according to any one of claims 1 to 6, characterized in that: The step of determining the abnormality detection result of the image to be detected according to the abnormality heat map of the image to be detected includes: Determining an anomaly score of the image to be detected based on the anomaly heat map of the image to be detected; Comparing the abnormality score of the image to be detected with a preset threshold to determine whether the image to be detected has an abnormality; And / or, based on the abnormal thermal map of the image to be detected, a binarization map of the image to be detected is determined, and the binarization map is used to display the position of the abnormal area in the image to be detected.

8. An image anomaly detection device for industrial scenes, characterized in that: include: An image acquisition unit, used to acquire an image to be detected in an industrial scene; An anomaly detection unit, used for inputting the image to be detected into an image anomaly detection model to determine an abnormal heat map of the image to be detected, wherein the image anomaly detection model is obtained by training based on normal image samples; The abnormality determination unit is used to determine the abnormality detection result of the image to be detected according to the abnormality heat map of the image to be detected.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the image anomaly detection method for industrial scenarios as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image anomaly detection method for industrial scenarios as described in any one of claims 1 to 7 is implemented.