Image Processing Method, Apparatus, Device, Storage Medium and Computer Program Product

By performing feature extraction, encoding and decoding on the image, and using a discriminator based on synthetic image samples to detect abnormal pixels in the image, the problem of low accuracy in image abnormality detection in the prior art is solved, and high-accuracy pixel-level abnormality detection is achieved.

CN117710301BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311692164.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10
Estimated Expiration
2043-12-08

AI Technical Summary

Technical Problem

Existing machine learning models based on unsupervised learning have low accuracy in image anomaly detection.

Method used

By performing feature extraction, encoding and decoding processing on the target image, image reconstruction features are acquired, and the difference between them and image extraction features is input into the discriminator to obtain probability information indicating an abnormality in the pixel. The discriminator is trained based on synthetic image samples.

Benefits of technology

The accuracy of image abnormality detection is improved, pixel-level abnormality detection is realized, and the number of training samples is expanded through unsupervised learning and the accuracy of detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117710301B_ABST
    Figure CN117710301B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus, device, storage medium and computer program product, relating to the field of artificial intelligence technology. The method includes: extracting features from a target image to obtain image extraction features; performing encoding and decoding processing on the image extraction features to obtain image reconstruction features; inputting the difference between the image reconstruction features and the image extraction features into a discriminator to obtain first recognition information output by the discriminator, where the first recognition information includes first information elements, and the first information elements are used to indicate the probability that the corresponding pixels are abnormal; and obtaining a recognition result for abnormal recognition of the target image based on the image extraction features, the image reconstruction features, and the first recognition information. The above solution combines the image features before and after reconstruction, as well as the discrimination result of the difference between the image features before and after reconstruction, to comprehensively determine the abnormality of each pixel in the target image, and can achieve pixel-level image anomaly detection based on AI technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, device, storage medium and computer program product. Background Art

[0002] Image anomaly detection refers to the technology of detecting whether there are abnormal or defective pixels in an image. Image anomaly detection can be achieved through artificial intelligence.

[0003] In related technologies, image anomaly detection can be achieved through artificial intelligence. For example, a machine learning model is trained in advance using normal images in an unsupervised learning manner. It is expected that the machine learning model will have a larger reconstruction error on abnormal images, thereby achieving anomaly detection.

[0004] However, in the above-mentioned related technologies, the accuracy of machine learning models based on unsupervised learning for image anomaly detection is not high. Summary of the invention

[0005] The embodiments of the present application provide an image processing method, apparatus, device, storage medium and computer program product, which can improve the accuracy of image anomaly detection. The technical solution is as follows:

[0006] In one aspect, an image processing method is provided, the method comprising:

[0007] Performing feature extraction on a target image to obtain image extraction features of the target image;

[0008] Encoding and decoding the image extraction features to obtain image reconstruction features;

[0009] Inputting the difference between the image reconstruction feature and the image extraction feature into a discriminator, obtaining first identification information output by the discriminator, wherein the first identification information includes first information elements corresponding to each pixel in the target image, and the first information elements are used to indicate the probability that the corresponding pixel has an abnormality; the discriminator is a machine learning model obtained by training based on synthetic image samples, and the synthetic image samples are images with abnormalities synthesized based on images without abnormalities;

[0010] Based on the image extraction features, the image reconstruction features and the first recognition information, a recognition result of abnormality recognition of the target image is obtained.

[0011] In one aspect, an image processing method is provided, the method comprising:

[0012] Based on the first image sample, first mask information, second mask information and a second image sample are acquired; the first mask information is used to indicate that there are no abnormal pixels in the first image sample, the second mask information is used to indicate that there are abnormal pixels in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information;

[0013] Inputting the first image sample into a feature extractor in an image anomaly detection model to obtain a first image extracted feature sample output by the feature extractor;

[0014] Inputting the first image extracted feature sample into a codec in the image anomaly detection model to obtain a first image reconstructed feature sample output by the codec;

[0015] Inputting the difference between the first image reconstructed feature sample and the first image extracted feature sample into the discriminator in the image anomaly detection model to obtain a first identification information sample output by the discriminator;

[0016] Inputting the second image sample into the feature extractor to obtain a second image extracted feature sample output by the feature extractor;

[0017] Inputting the second image extraction feature sample into the codec to obtain the second image reconstruction feature sample output by the codec;

[0018] Inputting the difference between the second image reconstructed feature sample and the second image extracted feature sample into the discriminator to obtain a second identification information sample output by the discriminator;

[0019] Obtaining a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first identification information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second identification information sample;

[0020] Updating the parameters of the image anomaly detection model according to the loss function value;

[0021] Among them, the image anomaly detection model is used to extract features of the target image, obtain image extraction features of the target image, encode and decode the image extraction features to obtain image reconstruction features, perform discrimination processing based on the difference between the image reconstruction features and the image extraction features, and obtain first identification information, wherein the first identification information includes first information elements corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel has an abnormality; the image extraction features, the image reconstruction features and the first identification information are used to obtain the recognition result of abnormality recognition of the target image.

[0022] In another aspect, an image processing device is provided, the device comprising:

[0023] A feature extraction module is used to extract features from a target image to obtain image extraction features of the target image;

[0024] A coding and decoding module, used for encoding and decoding the image extraction features to obtain image reconstruction features;

[0025] a reconstruction module, configured to input the difference between the image reconstruction feature and the image extraction feature into a discriminator, and obtain first identification information output by the discriminator, wherein the first identification information includes first information elements corresponding to each pixel in the target image, and the first information elements are used to indicate the probability that the corresponding pixel has an abnormality; the discriminator is a machine learning model obtained by training based on synthetic image samples, and the synthetic image samples are images with abnormalities synthesized based on images without abnormalities;

[0026] The recognition module is used to obtain a recognition result of abnormality recognition of the target image based on the image extraction feature, the image reconstruction feature and the first recognition information.

[0027] In a possible implementation, the identification module is used to:

[0028] Based on the difference between the image extraction feature and the image reconstruction feature, second identification information is obtained; the second identification information includes second information elements corresponding to each pixel in the target image, and the second information elements are used to indicate the probability that the corresponding pixel is abnormal;

[0029] Based on the first identification information and the second identification information, segmentation information in the identification result is obtained; the segmentation information includes third information elements corresponding to each pixel in the target image, and the third information elements are used to indicate the probability that the corresponding pixel has an abnormality.

[0030] In a possible implementation, the dimension of the difference between the image extraction feature and the image reconstruction feature is smaller than the number of pixels of the target image; the recognition module is used to take the L2 norm of the difference between the image extraction feature and the image reconstruction feature, and upsample the L2 norm to obtain the second recognition information.

[0031] In a possible implementation manner, the identification module is used to perform a weighted sum or a weighted average on the first identification information and the second identification information to obtain the segmentation information.

[0032] In a possible implementation, the recognition module is further used to obtain an image classification result in the recognition result based on the segmentation information, and the image classification result is used to indicate the probability that the target image is an abnormal image.

[0033] In a possible implementation manner, the recognition module is used to take the maximum value or standard deviation of each third information element in the segmentation information to obtain the image classification result.

[0034] In a possible implementation, the feature extraction module is used to input the target image into a feature extractor in an image anomaly detection model to obtain the image extraction features output by the feature extractor;

[0035] The codec module is used to input the image extraction feature into the codec in the image anomaly detection model to obtain the image reconstruction feature output by the codec, wherein the size of the image reconstruction feature is the same as the size of the image extraction feature;

[0036] The reconstruction module is used to input the difference between the image reconstruction feature and the image extraction feature into the discriminator in the image anomaly detection model to obtain the first recognition information output by the discriminator;

[0037] Among them, the image anomaly detection model is a machine learning model obtained by unsupervised training based on a first image sample; the first image sample is an image without anomalies.

[0038] In a possible implementation manner, the device further includes:

[0039] a sample processing module, configured to obtain first mask information, second mask information, and a second image sample based on the first image sample before the feature extraction module extracts features from the target image and obtains image extraction features of the target image; the first mask information is used to indicate that there are no abnormal pixels in the first image sample, the second mask information is used to indicate that there are abnormal pixels in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information;

[0040] A model processing module, used for inputting the first image sample into the feature extractor to obtain a first image extracted feature sample output by the feature extractor; inputting the first image extracted feature sample into the codec to obtain a first image reconstructed feature sample output by the codec; inputting the difference between the first image reconstructed feature sample and the first image extracted feature sample into the discriminator to obtain a first identification information sample output by the discriminator; inputting the second image sample into the feature extractor to obtain a second image extracted feature sample output by the feature extractor; inputting the second image extracted feature sample into the codec to obtain a second image reconstructed feature sample output by the codec; inputting the difference between the second image reconstructed feature sample and the second image extracted feature sample into the discriminator to obtain a second identification information sample output by the discriminator;

[0041] A loss acquisition module, used to acquire a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first identification information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second identification information sample;

[0042] A parameter updating module is used to update the parameters of the image anomaly detection model through the loss function value.

[0043] In a possible implementation, the loss acquisition module is used to:

[0044] Obtaining a first loss function value based on a difference between the first image extracted feature sample and the first image reconstructed feature sample;

[0045] Obtaining a second loss function value based on a difference between the first identification information sample and the first mask information;

[0046] Obtaining a third loss function value based on a difference between the second image extracted feature sample and the second image reconstructed feature sample;

[0047] Based on the difference between the second identification information sample and the second mask information, a fourth loss function value is obtained.

[0048] In another aspect, an image processing device is provided, the device comprising:

[0049] a sample processing module, configured to obtain first mask information, second mask information, and a second image sample based on a first image sample; the first mask information is used to indicate that there are no abnormal pixels in the first image sample, the second mask information is used to indicate that there are abnormal pixels in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information;

[0050] a model processing module, configured to input the first image sample into a feature extractor in an image anomaly detection model to obtain a first image extracted feature sample output by the feature extractor; input the first image extracted feature sample into a codec in the image anomaly detection model to obtain a first image reconstructed feature sample output by the codec; input the difference between the first image reconstructed feature sample and the first image extracted feature sample into a discriminator in the image anomaly detection model to obtain a first identification information sample output by the discriminator; input the second image sample into the feature extractor to obtain a second image extracted feature sample output by the feature extractor; input the second image extracted feature sample into the codec to obtain a second image reconstructed feature sample output by the codec; input the difference between the second image reconstructed feature sample and the second image extracted feature sample into the discriminator to obtain a second identification information sample output by the discriminator;

[0051] A loss acquisition module, used to acquire a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first identification information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second identification information sample;

[0052] A parameter updating module, used to update the parameters of the image anomaly detection model according to the loss function value;

[0053] Among them, the image anomaly detection model is used to extract features of the target image, obtain image extraction features of the target image, encode and decode the image extraction features to obtain image reconstruction features, perform discrimination processing based on the difference between the image reconstruction features and the image extraction features, and obtain first identification information, wherein the first identification information includes first information elements corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel has an abnormality; the image extraction features, the image reconstruction features and the first identification information are used to obtain the recognition result of abnormality recognition of the target image.

[0054] In a possible implementation, the loss acquisition module is used to:

[0055] Obtaining a first loss function value based on a difference between the first image extracted feature sample and the first image reconstructed feature sample;

[0056] Obtaining a second loss function value based on a difference between the first identification information sample and the first mask information;

[0057] Obtaining a third loss function value based on a difference between the second image extracted feature sample and the second image reconstructed feature sample;

[0058] Based on the difference between the second identification information sample and the second mask information, a fourth loss function value is obtained.

[0059] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the image processing method as described in the above-mentioned embodiments of the present application.

[0060] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the image processing method as described in the above-mentioned embodiments of the present application.

[0061] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image processing method described in the above embodiment.

[0062] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:

[0063] The scheme shown in the embodiment of the present application, after extracting the image extraction features from the target image, reconstructs the image extraction features by means of encoding and decoding to obtain the image reconstruction features, further uses a discriminator to perform abnormality discrimination on the difference between the image reconstruction features and the image extraction features to obtain first recognition information indicating the probability of whether each pixel in the target image has an abnormality, and finally, combines the image features before and after reconstruction and the first recognition information to comprehensively obtain the recognition result of abnormal recognition of the target image. In addition to using the image features before and after reconstruction to identify the abnormality of the pixels in the target image, the above scheme also uses a discriminator to discriminate the abnormality of the pixels in the target image by using the difference between the image features before and after reconstruction. On the one hand, combining the image features before and after reconstruction, , as well as the discrimination results of the differences between the image features before and after reconstruction, comprehensively determine the abnormalities of each pixel in the target image, so that pixel-level image anomaly detection can be achieved while ensuring the accuracy of image anomaly detection. On the other hand, the above-mentioned discriminator is trained by synthetic image samples synthesized from normal images. Therefore, the training process of the discriminator does not require manual collection and annotation of abnormal images. It can be trained through unsupervised learning, thereby expanding the number of samples available for discriminator training, making up for the difference between the effect of training the discriminator through synthetic image samples and the effect of training the discriminator through real abnormal images, ensuring the accuracy of the discriminator, and then ensuring the accuracy of detecting abnormal pixels in the image through the discriminator. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0065] Figure 1 is a system configuration diagram of an image processing system involved in the present application;

[0066] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment;

[0067] Figure 3 is a flowchart of an image processing method according to an exemplary embodiment;

[0068] Figure 4 is a flowchart of an image processing method according to an exemplary embodiment;

[0069] Figure 5is a flowchart of an image processing method according to an exemplary embodiment;

[0070] Figure 6 This is the unsupervised + discriminant reconstruction network architecture diagram involved in this application;

[0071] Figure 7 It is a test result comparison chart involved in this application;

[0072] Figure 8 It is a schematic diagram of the abnormality identification comparison involved in this application;

[0073] Fig. 9 is a structural block diagram of an image processing device provided by an exemplary embodiment of the present application;

[0074] Fig.10 is a structural block diagram of an image processing device provided by an exemplary embodiment of the present application;

[0075] Fig.11 It is a schematic diagram of the structure of a server provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0076] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0077] Before describing the various embodiments shown in the present application, several concepts involved in the present application are first introduced.

[0078] 1) AI (Artificial Intelligence)

[0079] AI is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level technology and software-level technology. Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0080] 2) Computer Vision (CV)

[0081] Computer vision is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision that uses cameras and computers to replace human eyes to identify, detect and measure targets, and further performs graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, virtual reality, augmented reality, simultaneous positioning and mapping, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0082] 3) ML (Machine Learning)

[0083] Machine learning is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0084] 4) Cloud technology

[0085] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0086] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, image websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, which needs to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.

[0087] 5) Artificial Intelligence Cloud Services

[0088] The so-called artificial intelligence cloud service is generally also called AIaaS (AI as a Service). This is the service mode of a mainstream artificial intelligence platform. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI theme mall: all developers can access and use one or more artificial intelligence services provided by the platform through the API interface. Some senior developers can also use the AI ​​framework and AI infrastructure provided by the platform to deploy and operate their own cloud artificial intelligence services.

[0089] Please refer to Figure 1 , which shows a system structure diagram of an image processing system involved in various embodiments of the present application. Figure 1 As shown, the system includes an image acquisition device 120 , a terminal 140 , and a server 160 ; optionally, the system may also include a database 180 .

[0090] The image acquisition device 120 may be a device for acquiring images and may have a built-in or external camera component.

[0091] The image acquisition device 120 may include an image output interface, such as a Universal Serial Bus (USB) interface, a High Definition Multimedia Interface (HDMI) interface, or an Ethernet interface, etc.; or, the above-mentioned image output interface may also be a wireless interface, such as a Wireless Local Area Network (WLAN) interface, a Bluetooth interface, etc.

[0092] Correspondingly, depending on the type of the above-mentioned image output interface, the operator may have a variety of ways to export the image, for example, importing the image to the terminal 140 via a wired or short-distance wireless method, or importing the image to the terminal 140 or the server 160 via a local area network or the Internet.

[0093] The terminal 140 may be a terminal device with certain processing capabilities and interface display functions, for example, the terminal 140 may be a mobile phone, a tablet computer, an e-book reader, smart glasses, a laptop computer, a desktop computer, and the like.

[0094] The terminal 140 may be a terminal used by a user or a terminal used by a developer.

[0095] When terminal 140 is implemented as a terminal used by developers, the developers can develop a machine learning model for image processing through terminal 140 and deploy the machine learning model to server 160 or a terminal used by a user.

[0096] When the terminal 140 is implemented as a terminal for use by users, an application that acquires images for abnormality recognition processing and presents recognition results can be installed in the terminal 140. The application can have built-in or call the above-mentioned machine learning model for image processing. After the terminal 140 acquires the image captured by the image acquisition device 120, it can obtain the abnormality recognition result obtained by processing the image through the above-mentioned application, and present the abnormality recognition result for the user's reference.

[0097] exist Figure 1 In the system shown, the terminal 140 and the image acquisition device 120 are physically separate entity devices. Optionally, in another possible implementation, when the terminal 140 is implemented as a terminal used by a user, the terminal 140 and the image acquisition device 120 may also be integrated into a single entity device; for example, the terminal 140 may be a smart phone with a built-in camera.

[0098] Among them, server 160 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.

[0099] Among them, the above-mentioned server 160 can be a server that provides background services for the application installed in the terminal 140. The background server can be responsible for version management of the application, background processing of images acquired by the application and returning processing results, background training of machine learning models developed by developers, etc.

[0100] The database 180 may be a Redis database, or may be another type of database. The database 180 is used to store various types of data.

[0101] Optionally, the terminal 140 is connected to the server 160 via a communication network. Optionally, the image acquisition device 120 is connected to the server 160 via a communication network. Optionally, the communication network is a wired network or a wireless network.

[0102] Optionally, the system may also include a management device ( Figure 1 (not shown), the management device is connected to the server 160 via a communication network. Optionally, the communication network is a wired network or a wireless network.

[0103] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment. The method can be executed by a computer device, for example, the computer device can be a server, or the computer device can be a terminal, or the computer device can include a server and a terminal, wherein the server can be the above-mentioned Figure 1 In the embodiment shown, the server 160, the terminal may be the above-mentioned Figure 1 In the embodiment shown, the terminal 140 is used by the user. The computer device can be implemented as a model application device for image anomaly recognition. Figure 2 As shown, the image processing method may include the following steps.

[0104] Step 210: extract features from the target image to obtain image extraction features of the target image.

[0105] In an embodiment of the present application, the computer device may perform convolution, pooling and other processing on the target image to extract image extraction features from the target image.

[0106] Among them, when the computer device extracts features of the target image, it can extract image features of multiple different scales (that is, the number of feature dimensions), and fuse the image features of multiple different scales to obtain image extraction features.

[0107] For example, a computer device can perform multiple levels of feature extraction on a target image. When extracting features at the first level, convolution and other operations are performed on the input target image to obtain image features at the first level. When extracting features at subsequent levels, convolution and other operations are performed on image features at the previous one or more levels. The dimension of the image features at the previous level can be greater than or equal to the dimension of the image features at the subsequent level.

[0108] Optionally, if the dimension of the image features at the previous level is larger than the dimension of the image features at the subsequent level, the computer device may unify the dimensions of the image features at multiple levels when fusing image features of multiple different scales, for example, unify them into the dimensions of the image features of the last level to obtain multiple image features of the same dimension, and then concatenate the multiple image features of the same dimension to obtain image extraction features.

[0109] Optionally, when fusing multiple image features of different scales, the computer device may also average or weighted average the features of each dimension among the multiple image features of the same dimension to obtain the image extraction feature.

[0110] Step 220: Encode and decode the image extraction features to obtain image reconstruction features.

[0111] In the embodiment of the present application, the above-mentioned process of encoding and decoding the image extraction features may refer to the process of crossing and reconstructing the features of each dimension in the image extraction features.

[0112] For example, a computer device can perform multi-level processing on features of each dimension in image extraction features. In each level of processing, weighting, convolution, pooling and other processing based on the attention mechanism are performed on the image features, and the processed image features are output to the next level to finally obtain image reconstruction features.

[0113] Step 230: Input the difference between the image reconstructed feature and the image extracted feature into the discriminator to obtain the first identification information output by the discriminator, wherein the first identification information includes first information elements corresponding to each pixel in the target image, and the first information elements are used to indicate the probability that the corresponding pixel has an abnormality.

[0114] Among them, the above discriminator is a machine learning model trained based on synthetic image samples, and the synthetic image samples are abnormal images synthesized based on images without abnormalities.

[0115] In the embodiments of the present application, abnormal images can be synthesized in advance from images without abnormalities, and then a discriminator is trained based on the abnormal images. The input of the discriminator is the difference between the image reconstruction feature and the image extraction feature, and the output of the discriminator is information indicating the probability of abnormality of each pixel in the image.

[0116] For example, the above target image may include H×W pixels, and the first recognition information includes H×W first information elements, and the H×W first information elements correspond one-to-one to the H×W pixels.

[0117] Among them, the above first information element may be a probability value, and the probability value represents the probability of abnormality of the corresponding pixel.

[0118] Alternatively, the above first information element may be an anomaly score, and the anomaly score may be positively or negatively correlated with the probability of abnormality of the corresponding pixel.

[0119] For example, taking the above first information element as the anomaly score as an example, the above first recognition information may be as shown in Table 1 below.

[0120] Table 1

[0121] 1 3 2 1 4 1 0 1 … 0 1 1 2 1 2 0 1 2 … 1 0 2 1 4 70 2 2 2 … 1 0 1 0 85 92 88 1 0 … 0 0 3 1 1 81 3 0 1 … 3 1 4 0 2 1 1 3 3 … 3 … … … … … … … … … … 2 2 2 1 1 0 1 1 … 2

[0122] As shown in Table 1 above, each blank corresponds to an anomaly score of a first information element, and the value range of the anomaly score is [0, 100]. Among them, the higher the anomaly score of the first information element, the higher the probability of abnormality of the corresponding pixel.

[0123] Step 240: Obtain an identification result for abnormal identification of the target image based on the image extraction feature, the image reconstruction feature, and the first recognition information.

[0124] In the embodiments of the present application, the computer device can determine the probability of abnormality of each pixel in the target image based on the difference between the image extraction feature and the image reconstruction feature, and then, in combination with the first recognition information, comprehensively obtain an identification result for abnormal identification of the target image.

[0125] In summary, in the solution shown in the embodiments of the present application, after extracting image extraction features from the target image, the image extraction features are reconstructed through an encoding and decoding method to obtain image reconstruction features. Further, a discriminator is used to perform anomaly discrimination on the difference between the image reconstruction features and the image extraction features to obtain first recognition information indicating the probability of whether each pixel in the target image is abnormal. Finally, by combining the image features before and after reconstruction and the first recognition information, the recognition result of anomaly recognition of the target image is comprehensively obtained. In addition to using the image features before and after reconstruction to identify the abnormal conditions of the pixels in the target image, the above solution also uses the discriminator to discriminate the abnormal conditions of the pixels in the target image based on the difference between the image features before and after reconstruction. On the one hand, by combining the image features before and after reconstruction and the discrimination result of the difference between the image features before and after reconstruction, the abnormal conditions of each pixel in the target image are comprehensively determined, so that pixel-level image anomaly detection can be realized while ensuring the accuracy of image anomaly detection. On the other hand, the above discriminator is trained by using the synthetic image samples obtained by synthesizing normal images. Therefore, the training process of the discriminator does not require manual collection and annotation of abnormal images, and it can be trained in an unsupervised learning manner, so that the number of available samples for training the discriminator can be expanded, and the difference between the training effect of the discriminator by using synthetic image samples and the training effect of the discriminator by using real abnormal images can be compensated, ensuring the accuracy of the discriminator, and then ensuring the accuracy of detecting abnormal pixels in the image by using the discriminator.

[0126] Based on Figure 2 the embodiments shown, please refer to Figure 3 , which is a schematic flowchart of an image processing method shown according to an exemplary embodiment. As Figure 3 shown, the above step 240 can be implemented as step 240a and step 240b.

[0127] Step 240a: Obtain second recognition information based on the difference between the image extraction features and the image reconstruction features; the second recognition information includes second information elements respectively corresponding to each pixel in the target image, and the second information element is used to indicate the probability that the corresponding pixel is abnormal.

[0128] For example, the second recognition information includes H×W second information elements. Among the H×W second information elements, each second information element corresponds to one pixel among the H×W pixels. That is to say, the second information elements in the second recognition information are in one-to-one correspondence with the pixels in the target image.

[0129] Among them, the above second information element can be a probability value, and the probability value represents the probability that the corresponding pixel is abnormal.

[0130] Alternatively, the above second information element may be an anomaly score, which may be positively or negatively correlated with the probability that the corresponding pixel is anomalous.

[0131] In some embodiments, the dimension of the difference between the above image extraction features and the image reconstruction features is less than the number of pixels of the target image; obtaining the second recognition information based on the difference between the image extraction features and the image reconstruction features includes: taking the L2 norm of the difference between the image extraction features and the image reconstruction features, and upsampling the L2 norm to obtain the second recognition information.

[0132] In the embodiments of the present application, the computer device may subtract the image extraction features and the image reconstruction features element by element to obtain the difference between the image extraction features and the image reconstruction features, and then take the L2 norm of the difference between the image extraction features and the image reconstruction features and perform upsampling to obtain the second recognition information.

[0133] For example, the dimensions of the above image extraction features and the image reconstruction features are both h×w, where h×w is less than H×W. For example, h = H / 16 and w = W / 16; wherein, the computer device subtracts the features on the same dimension between the image extraction features and the image reconstruction features and then calculates the L2 norm. The dimension of the L2 norm is also h×w. Then, the calculation result (i.e., the above L2 norm) is subjected to upsampling processing (such as upsampling through bilinear interpolation processing) to obtain the second anomaly information that also contains H×W elements (i.e., the above second information element).

[0134] For example, the dimensions of the above image extraction features and the image reconstruction features are both h×w; the above subtracting the image extraction features and the image reconstruction features element by element means subtracting the first information element in the image extraction features from the first information element in the image reconstruction features, subtracting the second information element in the image extraction features from the second information element in the image reconstruction features, and so on, to obtain h×w subtraction results.

[0135] In the embodiments of the present application, after the computer device performs interaction and reconstruction on the image extraction features, it detects whether each pixel is anomalous through the L2 norm of the error before and after the interaction and reconstruction, ensuring the feasibility of detecting anomalies in the reconstruction error of the image features.

[0136] Step 240b: Obtain the segmentation information in the recognition result based on the first recognition information and the second recognition information; the segmentation information contains third information elements respectively corresponding to each pixel in the target image, and the third information element is used to indicate the probability that the corresponding pixel is anomalous.

[0137] For example, the above segmentation information includes H×W third information elements, and the H×W third information elements correspond one-to-one to the H×W pixels in the target image.

[0138] Among them, the above third information element can be a probability value, and this probability value represents the probability that the corresponding pixel is abnormal.

[0139] Alternatively, the above third information element can be an anomaly score, and this anomaly score can be positively or negatively correlated with the probability that the corresponding pixel is abnormal.

[0140] In the embodiments of the present application, the number of dimensions of the first recognition information and the second recognition information is the same, both being H×W. When obtaining the segmentation information in the recognition result, the third information elements corresponding to each dimension can be obtained to obtain segmentation information that also has H×W elements.

[0141] For example, the above obtaining the third information elements corresponding to each dimension respectively can be to combine the information elements on the first dimension in the first recognition information with the information elements on the first dimension in the second recognition information to obtain the information elements on the first dimension in the segmentation information, and combine the information elements on the second dimension in the first recognition information with the information elements on the second dimension in the second recognition information to obtain the information elements on the second dimension in the segmentation information, and so on, to obtain the H×W information elements in the segmentation information.

[0142] Among them, the computer device comprehensively determines the abnormality of each pixel in the image by combining the reconstruction error of the image features and the result of the discriminator's discrimination of the abnormality of the pixels in the target image based on the difference between the image features before and after reconstruction, ensuring the accuracy of pixel-level anomaly detection.

[0143] In a possible implementation manner, the above obtaining the segmentation information in the recognition result based on the first recognition information and the second recognition information includes:

[0144] Performing weighted summation or weighted averaging on the first recognition information and the second recognition information to obtain the segmentation information.

[0145] For example, the computer device can perform weighted summation or weighted averaging on the information elements on the same dimension in the first recognition information and the second recognition information to obtain the weighted summation or weighted averaging results of H×W dimensions, and the weighted summation or weighted averaging results of these H×W dimensions are the above segmentation information.

[0146] For example, in the above-mentioned first identification information and second identification information, the information elements on the same dimension are weighted and summed or weighted averaged. It can be that the information element on the first dimension in the first identification information is weighted and summed or weighted averaged with the information element on the first dimension in the second identification information to obtain the information element on the first dimension in the segmentation information. The information element on the second dimension in the first identification information is weighted and summed or weighted averaged with the information element on the second dimension in the second identification information to obtain the information element on the second dimension in the segmentation information, and so on, to obtain the H×W information elements in the segmentation information.

[0147] In the embodiments of the present application, the computer device can flexibly set the weights of the first identification information and the second identification information by performing weighted sum or weighted average processing on the information elements of the same dimension in the first identification information and the second identification information, so as to ensure the flexibility and accuracy of the final identification result obtained by integrating the first identification information and the second identification information.

[0148] In a possible implementation manner, the method further includes:

[0149] Based on the segmentation information, obtain the image classification result in the identification result, where the image classification result is used to indicate the probability that the target image belongs to an abnormal image.

[0150] Wherein, the above image classification result can be a probability value, and this probability value represents the probability that the target image belongs to an abnormal image.

[0151] Alternatively, the above image classification result can be an anomaly score, and this anomaly score can be positively or negatively correlated with the probability that the target image belongs to an abnormal image.

[0152] In the embodiments of the present application, the above segmentation information is a pixel-level anomaly recognition result, that is to say, this segmentation information is used to indicate whether there is an anomaly in each pixel of the target image; on this basis, the computer device can also obtain an image-level anomaly recognition result based on this segmentation information, that is to say, the above image classification result is used to represent whether the target image belongs to an abnormal image as a whole, so as to improve the diversity of the image recognition result, and further improve the applicable scenarios of image anomaly recognition.

[0153] In a possible implementation manner, the above-mentioned obtaining the image classification result in the identification result based on the segmentation information includes:

[0154] Take the maximum value or standard deviation of each third information element in the segmentation information to obtain the image classification result.

[0155] In the embodiments of the present application, the computer device takes the maximum value or standard deviation of each element in the segmentation information as the abnormal detection result at the image level, providing an implementable solution for determining the pixel-level abnormal recognition result based on the pixel-level abnormal recognition result, and ensuring the accuracy of the abnormal recognition result at the image level.

[0156] Based on Figure 2 or Figure 3 the embodiments shown, please refer to Figure 4 , which is a schematic flowchart of an image processing method shown according to an exemplary embodiment. As Figure 4 shown, the above steps 210, 220, and 230 can be respectively implemented as steps 210a, 220a, and 230a.

[0157] Step 210a: Input the target image into the feature extractor in the image anomaly detection model to obtain the image extraction features output by the feature extractor.

[0158] For example, the dimension of the image data received at the input end of the feature extractor is 3×H×W, corresponding to H×W pixels in the target image, and each pixel corresponds to data of 3 channels (such as the values of R, G, and B). After the feature extractor processes the image data with a data dimension of 3×H×W, it outputs image extraction features with a dimension of h×w.

[0159] Step 220a: Input the image extraction features into the encoder-decoder in the image anomaly detection model to obtain the image reconstruction features output by the encoder-decoder, and the size of the image reconstruction features is the same as the size of the image extraction features.

[0160] For example, the dimension of the data received at the input end of the above encoder-decoder is h×w. After the encoder-decoder processes the image extraction features with a dimension of h×w, it can output image reconstruction features with a dimension of h×w.

[0161] Step 230a: Input the difference between the image reconstruction features and the image extraction features into the discriminator in the image anomaly detection model to obtain the first recognition information output by the discriminator.

[0162] For example, the dimension of the data received at the input end of the above discriminator is h×w. The discriminator processes the image reconstruction features with a dimension of h×w and can output the first recognition information with a dimension of H×W.

[0163] Among them, the above image anomaly detection model is a machine learning model obtained through unsupervised training based on the first image sample; the first image sample is an image without anomalies.

[0164] In an embodiment of the present application, a computer device processes a target image through a pre-trained image anomaly detection model, thereby ensuring the efficiency and accuracy of image processing.

[0165] Based on Figure 4 the embodiment shown, please refer to Figure 5 , which is a schematic flowchart of an image processing method shown according to an exemplary embodiment. This method can be executed by a computer device. For example, the computer device can be a server, or the computer device can also be a terminal, or the computer device can include a server and a terminal. Among them, the server can be the server 160 in the Figure 1 embodiment shown above, and the terminal can be the terminal 140 used by a developer in the Figure 1 embodiment shown above. The computer device can be implemented as a model training device for model training. As Figure 5 shown, before the above step 210, this method can include the following steps.

[0166] Step 510: Based on a first image sample, obtain first mask information, second mask information, and a second image sample; the first mask information is used to indicate pixels without anomalies in the first image sample, the second mask information is used to indicate pixels with anomalies in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information.

[0167] For example, the computer device can obtain the first mask information pre-set by a developer, or randomly generate the above first mask information. Among them, assuming that the dimension of the image sample is H×W, the above first mask information can be a two-dimensional matrix with a dimension of H×W. Each dimension in this two-dimensional matrix corresponds to a pixel in the image sample, and the value of each dimension in this two-dimensional matrix is used to indicate whether the corresponding pixel has an anomaly. For example, if the value of a certain dimension in this two-dimensional matrix is 1, it means that the corresponding pixel in the image sample has an anomaly. If the value of this dimension is 0, it means that the corresponding pixel in the image sample is normal. Similarly, the above second mask information can also be a two-dimensional matrix with a dimension of H×W, and the values of all dimensions in this two-dimensional matrix are 0.

[0168] In addition, the computer device can also synthesize the second image sample according to the above first mask information and the first image sample. For example, the computer device can modify the values of the pixels indicated by the first mask information as having anomalies in the first image sample. For example, the computer device can randomly replace the values of the pixels indicated by the first mask information as having anomalies in the first image sample to obtain the above second image sample.

[0169] Step 520: Input the first image sample into the feature extractor to obtain the first image extraction feature sample output by the feature extractor; input the first image extraction feature sample into the codec to obtain the first image reconstruction feature sample output by the codec; input the difference between the first image reconstruction feature sample and the first image extraction feature sample into the discriminator to obtain the first recognition information sample output by the discriminator; input the second image sample into the feature extractor to obtain the second image extraction feature sample output by the feature extractor; input the second image extraction feature sample into the codec to obtain the second image reconstruction feature sample output by the codec; input the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain the second recognition information sample output by the discriminator.

[0170] Among them, the process of each branch part (i.e., the above-mentioned feature extractor, codec, and discriminator) in the image anomaly detection model processing the image sample to obtain the recognition information sample is similar to the process of each branch part in the image anomaly detection model processing the target image to obtain the first recognition information, which will not be elaborated here.

[0171] Step 530: Obtain the loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample.

[0172] In the embodiment of the present application, during the model training process, the above-mentioned first mask information and second mask information are equivalent to the annotation information corresponding to the first image sample and the second image sample respectively. Based on this annotation information, the loss function can be calculated for the intermediate data and output results in the process of the image anomaly detection model processing the image sample.

[0173] In a possible implementation manner, the above-mentioned obtaining the loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample includes:

[0174] Obtain the first loss function value based on the difference between the first image extraction feature sample and the first image reconstruction feature sample;

[0175] Obtain the second loss function value based on the difference between the first recognition information sample and the first mask information;

[0176] Obtain the third loss function value based on the difference between the second image extraction feature sample and the second image reconstruction feature sample;

[0177] Obtain the fourth loss function value based on the difference between the second recognition information sample and the second mask information.

[0178] Among them, during the calculation of the loss function, the above first loss function value and third loss function value are the losses generated during the reconstruction of the features of the input image samples by the image anomaly detection model, and the above second loss function value and fourth loss function value are the losses generated during the prediction of the abnormal pixels in the input image samples by the image anomaly detection model.

[0179] Step 540: Update the parameters of the image anomaly detection model through the loss function value.

[0180] For example, the computer device can update the parameters of the feature extractor and the codec through the first loss function value and the third loss function value, and update the parameters of the feature extractor, the codec, and the discriminator through the second loss function value and the fourth loss function value.

[0181] This application proposes an unsupervised + discriminative reconstruction integrated method for unified anomaly detection. Specifically, this method proposes a pixel-level discriminator to further enhance the distinction between normal samples and abnormal samples and refine the reconstruction error. In the inference stage, the results of unsupervised reconstruction and the pixel-level discriminator are integrated to significantly improve pixel-level anomaly segmentation.

[0182] Please refer to Figure 6 , which shows the unsupervised + discriminative reconstruction network architecture involved in this application. As Figure 6 shown, this network architecture mainly includes three modules, namely: a feature extraction part 610, an unsupervised reconstruction part 620, and a pixel-level discriminative reconstruction part 630.

[0183] The following is an introduction to the above three parts of the network architecture in turn:

[0184] 1) Feature extraction: Following the existing anomaly detection work, we use a pre-trained model on ImageNet to extract image features. Given an input image For example, this application extracts multi-scale feature maps from the first to fourth stages of a pre-trained convolutional network (such as EfficientNet-b4) That is:

[0185]

[0186] Among them where c i is the channel number, h i × w iis the spatial size level feature map of the i-th one. Considering that the semantic of the feature map at the bottom layer is weak but the resolution is high, the performance improvement brought by this is limited but the computational cost is expensive. In order to combine features from different hierarchies The sizes of all feature maps are adjusted to the same size (h 4 × w 4 ), that is, the size of the smallest feature map. Then, these resized feature maps are concatenated in the channel dimension. The whole process can be expressed as:

[0187]

[0188] where, h = h 4 and w = w 4 .

[0189] where, the above pre-trained model can use any pre-trained model based on a convolutional neural network or a Transformer architecture.

[0190] 2) Unsupervised reconstruction: Unsupervised reconstruction follows the standard Transformer architecture and consists of an encoding part and a decoding part. First, a linear projection layer with positional embeddings added is used to reduce the dimension of the multi-level features F extracted from the pre-trained model, and then global interaction and reconstruction of the features are performed through encoding and decoding composed of a series of Transformer blocks. Here, each Transformer block consists of multi-head attention and a fully connected feed-forward network. Finally, a linear projection layer is used for the decoding output to perform a dimension increase operation to restore to the original input feature dimension, and the output is denoted as The reconstruction loss function calculates the mean square error (MSE) between the reconstructed feature and the original feature, that is:

[0191]

[0192] where, for the above unsupervised reconstruction part, the specific network structure can use Transformer, or a convolutional neural network, or a fully connected neural network, etc.

[0193] 3) Discriminative Reconstruction: Using only unsupervised reconstruction, the performance of anomaly segmentation is still very poor. This is not surprising because it is only trained on anomaly-free training data, which may lead to "weak discriminability" between normal and abnormal samples in the feature space. To train a discriminator to enhance the discrimination between normal and abnormal samples, the simplest method is to learn a binary classifier through normal images and defective images. Unfortunately, for unsupervised anomaly detection tasks, only normal images are available. In fact, this is also a good setting for actual industrial inspection applications because defective products are always rare and difficult to collect on a large scale. To train the discriminator, normal images can be used to synthesize defective images. This application learns a discriminator that synthesizes defective images at the pixel level. In addition, this application designs a lightweight discriminator to refine the reconstruction error of the dual-mask autoencoder.

[0194] Given the normal training image I n and the corresponding anomaly mask Y n , represent the synthesized defective image and anomaly mask as I s and Y s . Then, input the normal image and the synthesized defective image {I t |t = n, s} into the multi-level feature extractor, and export its multi-level features as {F t |t = n, s}. Then, reconstruct {F t |t = n, s} using the unsupervised reconstruction network, and represent the corresponding features as Here, this application uses the absolute element-wise subtraction of the original features and the reconstructed features to measure their differences, i.e.:

[0195]

[0196] The above discriminator is designed with multiple convolutional blocks to learn features, and then followed by a 1x1 convolutional layer to perform pixel segmentation. Here, each convolutional block consists of a 3x3 convolution, BatchNorm, ReLU, and a 2x2 transposed convolution. Input the absolute reconstruction error {E t |t = n, s} into the designed discriminator, and obtain the estimated anomaly map To calculate and the ground truth Y t between the losses, resize the size of to the size of Y t . Considering that abnormal pixels usually account for a minority in anomaly detection, Dice loss can be used to calculate the loss function, which is very effective for learning from extremely imbalanced data, i.e.:

[0197]

[0198] where (i, j) represents Yt or spatial position

[0199] The abnormal reasoning process is as follows:

[0200] Pixel-level abnormal segmentation: The result of abnormal segmentation is an abnormal score map, which assigns an abnormal score to each pixel. For unsupervised reconstruction, the abnormal score map is calculated as the upsampling result of the L 2 norm, as shown below:

[0201]

[0202] For discriminative reconstruction, the abnormal score map is predicted as Finally, S res and are combined as the final abnormal segmentation map, that is:

[0203]

[0204] where ω ∈ [0, 1] is the weight

[0205] Image-level abnormal classification: Abnormal classification aims to detect whether an image contains abnormal regions. By taking the maximum value or standard deviation of the pixel-level prediction error S in the spatial dimension, it is converted into an image-level abnormal score

[0206] Experimental settings: Select EfficientNet-b4 pre-trained on ImageNet as the base network to extract features. The input image size is 224×224, and 3-stage features are extracted and scaled to a unified size of 14×14. Finally, the features are concatenated in the channel dimension to form the original features with a dimension of 272×14×14. The encoder and decoder each use a Transformer block

[0207] The main experimental conclusions are as follows:

[0208] 1) This solution can solve the problem of inaccurate positioning and segmentation of defective pixels in unsupervised reconstruction methods

[0209] Figure 7 is a comparison diagram of test results involved in this application, as Figure 7 shown, which shows a comparison diagram of the training loss and test metrics of the encoding and decoding networks of unsupervised reconstruction and the proposed unsupervised + discriminative reconstruction on the MVTec dataset

[0210] The results show that the ordinary unsupervised reconstruction network performs well only in image-level defect detection. The unsupervised + discriminative reconstruction encoding and decoding network proposed in this patent has obvious performance improvement in pixel-level defect segmentation

[0211] 2) The performance of this solution is significantly improved in the unified anomaly detection on the real industrial anomaly detection dataset, especially in the pixel-level anomaly segmentation. Please refer to Figure 8 , which shows a schematic diagram of anomaly recognition comparison involved in this application. As Figure 8 shown, compared with the unsupervised anomaly detection method (the second column) on 15 categories of MVTec, the anomaly detection boundary of the method proposed in this solution (the third column) is finer and more accurate, approaching the ground truth of manual annotation (the fourth column).

[0212] Fig. 9 is a structural block diagram of an image processing device provided by an exemplary embodiment of this application. As Fig. 9 shown, the device includes the following parts:

[0213] A feature extraction module 901, configured to extract features from a target image to obtain image extraction features of the target image;

[0214] An encoding and decoding module 902, configured to perform encoding and decoding processing on the image extraction features to obtain image reconstruction features;

[0215] A reconstruction module 903, configured to input a difference between the image reconstruction features and the image extraction features into a discriminator to obtain first recognition information output by the discriminator. The first recognition information includes first information elements respectively corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel has an anomaly; the discriminator is a machine learning model trained based on synthetic image samples, and the synthetic image samples are images with anomalies synthesized based on images without anomalies;

[0216] A recognition module 904, configured to obtain a recognition result of anomaly recognition for the target image based on the image extraction features, the image reconstruction features, and the first recognition information.

[0217] In a possible implementation manner, the recognition module 904 is configured to:

[0218] Obtain second recognition information based on a difference between the image extraction features and the image reconstruction features. The second recognition information includes second information elements respectively corresponding to each pixel in the target image, and the second information element is used to indicate the probability that the corresponding pixel has an anomaly;

[0219] Obtain segmentation information in the recognition result based on the first recognition information and the second recognition information. The segmentation information includes third information elements respectively corresponding to each pixel in the target image, and the third information element is used to indicate the probability that the corresponding pixel has an anomaly.

[0220] In a possible implementation, the dimension of the difference between the image extraction feature and the image reconstruction feature is less than the number of pixels of the target image; the recognition module 904 is configured to take the L2 norm of the difference between the image extraction feature and the image reconstruction feature, and perform upsampling on the L2 norm to obtain the second recognition information.

[0221] In a possible implementation, the recognition module 904 is configured to perform weighted summation or weighted averaging on the first recognition information and the second recognition information to obtain the segmentation information.

[0222] In a possible implementation, the recognition module 904 is further configured to, based on the segmentation information, obtain the image classification result in the recognition result, and the image classification result is used to indicate the probability that the target image belongs to an abnormal image.

[0223] In a possible implementation, the recognition module 904 is configured to take the maximum value or standard deviation of each third information element in the segmentation information to obtain the image classification result.

[0224] In a possible implementation, the feature extraction module 901 is configured to input the target image into a feature extractor in the image anomaly detection model to obtain the image extraction feature output by the feature extractor;

[0225] The encoding and decoding module 902 is configured to input the image extraction feature into an encoder-decoder in the image anomaly detection model to obtain the image reconstruction feature output by the encoder-decoder, and the size of the image reconstruction feature is the same as the size of the image extraction feature;

[0226] The reconstruction module 903 is configured to input the difference between the image reconstruction feature and the image extraction feature into a discriminator in the image anomaly detection model to obtain the first recognition information output by the discriminator;

[0227] Wherein, the image anomaly detection model is a machine learning model obtained by unsupervised training based on a first image sample; the first image sample is an image without anomalies.

[0228] In a possible implementation, the device further includes:

[0229] A sample processing module, configured to obtain first mask information, second mask information, and a second image sample based on the first image sample before the feature extraction module extracts features from a target image to obtain image extraction features of the target image; the first mask information is used to indicate pixels without anomalies in the first image sample, the second mask information is used to indicate pixels with anomalies in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information;

[0230] A model processing module, configured to input the first image sample into the feature extractor to obtain a first image extraction feature sample output by the feature extractor; input the first image extraction feature sample into the codec to obtain a first image reconstruction feature sample output by the codec; input the difference between the first image reconstruction feature sample and the first image extraction feature sample into the discriminator to obtain a first recognition information sample output by the discriminator; input the second image sample into the feature extractor to obtain a second image extraction feature sample output by the feature extractor; input the second image extraction feature sample into the codec to obtain a second image reconstruction feature sample output by the codec; input the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain a second recognition information sample output by the discriminator;

[0231] A loss acquisition module, configured to obtain a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample;

[0232] A parameter update module, configured to update parameters of the image anomaly detection model through the loss function value.

[0233] In a possible implementation manner, the loss acquisition module is configured to:

[0234] Obtain a first loss function value based on the difference between the first image extraction feature sample and the first image reconstruction feature sample;

[0235] Obtain a second loss function value based on the difference between the first recognition information sample and the first mask information;

[0236] Obtain a third loss function value based on the difference between the second image extraction feature sample and the second image reconstruction feature sample;

[0237] Obtain a fourth loss function value based on the difference between the second recognition information sample and the second mask information.

[0238] In summary, in the solution shown in the embodiments of the present application, after extracting the image extraction features from the target image, the image extraction features are subjected to feature reconstruction through an encoding and decoding method to obtain image reconstruction features. Further, a discriminator is used to perform anomaly discrimination on the difference between the image reconstruction features and the image extraction features to obtain first recognition information indicating the probability of whether each pixel in the target image is abnormal. Finally, by combining the image features before and after reconstruction and the first recognition information, an identification result for anomaly recognition of the target image is comprehensively obtained. In addition to using the image features before and after reconstruction to identify the abnormal conditions of the pixels in the target image, the above solution also uses the discriminator to discriminate the abnormal conditions of the pixels in the target image based on the difference between the image features before and after reconstruction. On the one hand, by combining the image features before and after reconstruction and the discrimination result of the difference between the image features before and after reconstruction, the abnormal conditions of each pixel in the target image are comprehensively determined, so that pixel-level image anomaly detection can be realized while ensuring the accuracy of image anomaly detection. On the other hand, the above discriminator is trained through synthetic image samples obtained by normal image synthesis. Therefore, the training process of the discriminator does not require manual collection and annotation of abnormal images, and it can be trained in an unsupervised learning manner, thereby being able to expand the number of available samples during the training of the discriminator and make up for the difference between the effect of training the discriminator through synthetic image samples and the effect of training the discriminator through real abnormal images, ensuring the accuracy of the discriminator, and then ensuring the accuracy of detecting abnormal pixels in the image through the discriminator.

[0239] Fig.10 is a structural block diagram of an image processing device provided by an exemplary embodiment of the present application. As Fig. 9 shown, the device includes the following parts:

[0240] A sample processing module 1001, configured to obtain first mask information, second mask information, and a second image sample based on a first image sample; the first mask information is used to indicate pixels in the first image sample that are not abnormal, the second mask information is used to indicate pixels in the image sample that are abnormal, and the second image sample is an image synthesized based on the first image sample and the second mask information;

[0241] The model processing module 1002 is configured to input the first image sample into a feature extractor in an image anomaly detection model to obtain a first image extraction feature sample output by the feature extractor; input the first image extraction feature sample into an encoder-decoder in the image anomaly detection model to obtain a first image reconstruction feature sample output by the encoder-decoder; input the difference between the first image reconstruction feature sample and the first image extraction feature sample into a discriminator in the image anomaly detection model to obtain a first recognition information sample output by the discriminator; input the second image sample into the feature extractor to obtain a second image extraction feature sample output by the feature extractor; input the second image extraction feature sample into the encoder-decoder to obtain a second image reconstruction feature sample output by the encoder-decoder; input the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain a second recognition information sample output by the discriminator.

[0242] The loss acquisition module 1003 is configured to obtain a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample.

[0243] The parameter update module 1004 is configured to update the parameters of the image anomaly detection model through the loss function value.

[0244] Wherein, the image anomaly detection model is configured to extract features from a target image to obtain an image extraction feature of the target image, perform encoding and decoding processing on the image extraction feature to obtain an image reconstruction feature, and perform discrimination processing based on the difference between the image reconstruction feature and the image extraction feature to obtain first recognition information, where the first recognition information includes first information elements respectively corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel is abnormal; the image extraction feature, the image reconstruction feature, and the first recognition information are used to obtain a recognition result for anomaly recognition of the target image.

[0245] In a possible implementation manner, the loss acquisition module is configured to:

[0246] Obtain a first loss function value based on the difference between the first image extraction feature sample and the first image reconstruction feature sample;

[0247] Obtain a second loss function value based on the difference between the first recognition information sample and the first mask information;

[0248] Obtain a third loss function value based on the difference between the feature sample extracted from the second image and the second image reconstruction feature sample;

[0249] Obtain a fourth loss function value based on the difference between the second recognition information sample and the second mask information.

[0250] It should be noted that: For the image processing device provided in the above embodiment, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the image processing device provided in the above embodiment and the embodiment of the image processing method belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.

[0251] Fig.11 The structural schematic diagram of a server provided by an exemplary embodiment of the present application is shown. Specifically:

[0252] The server 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory (RAM) 1102 and a read only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The server 1100 also includes a mass storage device 1106 for storing an operating system 1113, an application program 1114, and other program modules 1115.

[0253] The mass storage device 1106 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1106 and its associated computer-readable medium provide non-volatile storage for the server 1100. That is to say, the mass storage device 1106 can include a computer-readable medium (not shown) such as a hard disk or a compact disc read only memory (CD-ROM) drive.

[0254] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will understand that computer storage media is not limited to the above several types. The above system memory 1104 and mass storage device 1106 can be collectively referred to as memory.

[0255] According to various embodiments of the present application, the server 1100 can also operate by connecting to a remote computer on the network through a network such as the Internet. That is, the server 1100 can be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105, or in other words, the network interface unit 1111 can also be used to connect to other types of networks or remote computer systems (not shown).

[0256] The above memory further includes one or more programs, and the one or more programs are stored in the memory and configured to be executed by the CPU.

[0257] Embodiments of the present application also provide a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the image processing method provided by the above method embodiments.

[0258] Embodiments of the present application also provide a computer-readable storage medium, on which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the image processing method provided by the above method embodiments.

[0259] Embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the image processing method described in any one of the above embodiments.

[0260] Optionally, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid state drives (SSD, Solid State Drives), or optical discs, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance Random Access Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory). The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0261] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The described program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be read-only memory, a magnetic disk, or an optical disc, etc.

[0262] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. An image processing method, characterized in that, the method includes: performing feature extraction on a target image to obtain image extraction features of the target image; performing encoding and decoding processing on the image extraction features to obtain image reconstruction features; subtracting the image extraction features and the image reconstruction features element by element to obtain a difference between the image extraction features and the image reconstruction features, performing upsampling after taking the L2 norm of the difference between the image extraction features and the image reconstruction features to obtain second recognition information; the dimension of the difference between the image extraction features and the image reconstruction features is less than the number of pixels of the target image, the second recognition information includes second information elements respectively corresponding to each pixel in the target image, and the second information element is used to indicate the probability that the corresponding pixel is abnormal, and the second information elements in the second recognition information correspond one by one to the pixels in the target image; inputting the difference between the image reconstruction features and the image extraction features into a discriminator to obtain first recognition information output by the discriminator; the first recognition information includes first information elements respectively corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel is abnormal, and the first information elements in the first recognition information correspond one by one to the pixels in the target image; the discriminator is a machine learning model trained based on synthetic image samples, and the synthetic image samples are abnormal images synthesized based on images without abnormalities; the dimension numbers of the first recognition information and the second recognition information are the same; integrating the information elements on the same dimension in the second recognition information and the first recognition information to obtain an integration result for each dimension, and using it as segmentation information in the recognition result for abnormal recognition of the target image; the segmentation information includes third information elements respectively corresponding to each pixel in the target image, and the third information element is used to indicate the probability that the corresponding pixel is abnormal, and each third information element corresponds one by one to each pixel in the target image, and the segmentation information is used to represent whether each pixel in the target image is abnormal; statistically processing each third information element in the segmentation information to obtain an image classification result, and the image classification result is used to represent whether the target image belongs to an abnormal image as a whole.

2. The method according to claim 1, characterized in that, the method further includes: performing weighted summation or weighted averaging on the first recognition information and the second recognition information to obtain the segmentation information.

3. The method according to claim 1, characterized in that, the method further includes: taking the maximum value or standard deviation of each third information element in the segmentation information to obtain the image classification result.

4. The method according to any one of claims 1 to 3, characterized in that, the performing feature extraction on a target image to obtain image extraction features of the target image includes: Input the target image into the feature extractor in the image anomaly detection model to obtain the image extraction features output by the feature extractor; The encoding and decoding process of the image extraction features to obtain image reconstruction features includes: Input the image extraction features into the encoder-decoder in the image anomaly detection model to obtain the image reconstruction features output by the encoder-decoder, and the size of the image reconstruction features is the same as that of the image extraction features; Input the difference between the image reconstruction features and the image extraction features into the discriminator to obtain the first recognition information output by the discriminator, including: Input the difference between the image reconstruction features and the image extraction features into the discriminator in the image anomaly detection model to obtain the first recognition information output by the discriminator; Wherein, the image anomaly detection model is a machine learning model obtained by unsupervised training based on the first image sample; the first image sample is an image without anomalies.

5. The method according to claim 4, characterized in that, Before extracting the features of the target image to obtain the image extraction features of the target image, it further includes: Based on the first image sample, obtain the first mask information, the second mask information, and the second image sample; the first mask information is used to indicate the pixels without anomalies in the first image sample, the second mask information is used to indicate the pixels with anomalies in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information; Input the first image sample into the feature extractor to obtain the first image extraction feature sample output by the feature extractor; Input the first image extraction feature sample into the encoder-decoder to obtain the first image reconstruction feature sample output by the encoder-decoder; Input the difference between the first image reconstruction feature sample and the first image extraction feature sample into the discriminator to obtain the first recognition information sample output by the discriminator; Input the second image sample into the feature extractor to obtain the second image extraction feature sample output by the feature extractor; Input the second image extraction feature sample into the encoder-decoder to obtain the second image reconstruction feature sample output by the encoder-decoder; Input the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain the second recognition information sample output by the discriminator; Obtain the loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample; Update the parameters of the image anomaly detection model through the loss function value.

6. The method according to claim 5, characterized in that, Obtaining a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample includes: Obtaining a first loss function value based on the difference between the first image extraction feature sample and the first image reconstruction feature sample; Obtaining a second loss function value based on the difference between the first recognition information sample and the first mask information; Obtaining a third loss function value based on the difference between the second image extraction feature sample and the second image reconstruction feature sample; Obtaining a fourth loss function value based on the difference between the second recognition information sample and the second mask information.

7. An image processing method Characterized in that The method includes: Based on a first image sample, obtaining first mask information, second mask information, and a second image sample; the first mask information is used to indicate pixels without anomalies in the first image sample, the second mask information is used to indicate pixels with anomalies in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information; Inputting the first image sample into a feature extractor in an image anomaly detection model to obtain a first image extraction feature sample output by the feature extractor; Inputting the first image extraction feature sample into an encoder-decoder in the image anomaly detection model to obtain a first image reconstruction feature sample output by the encoder-decoder; Inputting the difference between the first image reconstruction feature sample and the first image extraction feature sample into a discriminator in the image anomaly detection model to obtain a first recognition information sample output by the discriminator; the discriminator is a machine learning model trained based on synthetic image samples, and the synthetic image samples are images with anomalies synthesized based on images without anomalies; Inputting the second image sample into the feature extractor to obtain a second image extraction feature sample output by the feature extractor; Inputting the second image extraction feature sample into the encoder-decoder to obtain a second image reconstruction feature sample output by the encoder-decoder; Inputting the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain a second recognition information sample output by the discriminator; Obtaining a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample; Updating the parameters of the image anomaly detection model through the loss function value; Among them, the image anomaly detection model is used to extract features from the target image to obtain the image extraction features of the target image, perform encoding and decoding processing on the image extraction features to obtain image reconstruction features, and through the discriminator, perform discrimination processing based on the difference between the image reconstruction features and the image extraction features to obtain the first recognition information. The first recognition information includes first information elements corresponding to each pixel in the target image respectively, and the first information element is used to indicate the probability that the corresponding pixel is abnormal. The first information elements in the first recognition information are in one-to-one correspondence with the pixels in the target image; the difference between the image extraction features and the image reconstruction features is used to obtain the second recognition information, and the second recognition information is obtained by upsampling after taking the L2 norm of the difference between the image extraction features and the image reconstruction features. The second recognition information includes second information elements corresponding to each pixel in the target image respectively, and the second information element is used to indicate the probability that the corresponding pixel is abnormal. The second information elements in the second recognition information are in one-to-one correspondence with the pixels in the target image; the first recognition information and the second recognition information have the same number of dimensions, the dimension of the difference between the image extraction features and the image reconstruction features is less than the number of pixels of the target image, and the difference between the image reconstruction features and the image extraction features is obtained by subtracting the image extraction features and the image reconstruction features element by element; The second recognition information and the first recognition information are used to obtain the recognition result of anomaly recognition for the target image, and the determination process of the recognition result includes: Integrate the information elements on the same dimension in the second recognition information and the first recognition information to obtain the integration result of each dimension, which is used as the segmentation information in the recognition result of anomaly recognition for the target image; the segmentation information includes third information elements corresponding to each pixel in the target image respectively, and the third information element is used to indicate the probability that the corresponding pixel is abnormal. Each third information element is in one-to-one correspondence with each pixel in the target image, and the segmentation information is used to indicate whether each pixel in the target image is abnormal; Statistically analyze each third information element in the segmentation information to obtain an image classification result, and the image classification result is used to represent whether the target image belongs to an abnormal image as a whole.

8. An image processing device, Characterized in that, The device includes: A feature extraction module, configured to extract features from a target image to obtain the image extraction features of the target image; An encoding and decoding module, configured to perform encoding and decoding processing on the image extraction features to obtain image reconstruction features; A module for performing the following steps: subtracting the image extraction features and the image reconstruction features element by element to obtain the difference between the image extraction features and the image reconstruction features; An identification module, configured to perform upsampling after taking the L2 norm of the difference between the features extracted from the image and the features reconstructed from the image, so as to obtain second identification information; the dimension of the difference between the features extracted from the image and the features reconstructed from the image is smaller than the number of pixels of the target image, and the second identification information includes second information elements respectively corresponding to each pixel in the target image, and the second information element is used to indicate the probability that the corresponding pixel is abnormal, and the second information elements in the second identification information are in one-to-one correspondence with the pixels in the target image; A reconstruction module, configured to input the difference between the features reconstructed from the image and the features extracted from the image into a discriminator, and obtain first identification information output by the discriminator; the first identification information includes first information elements respectively corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel is abnormal, and the first information elements in the first identification information are in one-to-one correspondence with the pixels in the target image; the discriminator is a machine learning model trained based on synthetic image samples, and the synthetic image samples are abnormal images synthesized based on images without abnormalities; the dimension numbers of the first identification information and the second identification information are the same; The identification module is further configured to integrate the information elements on the same dimension in the second identification information and the first identification information to obtain an integration result for each dimension, and use the integration result as segmentation information in the identification result of abnormal identification of the target image; the segmentation information includes third information elements respectively corresponding to each pixel in the target image, and the third information element is used to indicate the probability that the corresponding pixel is abnormal, and each third information element is in one-to-one correspondence with each pixel in the target image, and the segmentation information is used to indicate whether each pixel in the target image is abnormal; by statistically analyzing each third information element in the segmentation information, an image classification result is obtained, and the image classification result is used to represent whether the target image belongs to an abnormal image as a whole.

9. The apparatus according to claim 8, wherein, the identification module is configured to perform weighted summation or weighted averaging on the first identification information and the second identification information to obtain the segmentation information.

10. The apparatus according to claim 8, wherein, the identification module is configured to: take the maximum value or standard deviation of each third information element in the segmentation information to obtain the image classification result.

11. The apparatus according to any one of claims 8 to 10, wherein, the feature extraction module is configured to input the target image into a feature extractor in an image anomaly detection model to obtain the image extraction features output by the feature extractor; the encoding and decoding module is configured to input the image extraction features into an encoder-decoder in the image anomaly detection model to obtain the image reconstruction features output by the encoder-decoder, and the size of the image reconstruction features is the same as the size of the image extraction features; The reconstruction module is configured to input the difference between the image reconstruction feature and the image extraction feature into the discriminator in the image anomaly detection model, and obtain the first recognition information output by the discriminator; Wherein, the image anomaly detection model is a machine learning model obtained by unsupervised training based on a first image sample; the first image sample is an image without anomalies.

12. The apparatus according to claim 11, wherein, The apparatus further includes: A sample processing module, configured to obtain first mask information, second mask information, and a second image sample based on the first image sample; the first mask information is used to indicate the pixels without anomalies in the first image sample, the second mask information is used to indicate the pixels with anomalies in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information; A model processing module, configured to input the first image sample into the feature extractor to obtain a first image extraction feature sample output by the feature extractor; input the first image extraction feature sample into the codec to obtain a first image reconstruction feature sample output by the codec; input the difference between the first image reconstruction feature sample and the first image extraction feature sample into the discriminator to obtain a first recognition information sample output by the discriminator; input the second image sample into the feature extractor to obtain a second image extraction feature sample output by the feature extractor; input the second image extraction feature sample into the codec to obtain a second image reconstruction feature sample output by the codec; input the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain a second recognition information sample output by the discriminator; A loss acquisition module, configured to obtain a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample; A parameter update module, configured to update the parameters of the image anomaly detection model through the loss function value.

13. The apparatus according to claim 12, wherein, The loss acquisition module is configured to: Obtain a first loss function value based on the difference between the first image extraction feature sample and the first image reconstruction feature sample; Obtain a second loss function value based on the difference between the first recognition information sample and the first mask information; Obtain a third loss function value based on the difference between the second image extraction feature sample and the second image reconstruction feature sample; Obtain a fourth loss function value based on the difference between the second recognition information sample and the second mask information.

14. An image processing apparatus, wherein, The apparatus includes: A sample processing module, configured to obtain first mask information, second mask information, and a second image sample based on a first image sample; the first mask information is used to indicate pixels without anomalies in the first image sample, the second mask information is used to indicate pixels with anomalies in the image sample, and the second image sample is an image synthesized based on the first image sample and the second mask information; A model processing module, configured to input the first image sample into a feature extractor in an image anomaly detection model to obtain a first image extraction feature sample output by the feature extractor; input the first image extraction feature sample into an encoder-decoder in the image anomaly detection model to obtain a first image reconstruction feature sample output by the encoder-decoder; input the difference between the first image reconstruction feature sample and the first image extraction feature sample into a discriminator in the image anomaly detection model to obtain a first recognition information sample output by the discriminator; the discriminator is a machine learning model trained based on synthesized image samples, and the synthesized image samples are images with anomalies synthesized based on images without anomalies; input the second image sample into the feature extractor to obtain a second image extraction feature sample output by the feature extractor; input the second image extraction feature sample into the encoder-decoder to obtain a second image reconstruction feature sample output by the encoder-decoder; input the difference between the second image reconstruction feature sample and the second image extraction feature sample into the discriminator to obtain a second recognition information sample output by the discriminator; A loss obtaining module, configured to obtain a loss function value based on the first image extraction feature sample, the first image reconstruction feature sample, the first recognition information sample, the second image extraction feature sample, the second image reconstruction feature sample, and the second recognition information sample; A parameter updating module, configured to update parameters of the image anomaly detection model through the loss function value; Among them, the image anomaly detection model is used to extract features from the target image to obtain the image extraction features of the target image, perform encoding and decoding processing on the image extraction features to obtain image reconstruction features, and through the discriminator, perform discrimination processing based on the difference between the image reconstruction features and the image extraction features to obtain first recognition information. The first recognition information includes first information elements respectively corresponding to each pixel in the target image, and the first information element is used to indicate the probability that the corresponding pixel has an anomaly. The first information elements in the first recognition information correspond one-to-one with the pixels in the target image; the difference between the image extraction features and the image reconstruction features is used to obtain second recognition information. The second recognition information is obtained by upsampling the L2 norm of the difference between the image extraction features and the image reconstruction features. The second recognition information includes second information elements respectively corresponding to each pixel in the target image, and the second information element is used to indicate the probability that the corresponding pixel has an anomaly. The second information elements in the second recognition information correspond one-to-one with the pixels in the target image; the first recognition information and the second recognition information have the same number of dimensions. The dimension of the difference between the image extraction features and the image reconstruction features is less than the number of pixels of the target image. The difference between the image reconstruction features and the image extraction features is obtained by subtracting the image extraction features and the image reconstruction features element by element; The second recognition information and the first recognition information are used to obtain the recognition result of anomaly recognition for the target image. The determination process of the recognition result includes: Integrate the information elements on the same dimension in the second recognition information and the first recognition information to obtain the integration result of each dimension, which is used as the segmentation information in the recognition result of anomaly recognition for the target image; the segmentation information includes third information elements respectively corresponding to each pixel in the target image, and the third information element is used to indicate the probability that the corresponding pixel has an anomaly. Each third information element corresponds one-to-one with each pixel in the target image, and the segmentation information is used to indicate whether each pixel in the target image has an anomaly; Statistically analyze each third information element in the segmentation information to obtain an image classification result, and the image classification result is used to represent whether the target image belongs to an abnormal image as a whole.

15. A computer device, Characterized in that, The computer device includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the image processing method according to any one of claims 1 to 7.

16. A computer-readable storage medium, Characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 7.

17. A computer program product, Characterized in that, including computer instructions which, when executed by a processor, implement the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Time series data detection method and device

    CN112000830A

  • Abnormal image detection method and device, model training method and device, equipment and medium

    CN116958020A