Anomaly detection method, model training method, device, equipment and medium
Patent Information
- Application Number
- CN202310290119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-03-13
AI Technical Summary
但在实际应用中,许多异常区域在彩色图像和灰度图像中的清晰度较低,或与正常区域的差异较小,易造成模型漏检,检测效果和精度较差
[0030]本申请的技术方案将待检测图像输入目标网络模型进行特征提取、深度估计和图像重构,得到图像特征、参考特征、深度估计图像和重构图像;基于待检测图像、待检测图像对应的深度图像、图像特征、参考特征、重构图像和深度估计图像进行图像异常分析,得到异常响应图;基于异常响应图生成待检测图像的异常检测结果;通过获取深度图像,并在网络模型中进行待检测图像的多特征提取、深度估计、图像重构的多任务检测,能够充分结合深度信息以及充分提取图像特征,以弥补待检测图像中单一特征信息的局限,通过多任务输出提高异常检测信息维度和实体特征表达准确性,以提高异常检测精度,降低漏检率。
Smart Images

Figure CN116958033B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an anomaly detection method, model training method, apparatus, device, and medium. Background Technology
[0002] Anomaly detection aims to identify abnormal regions in images and is widely used in fields such as industrial quality inspection and medical image analysis. Related technologies typically employ deep neural networks to identify defects, lesions, and other anomalies in color or grayscale images. However, due to the scarcity of abnormal samples, ordinary stacked data-based deep learning methods cannot be directly applied. Therefore, normal color or grayscale image samples are usually used for network training to model normal data and achieve anomaly recognition. But in practical applications, many abnormal regions have low clarity in color and grayscale images, or only slight differences from normal regions, easily leading to missed detections and poor detection performance and accuracy. Summary of the Invention
[0003] This application provides an anomaly detection method, model training method, apparatus, device, and medium, which can significantly improve the accuracy of anomaly detection and reduce the false negative rate.
[0004] On the one hand, this application provides an anomaly detection method, the method comprising:
[0005] Obtain the image to be detected and the corresponding depth image of the image to be detected;
[0006] The image to be detected is input into the target network model for feature extraction, depth estimation, and image reconstruction to obtain image features, reference features, depth-estimated image, and reconstructed image.
[0007] Image anomaly analysis is performed based on the image to be detected, the depth image, the image features, the reference features, the reconstructed image, and the depth estimation image to obtain an anomaly response map;
[0008] Anomaly detection results of the image to be detected are generated based on the anomaly response map.
[0009] On the other hand, a model training method is provided, the method comprising:
[0010] Obtain a training set, which includes multiple pairs of sample images, each pair of sample images including a positive sample image and a sample depth image corresponding to the positive sample image;
[0011] The positive sample image is input into a preset neural network for feature extraction, depth estimation and image reconstruction to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image;
[0012] The model loss is calculated based on the positive sample image, the sample depth image, sample image features, sample reference features, sample depth estimation image, and sample reconstruction image.
[0013] The preset neural network is trained based on the model loss to obtain the target network model.
[0014] On the other hand, an anomaly detection device is provided, the device comprising:
[0015] Image acquisition module: used to acquire the image to be detected and the corresponding depth image;
[0016] Detection module: used to input the image to be detected into the target network model for feature extraction, depth estimation and image reconstruction, to obtain image features, reference features, depth estimated image and reconstructed image;
[0017] Anomaly analysis module: used to perform image anomaly analysis based on the image to be detected, the depth image, the image features, the reference features, the reconstructed image, and the depth estimation image, and obtain an anomaly response map;
[0018] Result generation module: used to generate anomaly detection results for the image to be detected based on the anomaly response map.
[0019] On the other hand, a training apparatus for a target network model is provided, the apparatus comprising:
[0020] Sample acquisition module: used to acquire a training set, which includes multiple sample image pairs, and each sample image pair includes a positive sample image and a sample depth image corresponding to the positive sample image;
[0021] Neural network module: used to input the positive sample image into a preset neural network for feature extraction, depth estimation and image reconstruction, to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image;
[0022] Loss calculation module: used to calculate the model loss based on the positive sample image, the sample depth image, sample image features, sample reference features, sample depth estimation image and sample reconstruction image;
[0023] Training module: used to train the preset neural network based on the model loss to obtain the target network model.
[0024] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the anomaly detection method or the model training method as described above.
[0025] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the anomaly detection method or the model training method described above.
[0026] On the other hand, a server is provided, the server including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the anomaly detection method or the model training method as described above.
[0027] On the other hand, a terminal is provided, the terminal including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the anomaly detection method or the model training method as described above.
[0028] On the other hand, a computer program product or computer program is provided, which includes computer instructions that, when executed by a processor, implement the anomaly detection method as described above or the model training method as described above.
[0029] The anomaly detection method, model training method, apparatus, equipment, storage medium, server, terminal, computer program, and computer program product provided in this application have the following technical effects:
[0030] The technical solution of this application inputs the image to be detected into a target network model for feature extraction, depth estimation, and image reconstruction to obtain image features, reference features, depth-estimated images, and reconstructed images. Based on the image to be detected, the corresponding depth image, image features, reference features, reconstructed images, and depth-estimated images, image anomaly analysis is performed to obtain an anomaly response map. Anomaly detection results for the image to be detected are generated based on the anomaly response map. By acquiring depth images and performing multi-task detection of multi-feature extraction, depth estimation, and image reconstruction of the image to be detected in the network model, depth information and image features can be fully combined to compensate for the limitations of single feature information in the image to be detected. By improving the dimensionality of anomaly detection information and the accuracy of entity feature expression through multi-task output, the anomaly detection accuracy is improved and the false negative rate is reduced. Attached Figure Description
[0031] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application;
[0033] Figure 2 This is a flowchart illustrating an anomaly detection method provided in an embodiment of this application;
[0034] Figure 3 This is a flowchart illustrating another anomaly detection method provided in an embodiment of this application;
[0035] Figure 4 This is a flowchart illustrating another anomaly detection method provided in an embodiment of this application;
[0036] Figure 5 This is a schematic diagram of the framework structure of a target network model provided in an embodiment of this application;
[0037] Figure 6 This is a schematic flowchart of a model training method provided in an embodiment of this application;
[0038] Figure 7 This is a schematic diagram of the frame of an anomaly detection device provided in an embodiment of this application;
[0039] Figure 8 This is a schematic diagram of the framework of a model training device provided in an embodiment of this application;
[0040] Figure 9 This is a hardware structure block diagram of an electronic device for an anomaly detection method provided in an embodiment of this application. Detailed Implementation
[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or sub-modules is not necessarily limited to those steps or sub-modules explicitly listed, but may include other steps or sub-modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0043] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0044] Depth estimation: This refers to the task of predicting the distance from each pixel to the camera when given an RGB image.
[0045] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0046] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0047] Computer vision (CV) is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0048] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely used in many fields. The solutions provided in the embodiments of this application involve artificial intelligence technologies such as machine learning / deep learning and natural language processing, which are specifically illustrated through the following embodiments.
[0049] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application, such as... Figure 1 As shown, the application environment may include at least terminal 01 and server 02. In practical applications, terminal 01 and server 02 can be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions on this.
[0050] In this application embodiment, server 02 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0051] Specifically, cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology can be applied to various fields, such as medical cloud, cloud IoT, cloud security, cloud education, cloud conferencing, AI cloud services, cloud applications, cloud calling, and cloud social networking. Based on the cloud computing business model, cloud technology distributes computing tasks across a resource pool composed of numerous computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, the resources in the "cloud" are infinitely scalable, readily available, on-demand, expandable, and pay-as-you-go. As a provider of basic cloud computing capabilities, a cloud resource pool (referred to as a cloud platform, generally called IaaS (Infrastructure as a Service)) platform is established, deploying various types of virtual resources within the pool for external customers to choose from. The cloud resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, and network devices.
[0052] Based on logical function, a PaaS (Platform as a Service) layer can be deployed on top of the IaaS layer, and a SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, SaaS can be deployed directly on top of IaaS. PaaS is a platform for running software, such as databases and web containers. SaaS refers to various types of business software, such as web portals and bulk SMS senders. Generally speaking, SaaS and PaaS are upper layers compared to IaaS.
[0053] Specifically, the server 02 mentioned above may include physical devices, such as network communication submodules, processors, and memory, and may also include software running on the physical devices, such as applications.
[0054] Specifically, terminal 01 may include physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart voice interaction devices, smart home appliances, smart wearable devices, and in-vehicle terminal devices, and may also include software running on the physical devices, such as applications.
[0055] In this embodiment, terminal 01 can be used to acquire or receive an image to be detected and a depth image, and call a target network model to perform anomaly detection to obtain anomaly detection results. Alternatively, it can send the image to be detected and the depth image to server 02, so that server 02 can call the target network model to perform anomaly detection to obtain anomaly detection results. Server 02 can be used to provide anomaly detection services to obtain anomaly detection results for the image to be detected. Specifically, server 02 can also be used to provide model training services for a preset neural network to obtain a target network model, and can also be used to store training sets and model training data, etc.
[0056] Specifically, the anomaly detection method of this application can be applied to various application scenarios of anomaly detection, such as detecting defects and damage in input images in industrial quality inspection, locating lesions in medical images, or locating risk objects in images of physical entities such as buildings.
[0057] Furthermore, it is understandable that Figure 1 The example shown is merely an application environment for an anomaly detection method. This application environment may include more or fewer nodes, and this application does not impose any restrictions here.
[0058] The application environment involved in this application embodiment, or the terminal 01 and server 02 in the application environment, can be a distributed system formed by connecting clients and multiple nodes (any form of computing device accessing the network, such as servers and user terminals) through network communication. The distributed system can be a blockchain system, which can provide the aforementioned anomaly detection service and data storage service, etc.
[0059] The following describes an anomaly detection method based on the aforementioned application environment. This method can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving. For example, it can be applied to industrial AI quality inspection projects using depth cameras, achieving the detection of abnormal defects using only normal samples for training. Please refer to... Figure 2 , Figure 2 This is a flowchart illustrating an anomaly detection method provided in an embodiment of this application. This specification provides method operation steps as shown in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or server product execution, the method can be executed sequentially according to the embodiments or drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as... Figure 2 As shown, the method may include the following steps S201-S207.
[0060] S201: Obtain the image to be detected and the corresponding depth image.
[0061] In this embodiment, the image to be detected and the depth image are images including the object to be detected, which can be, for example, a product, an organ, or other entity. The image to be detected is an RGB image captured of the object to be detected, such as one generated from an image acquired by a 2D camera. The depth image can be an image containing three-dimensional depth information, such as one generated from an image acquired by a 3D camera. The image to be detected and the depth image can be preprocessed images of the same size, and the same pixel in the image to be detected and the depth image correspond to the same entity location.
[0062] S203: Input the image to be detected into the target network model for feature extraction, depth estimation and image reconstruction to obtain image features, reference features, depth estimated image and reconstructed image.
[0063] In this embodiment, the target network model is obtained by using positive sample images as input to a preset neural network, and by performing unsupervised training on the preset neural network based on the output sample image features, sample reference features, sample depth estimation images, and sample reconstructed images, as well as the loss of the positive sample images and the corresponding sample depth images. Both the positive sample images and the sample depth images are positive samples without abnormal regions.
[0064] The pre-defined neural network can include a feature extraction task branch, a depth estimation task branch, and a reconstruction task branch. The model loss can be generated based on distillation loss, depth estimation loss, and reconstruction loss. Specifically, the model parameters of the feature extraction task branch of the pre-defined neural network are adjusted using distillation loss, the model parameters of the depth estimation task branch of the pre-defined neural network are adjusted using depth estimation loss, and the model parameters of the reconstruction task branch of the pre-defined neural network are adjusted using reconstruction loss to obtain an updated pre-defined neural network. This network is then iteratively trained until the training termination condition is met, resulting in the target network model. The feature extraction task branch, depth estimation task branch, and reconstruction task branch all include a shared backbone network.
[0065] Specifically, the reference features and image features are aligned and are both deep features, i.e., semantic features, used to characterize the semantic information of the image to be detected. Both the reference features and image features include feature maps output by at least one network layer; in some embodiments, they include feature maps output by two or more network layers.
[0066] In some embodiments, please refer to Figure 5 The target network model includes a backbone network, auxiliary networks, a deep decoder, and a reconstruction decoder. The auxiliary networks are the teacher models corresponding to the backbone network. Please refer to the relevant documentation for further details. Figure 3S203 may include S2031-S2034:
[0067] S2031: Input the image to be detected into the backbone network for feature extraction to obtain image features;
[0068] S2032: Input the image to be detected into the auxiliary network for feature extraction to obtain reference features;
[0069] S2033: Input the image features into the depth decoder to perform depth estimation and obtain the depth estimated image;
[0070] S2034: Input the image features into the reconstruction decoder to reconstruct the image and obtain the reconstructed image.
[0071] As shown in the figure, the backbone network and auxiliary network form a feature distillation task branch, the backbone network and deep decoder form a depth estimation task branch, and the backbone network and reconstruction decoder form an image reconstruction task branch. The backbone network and auxiliary network can have the same network structure. The teacher model is used to extract reference features and guide the feature extraction training of the backbone network. In some embodiments, the auxiliary network is a classification model pre-trained based on a pre-trained image set, such as a deep neural network trained on ImageNet. The deep decoder decodes the features extracted by the backbone network into a depth estimation image, and the depth image and the depth estimation image have the same size. The reconstruction decoder decodes the features extracted by the backbone network into a reconstructed image, and the reconstructed image has the same image category and size as the image to be detected, such as both being RGB images. This enables anomaly detection scenarios with RGBD as input, obtaining multi-dimensional outputs through multiple task branches. This is beneficial for improving the effectiveness of distillation and reconstruction learning during training by combining depth information, and can achieve multiple detections of anomalies such as defects through image feature extraction, reference feature extraction, difference analysis, and depth estimation, thereby improving detection performance.
[0072] S205: Perform image anomaly analysis based on the image to be detected, depth image, image features, reference features, reconstructed image, and depth estimation image to obtain an anomaly response map.
[0073] In this embodiment of the application, the anomaly response map can be used to characterize the abnormal distribution or abnormal region of the image to be detected. The anomaly response map and the image to be detected have the same size. It includes the anomaly evaluation index of each pixel in the image to be detected. The higher the anomaly evaluation index, the greater the confidence that the pixel belongs to the abnormal region, and vice versa.
[0074] In some embodiments, please refer to Figure 4 S205 may include S301-S309:
[0075] S301: Perform pixel feature difference analysis based on image features and reference features to obtain the feature difference results corresponding to the pixels in the image to be detected;
[0076] S303: Perform pixel difference analysis on the image to be detected and the reconstructed image to obtain the reconstruction difference results corresponding to the pixels;
[0077] S305: Perform pixel depth difference analysis on the depth image and the depth estimation image to obtain the depth difference results corresponding to the pixels;
[0078] S307: Based on feature difference results, depth difference results, and reconstruction difference results, determine the anomaly evaluation index for each pixel in the image to be detected;
[0079] S309: Generate an anomaly response map based on the anomaly evaluation index of each pixel.
[0080] Specifically, the anomaly evaluation metric can be an anomaly score for each pixel. A heatmap is generated based on the anomaly score of each pixel in the image to be detected, resulting in an anomaly response map. The anomaly score in the response map corresponds one-to-one with the pixel in the image to be detected. The anomaly evaluation metric for a single pixel can be the statistical value of the feature difference result, reconstruction difference result, and depth difference result for that pixel, such as a weighted sum or weighted average of these results. Thus, by combining the outputs of multiple task branches to determine various difference analysis results, an anomaly response map that integrates multi-dimensional anomaly information can be obtained, improving the accuracy of anomaly evaluation for each pixel and thereby optimizing the anomaly detection effect.
[0081] The image features and the reference features have the same feature size and number of feature maps, including the feature map corresponding to each pixel in the image to be detected.
[0082] In some embodiments, S301 may specifically include: performing a difference analysis on the feature map corresponding to the pixel at the same location in the image features and its corresponding feature map in the reference features to obtain the feature difference result of each pixel in the image to be detected.
[0083] In one embodiment, the feature extraction task branch is constructed based on feature distillation. The backbone network and auxiliary network use the same ResNet50 structure, including 5 network sub-modules (stages). The input size is 3×256×256. The feature size of stage 0 is 64×128×128, the feature size of stage 1 is 256×64×64, the feature size of stage 2 is 512×32×32, the feature size of stage 3 is 1024×16×16, and the feature size of stage 4 is 2048×8×8. The feature difference result can be the feature distance. Accordingly, the feature difference result D1 can be calculated using the following formulas (1) and (2), where the size of the image to be detected is H×W, and [h,w] are the two-dimensional coordinates of the pixel in the image to be detected. This represents the feature of the backbone network (S) at stage k, at position [h, w] (feature size is cx1x1). Let C represent the feature of the teacher network (T) at stage k, position [h, w], where K is the total number of submodules. k (h,w) represents the feature distance between the features of the backbone network and the features of the auxiliary network at stage k, position h,w.
[0084]
[0085] In one embodiment, D1 can be obtained by summing the feature distances calculated from the features in all the stages mentioned above. In another embodiment, D1 can be obtained by summing the feature distances calculated from the features in stage 3 and stage 4, so that the feature difference results fully reflect the differences in deep semantic features and avoid interference from shallow texture features, then k=3 and K=4.
[0086] In other embodiments, instead of calculating the feature difference results for each individual pixel, feature difference calculation is performed only on foreground pixels in the image to be detected. Accordingly, the method also includes a foreground and background segmentation step for the image to be detected, specifically including: performing image segmentation on the image to be detected based on the pixel depth of the depth image to obtain an image segmentation result, which is used to characterize whether a pixel in the image to be detected is a foreground pixel or a background pixel. A depth threshold can be set, and the pixel depth of each pixel in the depth image can be compared with the depth threshold, using the depth threshold as the dividing point to separate the foreground and background in the image. Specifically, foreground pixels can be labeled with a value of 1, and background pixels can be labeled with a value of 0. This achieves the key region selection of the image to be detected, so that the difference analysis is focused on the core entity objects in the image, reducing background interference, thereby reducing the amount of data for analysis and computation, and improving the complexity of anomaly detection and localization calculations.
[0087] Further, S301 can specifically include: extracting the first pixel feature of the foreground pixel from the image features, and extracting the second pixel feature of the foreground pixel from the reference features; performing difference analysis on the first pixel feature and the second pixel feature of the foreground pixel at the same position to obtain the feature difference result. The first pixel feature and the second pixel feature can be the feature map corresponding to the aforementioned pixel, such as the features in stage3 and stage4 mentioned above. The difference analysis method is similar to that mentioned above, except that a background mask constraint is added, which can utilize depth map information while performing feature distillation, so that the difference analysis is performed on the core object in the foreground rather than the background, reducing the amount of computation and resource consumption, and improving the positioning accuracy. Taking the aforementioned backbone network and auxiliary network structure as an example, the feature difference result can be calculated using the aforementioned formula (1) and the following formula (3), where M(h,w) refers to whether the location [h,w] is the foreground, and is 1 if it is the foreground and 0 if it is not the foreground.
[0088]
[0089] Understandably, the training data used during training are all positive samples, i.e., defect-free images. Therefore, the shared backbone network, under the guidance of the teacher model (auxiliary network), learns the ability to extract features from normal images. This results in different feature extraction effects for normal and abnormal pixel regions, meaning the obtained features exhibit differences. Correspondingly, the larger the difference value corresponding to the above feature differences, the higher the probability that the pixel belongs to an abnormal region, and the higher the probability that the image to be detected is abnormal; conversely, the smaller the difference value, the lower the probability.
[0090] The image to be detected and the reconstructed image have the same image size. The reconstruction difference result is used to characterize the degree of difference between the image to be detected and the reconstructed image. Similar to the above, the higher the degree of pixel difference, the higher the probability that the pixel belongs to an abnormal region, and the higher the probability that the image to be detected is abnormal, and vice versa. In some embodiments, the reconstruction difference result includes reconstructed pixel distance and reconstructed feature distance. Accordingly, S303 may include: extracting the first pixel of each pixel in the image to be detected and the second pixel of each pixel in the reconstructed image; performing difference analysis on the first pixel and the second pixel of the same pixel to obtain the reconstructed pixel distance; the reconstructed pixel distance can be the pixel difference between the first pixel and the second pixel; performing feature extraction on the image to be detected and the reconstructed image respectively to obtain the first feature and the second feature; performing feature difference analysis on each pixel based on the first feature and the second feature to obtain the reconstructed feature distance; the reconstructed feature distance is the feature distance between the first feature and the second feature. In this way, by describing the reconstruction difference from the image itself and the feature scale through the reconstructed pixel distance and the reconstructed feature distance, the comprehensiveness of the difference analysis information is improved, thereby improving the accuracy of the pixel difference score.
[0091] In one embodiment, the reconstructed pixel distance D2 and the reconstructed feature distance D3 can be calculated using the following formulas (4) and (5), where I r Represents the reconstructed image, I s Represents the image to be detected, I r (h,w) represents the pixel value (second pixel) at position [h,w] in the reconstructed image. s (h,w) represents the pixel value (first pixel) at position [h,w] in the image to be detected; D2 calculates the distance between the reconstructed image and the image to be detected in terms of pixel value; F i F represents the i-th layer feature in the first and second features, where n is the total number of layers. i (I r (h,w) represents the feature of the i-th layer at position [h,w] in the first feature corresponding to the reconstructed image, F i (I s (h,w) represents the feature of the i-th layer at position [h,w] in the first feature of the image to be detected.
[0092] D2=|I r (h,w)-I s (h,w)| (4)
[0093]
[0094] In other embodiments, unlike the difference analysis of the whole image, pixel filtering is performed through a depth mask to calculate the difference of foreground pixels. Accordingly, S303 may include: extracting the first pixel of the foreground pixel in the image to be detected and the second pixel of the foreground pixel in the reconstructed image; performing difference analysis on the first pixel and the second pixel of the same foreground pixel to obtain the reconstructed pixel distance; the reconstructed pixel distance can be the pixel difference between the first pixel and the second pixel; performing feature extraction on the image to be detected and the reconstructed image respectively to obtain the first feature and the second feature; performing feature difference analysis on the foreground pixels based on the first feature and the second feature to obtain the reconstructed feature distance; the reconstructed feature distance is the feature distance between the first feature and the second feature. In one embodiment, the reconstructed pixel distance D2 can be calculated using the following formula (5), where M refers to whether the pixel at [h,w] is a foreground pixel. If it is, the value is 1; otherwise, it is 0. M can be obtained by using the threshold of the depth map obtained from the acquisition. After adding M for filtering, interference caused by the background can be avoided.
[0095] D2=|(I r (h,w)-I s (h,w))·M(h,w)| (6)
[0096]
[0097] In one embodiment, the reconstruction task branch consists of a backbone network and a reconstruction decoder. The backbone network structure is as described above, and the structure of the reconstruction decoder is basically the same as that of the depth decoder, except that the deconvolution of the last layer is different. In the reconstruction encoder, the output size of the deconvolution of the last layer is 3×256×256.
[0098] Feature extraction for both the image to be detected and the reconstructed image can be achieved using the same feature extraction network. Specifically, it can employ a deep neural network model independent of the aforementioned reconstruction task branch, such as the VGG19 model pre-trained on ImageNet. This allows for reliable difference analysis by using the same network for feature extraction.
[0099] The depth image and the depth estimation image have the same image size. The depth difference result is used to characterize the depth difference between the same location in the depth image and the depth estimation image. Specifically, the depth difference result can include depth distance and gradient difference information. S305 can specifically include: performing pixel distance analysis based on the depth image and the depth estimation image to obtain the depth distance corresponding to the pixel; the depth distance can be the difference between the pixel depth of the same location in the depth image and the pixel depth in the depth estimation image; obtaining the first depth gradient information corresponding to the depth image and the second depth gradient information corresponding to the depth estimation image, the depth gradient information is used to characterize the depth change rate between a single pixel and pixels in a surrounding preset area; performing gradient difference analysis based on the first depth gradient information and the second depth gradient information to obtain the gradient difference information corresponding to the pixel; the gradient difference information is used to characterize the difference between the depth gradient of the same location in the depth image and the depth gradient in the depth estimation image. In this way, the multidimensional difference information between the depth image and the depth estimation image is expressed through depth distance and gradient difference information to highlight the overall distribution difference and edge difference of the entity. The backbone network of the depth estimation module is trained with an auxiliary network as the teacher model and is trained with defect-free positive samples. Therefore, it learns feature extraction for normal images, and the corresponding depth decoder can also learn depth estimation for normal images. Defects and other anomalies have more obvious depth features in depth images. The greater the difference between the depth estimation image output by the depth estimation task branch and the depth image, the higher the probability that the pixel belongs to the abnormal region and the higher the probability that there is an anomaly in the image to be detected, and vice versa.
[0100] In one embodiment, the depth distance D4 and gradient difference information D5 can be calculated using the following formulas (8)-(10), wherein, in formula (8), D p D represents the depth estimation image. g Represents a depth image, D p(h,w) represents the pixel value at location [h,w] in the depth estimation image, D g (h,w) represents the pixel value at position [h,w] in the depth image, and D4 calculates the distance error between the depth estimation image and the depth image; in formula (9), D represents the depth map, * represents the convolution operation, h x It is a convolution kernel that extracts the gradient in the x-direction, which can be, for example, a 3x3 convolution kernel [[1,0,-1],[1,0,-1],[1,0,-1]], h y It is the convolution kernel that extracts the gradient in the y-direction, which can be, for example, a 3x3 convolution kernel [[1,1,1],[0,0,0],[-1,-1,-1]]; in formula (10), g(D p (h,w)) represents the second depth gradient information, and g(D) g (h,w)) represents the first depth gradient information, c is a constant, and its function is to prevent the denominator from being 0. D5 calculates the difference in depth changes between the depth estimation image and the depth image.
[0101] D4=|D p (h,w)-D g (h,w)| (8)
[0102]
[0103] In one embodiment, the depth estimation task branch consists of a backbone network and a depth decoder. The backbone network is shared and has the same structure as in the previous example. The depth decoder consists of four convolutional combinations (Conv2d+BN+ReLU), four bilinear interpolations (UpSample), and one deconvolution. The relevant feature size changes as follows: the output feature map of the backbone network is 2048×8×8. After the first convolutional combination, it becomes 1024×8×8. After the first bilinear interpolation, it becomes 1024×16×16. Then, after the second convolutional combination and the second bilinear interpolation, it becomes 512×32×32. Then, after the third convolutional combination and the third bilinear interpolation, it becomes 256×64×64. Then, after the fourth convolutional combination and the fourth bilinear interpolation, it becomes 64×128×128. Finally, a deconvolution is used, and the size becomes 1×256×256, resulting in the depth estimated image.
[0104] In some embodiments, such as Figure 5 As shown, the 512×32×32 feature map output after the second convolution combination and the second bilinear interpolation will be added to the 512×32×32 output of the backbone network stage2, so that the depth decoder can obtain the shallow detail features of the image to be detected, thereby improving the accuracy of depth estimation.
[0105] The obtained feature difference results, reconstructed pixel distance, reconstructed feature distance, depth distance, and gradient difference information are statistically processed to obtain an anomaly evaluation index for each pixel. Specifically, the anomaly evaluation index can be a weighted sum or weighted average of the feature difference results, reconstructed pixel distance, reconstructed feature distance, depth distance, and gradient difference information. Understandably, the larger the anomaly evaluation index, the greater the probability that the pixel is an anomaly, i.e., the greater the probability of an anomaly in the image to be detected, and vice versa.
[0106] In one embodiment, the anomaly evaluation index D can be calculated using the following formula (11), where a, b, and d are constants, for example, a is 2 and b and d are 1.
[0107] D= aD1+b(D2+D3)+d (D4+D5) (11)
[0108] Understandably, when the feature difference results and reconstruction difference results are obtained through the background mask, the values of the feature difference results and reconstruction difference results of the corresponding background pixels are all set to 0.
[0109] In conjunction with the above scheme, in some embodiments, a normal vector estimation task branch can be added, consisting of a backbone network and a normal vector decoder. During the detection process, it is also necessary to acquire the normal vector image corresponding to the image to be detected. The two images have the same size, and pixels at the same position represent the same location of the entity. The normal vector image can be generated based on the image acquired by the normal vector camera. Similar to the aforementioned depth estimation task branch, the normal vector decoder in this branch is used to input image features to perform normal vector map estimation, obtaining a normal vector estimation image. Then, the normal vector estimation image and the normal vector image are compared to obtain the normal vector difference result. It is understood that the normal vector difference result can include pixel normal vector distance and normal vector gradient difference information, and the acquisition methods are similar to those of the aforementioned depth distance and gradient difference information, respectively, and will not be repeated here. The model structure of the target network model in this application can be extended based on actual needs, with good generalization performance and wide applicability.
[0110] S207: Generate anomaly detection results for the image to be detected based on the anomaly response map.
[0111] In this embodiment, the anomaly detection result is used to characterize whether the image to be detected is an abnormal image or a normal image, and to characterize the location of abnormal regions in the image to be detected. The anomaly detection result may include an image category and a binary image. The image category includes normal and abnormal. In the binary image, regions with a pixel value of 1 are abnormal regions, and regions with a pixel value of 0 are normal regions. Accordingly, after obtaining the anomaly response map, the maximum value among its various anomaly evaluation indicators is determined. The maximum value is compared with an indicator threshold. If it exceeds the indicator threshold, it indicates that the image category of the image to be detected is abnormal; otherwise, it is normal. The pixels corresponding to the anomaly evaluation indicators that exceed the indicator threshold are determined as abnormal regions.
[0112] In summary, by acquiring depth images and performing multi-task detection—including multi-feature extraction, depth estimation, and image reconstruction—on the network model, depth information and image features can be fully combined to compensate for the limitations of single feature information in the image. Multi-task output improves the dimensionality of anomaly detection information and the accuracy of entity feature representation, thereby increasing anomaly detection precision and reducing the false negative rate. In practical applications, this application can effectively identify unclear defect structures on 2D cameras, such as detecting pits that are almost indistinguishable from normal samples in color. Furthermore, the proposed method of hybrid training and anomaly detection using depth estimation, image reconstruction, and feature distillation, where the three tasks share a backbone network and different branches complete their respective specific tasks, is adaptable to anomaly detection tasks under RGBD data. For example, in industrial scenarios where RGB and depth maps can be acquired simultaneously, defect detection of industrial products can be achieved using only defect-free data for training, significantly improving detection performance.
[0113] Based on some or all of the above implementation methods, please refer to Figure 6 This application also provides a model training method, which may specifically include the following steps S401-S407:
[0114] S401: Obtain the training set, which includes multiple sample image pairs. Each sample image pair includes a positive sample image and the corresponding sample depth image.
[0115] Specifically, the positive sample images are similar to the aforementioned images to be detected, and the sample depth images are similar to the aforementioned depth images. All sample image pairs are positive sample pairs, meaning that the positive sample images are all normal images without any abnormal regions, and the corresponding sample depth images are also normal images.
[0116] In one embodiment, training can use 500 pairs of defect-free data, namely 500 RGB images and corresponding 500 depth images. Before inputting the data into the network for training, all images are scaled to a size of 256×256 pixels, and the pixel values are normalized to 0 to 1.
[0117] S403: Input the positive sample image into the preset neural network for feature extraction, depth estimation and image reconstruction to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image.
[0118] Specifically, the network structure of the pre-defined neural network is similar to that described above; see the specific example below. Figure 5 The sample image features, sample reference features, sample depth estimation images, and sample reconstructed images are similar to the aforementioned image features, reference features, depth estimation images, and reconstructed images, and will not be described again.
[0119] S405: The loss is calculated based on the positive sample image, sample depth image, sample image features, sample reference features, sample depth estimation image, and sample reconstruction image to obtain the model loss.
[0120] S407: Train a preset neural network based on model loss to obtain the target network model.
[0121] Specifically, the model parameters of a pre-defined neural network are updated based on model loss to perform iterative training until the training termination condition is met. The updated pre-defined neural network model that meets the training termination condition is then identified as the target network model. Understandably, during an iteration, shared network parameters, such as the backbone network parameters, can be updated based on model loss, and the specific parameters of each branch can be updated based on the loss of each branch, such as updating the deep decoder network parameters based on the second loss and updating the reconstruction decoder network parameters based on the third loss. The training termination condition can be met when the difference in model loss between two iterations is less than or equal to the model loss threshold, or when the number of iterations reaches a pre-defined number of iterations. Understandably, by setting up a backbone network and corresponding teacher models (auxiliary networks), as well as a deep decoder and a reconstruction decoder, joint training of distillation learning, depth estimation, and reconstruction is performed to achieve information sharing and leverage the strengths of each. The foreground and background information extracted from the depth map can assist the learning of distillation and reconstruction tasks, and the distillation and reconstruction tasks synergistically help the learning of the depth estimation task. Furthermore, the efficiency of the hybrid training framework is significantly higher than that of three independent task models, both in anomaly detection and training.
[0122] In some embodiments, S405 may include the following S4051-S4054:
[0123] S4051: Calculate the distillation loss based on the sample image features and sample reference features to obtain the first loss.
[0124] Specifically, the calculation of distillation loss is related to the calculation of the aforementioned feature difference results. For the entire positive sample image, the feature distances of each pixel are summed and averaged to obtain the first loss. In one embodiment, the first loss L... KDIt can be calculated using the aforementioned formula (1) and the following formula (12), where the size of the positive sample image is H×W, [h,w] is the two-dimensional coordinate of the pixel in the positive sample image, and H k W is the size of the positive sample image in the H direction. k H is the size of the positive sample image in the W direction. k W k This represents the total number of pixels in the positive sample image. In another embodiment, foreground optimization is performed by incorporating depth information to avoid background interference; correspondingly, the first loss L... KD It can be calculated using the following formula (12). In this way, while distilling the image features, the depth map information can be used, so that the model can focus on the core object rather than the background during distillation learning.
[0125]
[0126] Distillation loss can be calculated based on deep features, i.e. semantic features. In the ResNet50 structure, stage3 and stage4 features can be used for calculation, so k=3 and K=4.
[0127] S4052: Calculate the depth estimation loss based on the depth image and the sample depth estimation image to obtain the second loss.
[0128] Specifically, the loss function of the second loss may include depth distance loss and gradient similarity loss. The calculation of depth estimation loss is related to the calculation of the aforementioned depth difference results. For the entire sample depth image, the depth distance and gradient difference information of each pixel are obtained, summed, and averaged to obtain the depth distance loss and gradient similarity loss. In one embodiment, the depth distance loss L... d The gradient similarity loss L can be calculated using the following formula (13) or (14). g It can be calculated using the following formula (15), where D p The image representing the depth estimation of the sample, D g Represents the depth image of the sample.
[0129] L d =|D p -D g | (13)
[0130]
[0131] S4053: Calculate the reconstruction loss based on the positive sample image and the reconstructed sample image to obtain the third loss.
[0132] Specifically, the loss function of the third loss can include reconstruction loss and perceptual loss. The calculation of the third loss is related to the calculation of the reconstruction difference results mentioned above. For the entire positive sample image, the reconstructed pixel distance and reconstructed feature distance of each pixel are summed and averaged to obtain the reconstruction loss and perceptual loss. In one embodiment, the reconstruction loss L r The perceived loss L can be calculated using the following formula (16) or (17). p It can be calculated using the aforementioned formula (18), where I r To reconstruct the image for the sample, I s For positive sample images, L p Essentially, it represents the similarity between the reconstructed image and the positive sample image at the feature level.
[0133] L r =|I r -I s | (16)
[0134]
[0135] In another embodiment, foreground optimization is performed using depth information to avoid background interference; correspondingly, the reconstruction loss L... r The perceived loss L can be calculated using the following formula (19) or (20). p It can be calculated using the aforementioned formula (21).
[0136] L r =|(I r -I s )·M|(19)
[0137]
[0138] S4054: The model loss is obtained by fusing the first loss, the second loss, and the third loss.
[0139] Specifically, the fusion here can be a weighted summation operation. In one embodiment, the model loss L can be calculated using the following formula (22). For example, a is 2, and b and c are 1.
[0140] L = a L KD +b(L d + L g) +d (L r + L p ) (twenty two)
[0141] In this way, by training the neural network with features that can characterize the feature extraction error, depth estimation error, and reconstruction error of multi-task branches, it can learn and optimize the feature extraction, depth estimation, and reconstruction knowledge of normal data under the guidance of the teacher model, thereby widening the gap in its relevant output for abnormal data and improving the accuracy of anomaly detection.
[0142] In some embodiments, the auxiliary network can be a classification model pre-trained using ImageNet, and the weights of the teacher network are fixed and not updated throughout the training process of the anomaly detection model.
[0143] This application proposes a framework for training three task branches in a hybrid manner. During training, the three tasks can be trained simultaneously, or the number of specific tasks can be adjusted in a single training session, such as performing only depth estimation and reconstruction training, or performing only depth estimation and feature distillation training. The resulting target network model has excellent detection performance for inconspicuous defects in images such as RGB in anomaly detection applications, and the false negative rate is significantly reduced.
[0144] This application embodiment also provides an anomaly detection device 700, such as Figure 7 As shown, Figure 7 The diagram shows a structural schematic of an anomaly detection device provided in an embodiment of this application. The device may include the following modules.
[0145] Image acquisition module 11: used to acquire the image to be detected and the corresponding depth image;
[0146] Detection module 12: Used to input the image to be detected into the target network model for feature extraction, depth estimation and image reconstruction, to obtain image features, reference features, depth estimated image and reconstructed image;
[0147] Anomaly Analysis Module 13: Used to perform image anomaly analysis based on the image to be detected, depth image, image features, reference features, reconstructed image, and depth estimation image to obtain anomaly response map;
[0148] Result generation module 14: Used to generate anomaly detection results for the image to be detected based on the anomaly response map.
[0149] In some embodiments, the target network model includes a backbone network, an auxiliary network, a deep decoder, and a reconstruction decoder, with the auxiliary network being the teacher model corresponding to the backbone network; the detection module 12 may include:
[0150] The first extraction submodule is used to input the image to be detected into the backbone network for feature extraction to obtain image features;
[0151] The second extraction submodule is used to input the image to be detected into the auxiliary network for feature extraction to obtain reference features;
[0152] Depth estimation submodule: Used to input image features into the depth decoder for depth estimation, and obtain a depth-estimated image;
[0153] The reconstruction submodule is used to input image features into the reconstruction decoder to reconstruct the image and obtain the reconstructed image.
[0154] In some embodiments, the apparatus further includes:
[0155] Sample acquisition module: used to acquire the training set, which includes multiple sample image pairs. Each sample image pair includes a positive sample image and the corresponding sample depth image.
[0156] Neural network module: used to input positive sample images into a preset neural network for feature extraction, depth estimation and image reconstruction, to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image;
[0157] Loss calculation module: used to calculate the model loss based on positive sample images, sample depth images, sample image features, sample reference features, sample depth estimation images, and sample reconstructed images;
[0158] Training module: Used to train a preset neural network based on model loss to obtain the target network model.
[0159] In some embodiments, the anomaly analysis module 13 includes:
[0160] Feature Analysis Submodule: Used to perform pixel feature difference analysis based on image features and reference features to obtain the feature difference results corresponding to the pixels in the image to be detected;
[0161] The depth analysis submodule is used to perform pixel depth difference analysis on depth images and depth estimation images to obtain the depth difference results corresponding to the pixels.
[0162] Pixel Analysis Submodule: Used to perform pixel difference analysis on the image to be detected and the reconstructed image, and obtain the reconstruction difference results corresponding to the pixels;
[0163] Evaluation index determination submodule: used to determine the anomaly evaluation index of each pixel in the image to be detected based on feature difference results, depth difference results, and reconstruction difference results;
[0164] Response map generation submodule: Used to generate anomaly response maps based on anomaly evaluation metrics for each pixel.
[0165] In some embodiments, the reconstruction difference results include reconstructed pixel distance and reconstructed feature distance; the pixel analysis submodule may include:
[0166] First extraction unit: used to extract the first pixel of each pixel in the image to be detected and the second pixel of each pixel in the reconstructed image;
[0167] First pixel difference analysis unit: used to perform difference analysis on the first pixel and the second pixel of a pixel point at the same position to obtain the reconstructed pixel distance;
[0168] First feature extraction unit: used to extract features from the image to be detected and the reconstructed image respectively to obtain first feature and second feature;
[0169] First feature difference analysis unit: used to perform feature difference analysis on each pixel based on the first feature and the second feature, and obtain the reconstructed feature distance.
[0170] In some embodiments, the apparatus further includes an image segmentation module: used to perform image segmentation on the image to be detected based on the pixel depth of the depth image, to obtain an image segmentation result, the image segmentation result being used to characterize whether the pixels in the image to be detected are foreground pixels or background pixels;
[0171] Accordingly, the feature analysis submodule includes:
[0172] Pixel feature extraction unit: used to extract the first pixel feature of the foreground pixel from the image features, and to extract the second pixel feature of the foreground pixel from the reference features;
[0173] Feature difference analysis unit: used to perform difference analysis on the first pixel feature and the second pixel feature of the same foreground pixel point to obtain feature difference results.
[0174] In some embodiments, the pixel analysis submodule includes:
[0175] Pixel extraction unit: used to extract the first pixel of the foreground pixels in the image to be detected and the second pixel of the foreground pixels in the reconstructed image;
[0176] Pixel difference analysis unit: used to perform difference analysis on the first and second pixels of the same foreground pixel point to obtain the reconstructed pixel distance;
[0177] Feature extraction unit: used to extract features from the image to be detected and the reconstructed image respectively, to obtain the first feature and the second feature;
[0178] Reconstruction Difference Analysis Unit: Used to perform feature difference analysis on foreground pixels based on the first feature and the second feature, and obtain the reconstructed feature distance.
[0179] In some embodiments, the depth difference results include depth distance and gradient difference information; the depth analysis submodule may include:
[0180] Pixel distance analysis unit: used to perform pixel distance analysis based on depth image and depth estimation image to obtain the depth distance corresponding to the pixel;
[0181] Gradient information acquisition unit: used to acquire the first depth gradient information corresponding to the depth image and the second depth gradient information corresponding to the depth estimation image;
[0182] Gradient difference analysis unit: used to perform gradient difference analysis based on the first depth gradient information and the second depth gradient information to obtain the gradient difference information corresponding to the pixel.
[0183] This application also provides a model training device 800, such as... Figure 8 As shown, Figure 8 The diagram shows a structural schematic of a model training device provided in an embodiment of this application. The device may include the following modules.
[0184] Sample acquisition module 21: used to acquire training set, which includes multiple sample image pairs, and each sample image pair includes a positive sample image and a sample depth image corresponding to the positive sample image;
[0185] Neural network module 22: used to input positive sample images into a preset neural network for feature extraction, depth estimation and image reconstruction, to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image;
[0186] Loss calculation module 23: used to calculate the model loss based on positive sample images, sample depth images, sample image features, sample reference features, sample depth estimation images, and sample reconstructed images;
[0187] Training module 24: Used to train a preset neural network based on model loss to obtain the target network model.
[0188] In some embodiments, training module 24 may include:
[0189] First Loss Submodule: Used to calculate the distillation loss based on sample image features and sample reference features to obtain the first loss;
[0190] The second loss submodule is used to calculate the depth estimation loss based on the depth image and the sample depth estimation image to obtain the second loss.
[0191] The third loss submodule is used to calculate the reconstruction loss based on the positive sample image and the reconstructed sample image, and obtain the third loss.
[0192] Loss fusion submodule: used to fuse the first loss, second loss and third loss to obtain the model loss.
[0193] It should be noted that the above-described device embodiments and method embodiments are based on the same implementation methods.
[0194] This application provides an anomaly detection device. The scheduling device can be a terminal or a server, including a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or at least one program is loaded and executed by the processor to implement the anomaly detection method provided in the above method embodiments.
[0195] Memory is used to store software programs and modules. The processor executes these stored software programs and modules to perform various functional applications and detect anomalies. Memory can primarily consist of a program storage area and a data storage area. The program storage area stores the operating system, application programs required for functionality, etc.; the data storage area stores data created based on device usage, etc. Furthermore, memory can include high-speed random access memory (RAM) and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.
[0196] The methods and embodiments provided in this application can be executed in electronic devices such as mobile terminals, computer terminals, servers, or similar computing devices. Figure 9 This is a hardware structure block diagram of an electronic device for an anomaly detection method provided in an embodiment of this application. For example... Figure 9 As shown, the electronic device 900 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 910 (CPUs 910 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 930 for storing data, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage media 920 may be temporary or persistent storage. The program stored in the storage media 920 may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the CPU 910 may be configured to communicate with the storage media 920 and execute a series of instruction operations in the storage media 920 on the electronic device 900. The electronic device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server. TM Mac OS XTM Unix TM Linux™, FreeBSD™, etc.
[0197] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 940 may be a radio frequency (RF) module for wireless communication with the Internet.
[0198] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 900 may also include... Figure 9 The more or fewer components shown, or having the same Figure 9 The different configurations shown.
[0199] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing an anomaly detection method in the method embodiments. The at least one instruction or the at least one program is loaded and executed by the processor to implement the anomaly detection method provided in the above method embodiments.
[0200] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0201] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.
[0202] The anomaly detection method, model training method, apparatus, device, storage medium, server, terminal, and program product provided in this application, as described above, enable the following technical solutions: The image to be detected is input into a target network model for feature extraction, depth estimation, and image reconstruction, resulting in image features, reference features, a depth-estimated image, and a reconstructed image. Based on the image to be detected, its corresponding depth image, image features, reference features, reconstructed image, and depth-estimated image, image anomaly analysis is performed to obtain an anomaly response map. Anomaly detection results for the image to be detected are generated based on the anomaly response map. By acquiring a depth image and performing multi-task detection of multi-feature extraction, depth estimation, and image reconstruction of the image to be detected within the network model, depth information and image features can be fully combined to compensate for the limitations of single feature information in the image to be detected. Multi-task output improves the dimensionality of anomaly detection information and the accuracy of entity feature representation, thereby increasing anomaly detection accuracy and reducing the false negative rate.
[0203] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0204] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device, equipment, and storage medium embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0205] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing the relevant hardware to implement them. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0206] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An anomaly detection method, characterized in that, The method includes: Obtain the image to be detected and the corresponding depth image of the image to be detected; The image to be detected is input into the target network model for feature extraction, depth estimation, and image reconstruction to obtain image features, reference features, depth-estimated image, and reconstructed image. Pixel feature difference analysis is performed based on the image features and the reference features to obtain the feature difference results corresponding to the pixels in the image to be detected. Pixel difference analysis is performed on the image to be detected and the reconstructed image to obtain the reconstruction difference results corresponding to the pixels; Perform pixel depth difference analysis on the depth image and the depth estimation image to obtain the depth difference results corresponding to the pixel points; Based on the feature difference results, the depth difference results, and the reconstruction difference results, an anomaly evaluation index for each pixel in the image to be detected is determined. An anomaly response map is generated based on the anomaly evaluation index of each pixel. The anomaly response map is used to characterize the abnormal distribution or abnormal region of the image to be detected. Anomaly detection results for the image to be detected are generated based on the anomaly response map.
2. The method according to claim 1, characterized in that, The target network model includes a backbone network, an auxiliary network, a depth decoder, and a reconstruction decoder, wherein the auxiliary network is the teacher model corresponding to the backbone network; the step of inputting the image to be detected into the target network model for feature extraction, depth estimation, and image reconstruction to obtain image features, reference features, a depth-estimated image, and a reconstructed image includes: The image to be detected is input into the backbone network for feature extraction to obtain the image features; The image to be detected is input into the auxiliary network for feature extraction to obtain the reference features; The image features are input into the depth decoder for depth estimation to obtain the depth-estimated image; The image features are input into the reconstruction decoder to reconstruct the image, resulting in the reconstructed image.
3. The method according to claim 1, characterized in that, The target network model was trained using the following method: Obtain a training set, which includes multiple pairs of sample images, each pair of sample images including a positive sample image and a sample depth image corresponding to the positive sample image; The positive sample image is input into a preset neural network for feature extraction, depth estimation and image reconstruction to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image; The model loss is calculated based on the positive sample image, the sample depth image, the sample image features, the sample reference features, the sample depth estimation image, and the sample reconstruction image. The preset neural network is trained based on the model loss to obtain the target network model.
4. The method according to any one of claims 1-3, characterized in that, The anomaly evaluation index is the anomaly score of each pixel, and the anomaly response map is a heatmap generated based on the anomaly score of each pixel in the image to be detected.
5. The method according to claim 1, characterized in that, The reconstructed difference results include reconstructed pixel distance and reconstructed feature distance; the pixel difference analysis of the image to be detected and the reconstructed image to obtain the reconstructed difference results corresponding to the pixels includes: Extract the first pixel of each pixel in the image to be detected and the second pixel of each pixel in the reconstructed image; The reconstructed pixel distance is obtained by performing a difference analysis on the first and second pixels of pixels at the same location; Feature extraction is performed on the image to be detected and the reconstructed image respectively to obtain the first feature and the second feature; Based on the first feature and the second feature, feature difference analysis is performed on each pixel to obtain the reconstructed feature distance.
6. The method according to claim 1, characterized in that, The method further includes: The image to be detected is segmented based on the pixel depth of the depth image to obtain an image segmentation result. The image segmentation result is used to characterize whether the pixels in the image to be detected are foreground pixels or background pixels. The step of performing pixel feature difference analysis based on the image features and the reference features to obtain the feature difference results corresponding to the pixels in the image to be detected includes: Extract the first pixel feature of the foreground pixel from the image features, and extract the second pixel feature of the foreground pixel from the reference features; The difference analysis is performed on the first pixel feature and the second pixel feature of the foreground pixel at the same position to obtain the feature difference result.
7. The method according to claim 1, characterized in that, The method further includes: The image to be detected is segmented based on the pixel depth of the depth image to obtain an image segmentation result. The image segmentation result is used to characterize whether the pixels in the image to be detected are foreground pixels or background pixels. The step of performing pixel difference analysis on the image to be detected and the reconstructed image to obtain the reconstruction difference results corresponding to the pixels includes: Extract the first pixel of the foreground pixel in the image to be detected and the second pixel of the foreground pixel in the reconstructed image; Perform a difference analysis on the first and second pixels of the foreground pixels at the same location to obtain the reconstructed pixel distance; Feature extraction is performed on the image to be detected and the reconstructed image respectively to obtain the first feature and the second feature; Based on the first feature and the second feature, feature difference analysis of the foreground pixels is performed to obtain the reconstructed feature distance.
8. The method according to claim 1, characterized in that, The depth difference result includes depth distance and gradient difference information; the step of performing pixel depth difference analysis on the depth image and the depth estimation image to obtain the depth difference result corresponding to the pixel includes: Pixel distance analysis is performed based on the depth image and the depth estimation image to obtain the depth distance corresponding to the pixel. Obtain the first depth gradient information corresponding to the depth image and the second depth gradient information corresponding to the depth estimation image; Gradient difference analysis is performed based on the first depth gradient information and the second depth gradient information to obtain the gradient difference information corresponding to the pixel.
9. A model training method, characterized in that, The method includes: Obtain a training set, which includes multiple pairs of sample images, each pair of sample images including a positive sample image and a sample depth image corresponding to the positive sample image; The positive sample image is input into a preset neural network for feature extraction, depth estimation and image reconstruction to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image; The model loss is calculated based on the positive sample image, the sample depth image, the sample image features, the sample reference features, the sample depth estimation image, and the sample reconstruction image. The preset neural network is trained based on the model loss to obtain a target network model, which serves as the target network model in the anomaly detection method as described in any one of claims 1-8.
10. The method according to claim 9, characterized in that, The loss calculation based on the positive sample image, the sample depth image, sample image features, sample reference features, sample depth estimation image, and sample reconstruction image yields the model loss, which includes: Distillation loss is calculated based on the sample image features and the sample reference features to obtain the first loss; The depth estimation loss is calculated based on the depth image and the sample depth estimation image to obtain the second loss; The reconstruction loss is calculated based on the positive sample image and the reconstructed sample image to obtain the third loss; The model loss is obtained by fusing the first loss, the second loss, and the third loss.
11. An anomaly detection device, characterized in that, The device includes: Image acquisition module: used to acquire the image to be detected and the corresponding depth image; Detection module: used to input the image to be detected into the target network model for feature extraction, depth estimation and image reconstruction, to obtain image features, reference features, depth estimated image and reconstructed image; Anomaly analysis module: used to perform image anomaly analysis based on the image to be detected, the depth image, the image features, the reference features, the reconstructed image, and the depth estimation image, and obtain an anomaly response map; Result generation module: used to generate anomaly detection results for the image to be detected based on the anomaly response map; The anomaly analysis module includes: Feature Analysis Submodule: Used to perform pixel feature difference analysis based on image features and reference features to obtain the feature difference results corresponding to the pixels in the image to be detected; The depth analysis submodule is used to perform pixel depth difference analysis on depth images and depth estimation images to obtain the depth difference results corresponding to the pixels. Pixel Analysis Submodule: Used to perform pixel difference analysis on the image to be detected and the reconstructed image, and obtain the reconstruction difference results corresponding to the pixels; Evaluation index determination submodule: used to determine the anomaly evaluation index of each pixel in the image to be detected based on feature difference results, depth difference results, and reconstruction difference results; The response map generation submodule is used to generate an anomaly response map based on the anomaly evaluation index of each pixel. The anomaly response map is used to characterize the abnormal distribution or abnormal region of the image to be detected.
12. A model training device, characterized in that, The device includes: Sample acquisition module: used to acquire a training set, which includes multiple sample image pairs, and each sample image pair includes a positive sample image and a sample depth image corresponding to the positive sample image; Neural network module: used to input the positive sample image into a preset neural network for feature extraction, depth estimation and image reconstruction, to obtain sample image features, sample reference features, sample depth estimation image and sample reconstruction image; Loss calculation module: used to calculate the model loss based on the positive sample image, the sample depth image, the sample image features, the sample reference features, the sample depth estimation image, and the sample reconstruction image; Training module: used to train the preset neural network based on the model loss to obtain the target network model, wherein the target network model is the target network model in the anomaly detection method as described in any one of claims 1-8.
13. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the anomaly detection method as described in any one of claims 1-8 or the model training method as described in any one of claims 9-10.
14. A computer device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the anomaly detection method as described in any one of claims 1-8 or the model training method as described in any one of claims 9-10.
Citation Information
Patent Citations
Airfield pavement defect detection and state evaluation method
CN114882367A