Method and apparatus for image anomaly detection and computer device
The multi-stage neural network method for image anomaly detection addresses the limitations of manual inspection by enabling rapid and accurate detection and classification of defects in industrial settings, enhancing efficiency and reducing deployment cycles.
Patent Information
- Application Number
- PCT/CN2023/127824
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Conventional image anomaly detection in industrial settings relies heavily on manual visual inspection, which is labor-intensive, prone to human error, and limited in accuracy, making it difficult to efficiently detect and classify defects in a timely manner.
A method and apparatus for image anomaly detection using a multi-stage neural network approach, where a first neural network identifies images and trains a second neural network based on abnormal images, and a third neural network is trained for precise anomaly location, enabling progressive visual detection and optimization.
This approach allows for rapid deployment of a vision system, reduces manual annotation requirements, and enables efficient detection and classification of multiple types of defects, thereby improving detection accuracy and reducing deployment cycles.
Smart Images

Figure CN2023127824_08052025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR IMAGE ANOMALY DETECTION AND COMPUTER DEVICETECHNICAL FIELD
[0001] This application relates to the field of image recognition, and in particular, to a method and apparatus for image anomaly detection, a computer device, and a storage medium.BACKGROUND
[0002] In recent years, with the development of an artificial intelligence technology and the increase in the demand for industrial automation, a machine vision system based on a deep learning technology becomes more and more widely used in an industrial site, and product appearance detection is more frequently used. The appearance detection refers to detection used for confirming a foreign matter, defect, and blemish on a surface of a part and product, which can prevent inflow of a defective product. Conventional appearance detection mainly relies on manual visual detection, but the manual visual detection has an accuracy limitation, consumes a lot of labor costs, and may also lead to an accuracy deviation and a human error due to personal differences.SUMMARY
[0003] This summary is provided to introduce some concepts being selected in a simplified form that are further described below in the detailed description. This summary is not intended to identify any key feature or essential feature of the claimed subject, nor is intended to be used for assisting in determining the scope of the claimed subject.
[0004] Based on this, this application discloses a method for image anomaly detection, including: identifying an image by using a first neural network; training a second neural network based on an identified abnormal image; using the first neural network and the second neural network jointly to identify a category of the abnormal image and label the abnormal image with a box; training a third neural network by using the category of the image and the labeling box; and detecting an anomaly location of the image by using the first neural network, the second neural network, and the third neural network.
[0005] According to the foregoing method, step-by-step progressive visual detection can be achieved, a deployment cycle can be shortened, and detection of a plurality of types of defections can be completed through optimization iteration.
[0006] Further, the identifying an image by using a first neural network, training a second neural network based on an identified abnormal image includes: identifying an image by using an autoencoder neural network; and training a convolutional neural network based on an identified abnormal image.
[0007] Through the foregoing method, a trained convolutional neural network may be used for identifying a type of image anomaly, to further perform detailed classification.
[0008] Further, the identifying an image by using an autoencoder neural network includes: identifying an image by using the autoencoder neural network and scoring based on a similarity.
[0009] Through the foregoing method, a normal image and an abnormal image can be identified by scoring.
[0010] Further, the training a third neural network by using the category of the image and the labeling box includes:
[0011] training a detection neural network by using the category of the image and the labeling box.
[0012] Through the foregoing method, a trained neural network may be used for identifying a type of image anomaly, and locating an abnormal part.
[0013] This application also discloses an apparatus for image anomaly detection, including: a first neural network module, configured to identify an image by using a first neural network, and train a second neural network based on an identified abnormal image; a second neural network module, configured to use the first neural network and the second neural network jointly to identify a category of the abnormal image and label the abnormal image with a box, and train a third neural network by using the category of the image and the labeling box; and a third neural network module, configured to detect an anomaly location of the image by using the first neural network, the second neural network, and the third neural network.
[0014] This application also provides a computer device, including a memory and a processor. The memory stores a computer program. The foregoing method is implemented when the processor executes the computer program.
[0015] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The foregoing method is implemented when the computer program is executed by a processor.
[0016] This application also provides a computer program product. The computer program product is tangibly stored on a computer-readable medium and includes computer-executable instructions. At least one processor is enabled to perform the foregoing method when the computer-executable instructions are executed.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Implementation of the present disclosure are described in a form of example and not in a form of limitation in the accompanying drawings. Similar reference numerals in the accompanying drawings indicate same or similar parts.
[0018] FIG. 1 is a schematic diagram of a data flow of image anomaly detection according to an embodiment of this application.
[0019] FIG. 2 is a schematic flowchart of a method for image anomaly detection according to an embodiment of this application.
[0020] FIG. 3 is a schematic flowchart of a method for image anomaly detection according to an embodiment of this application.
[0021] FIG. 4 is a schematic flowchart of a method for image anomaly detection according to an embodiment of this application.
[0022] FIG. 5 is a schematic diagram of an apparatus for image anomaly detection according to an embodiment of this application.
[0023] FIG. 6 is a schematic diagram of a computer device for image anomaly detection according to an embodiment of this application.
[0024] Reference numerals are as follows:
[0025] S401 to S403: steps
[0026] 500: apparatus
[0027] 501: module
[0028] 502: module
[0029] 503: module
[0030] 504: module
[0031] 600: computer device
[0032] 602: processor
[0033] 604: memoryDETAILED DESCRIPTION
[0034] Numerous specific details are set forth in the following specification for a purpose of explanation. However, it may be understood that this application may be implemented without these specific details. In other examples, a known circuit, structure, and technology are not disclosed in detail, not to influence understanding on the specification.
[0035] References throughout this specification to "an implementation, ""implementation, ""an example implementation, ""some implementations, ""various implementations, "and the like, mean that a described implementation of this application may include particular features, structures, or characteristics. However, this does not imply that each implementation needs to contain these specific features, structures, or characteristics. In addition, some implementations may have some, all, or none of the features described for other implementations.
[0036] As mentioned in the background, there are also some solutions in the conventional technology. For example, a machine vision system based on deep learning can implement standardized continuous automatic detection, to increase product output. In a process of building a machine vision system, defects first need to be classified according to a standard. Then an image acquisition solution is designed. Finally, an algorithm is designed and implemented. In this process, a large amount of data is needed for testing. In a solution based on deep learning, a large amount of data is needed for training. It may be learned from the above that although the machine vision system based on deep learning can detect product defects more efficiently and accurately, the machine vision system needs a lot of preliminary preparation work and a long deployment cycle. Due to the complexity of an industrial manufacturing production process, there are many kinds of product defects. These factors cause an ordinary machine vision system to be unable to launch quickly to generate economic benefits, and it is difficult to iterate after deployment.
[0037] As mentioned above, generally, a defect detection system based on the deep learning technology may select a neural network for a specific defect detection task, and then the neural network is deployed after performing processes such as data collection, data annotation, and model training. An implementation scenario of this application may be that a step-by-step progressive visual detection system is provided in this application, a deployment cycle can be shortened, and detection of a plurality of types of defections can be completed through optimization iteration. In other words, this application provides a step-by-step progressive defect detection method based on an artificial intelligence technology, which can be launched quickly to generate benefits, and a step-by-step progressive method is used to continuously improve a perception ability of a visual system.
[0038] As shown in FIG. 4, specifically, this application discloses a method for image anomaly detection, including:
[0039] S401: Identify an image by using a first neural network, and train a second neural network based on an identified abnormal image.
[0040] In some embodiments, the first neural network may be an autoencoder neural network.
[0041] Generally, an acquired normal sample is first used to train an autoencoder neural network based on a memory mechanism without manual annotation, and then the network is deployed to identify anomalies through a similarity scoring mechanism in a grid.
[0042] Specifically, a normal or standard image is used to train the autoencoder neural network at this time, so that the autoencoder neural network has a function of identifying normal and abnormal images.
[0043] Further, in some embodiments, the autoencoder neural network has a function of recovering (recover) an identified image. After a recovered image is obtained, compared with an image previously trained and learned, it is determined whether an image to be identified is normal image or abnormal image. A recovered image can be compared with a trained image by using the similarity scoring. This further serves as evidence for determining of the autoencoder neural network.
[0044] In some embodiments, the second neural network may be a convolutional neural network (convolutional neural network) .
[0045] After the system runs for a period of time, some defective and abnormal data is identified. The data can be manually annotated after manual verification. Next, the convolutional neural network is trained by using the data and then deployed into a system. The system includes two neural networks that identify a defect jointly.
[0046] S402: Use the first neural network and the second neural network jointly to identify a category of the abnormal image and label the abnormal image with a box; and train a third neural network by using the category of the image and the labeling box.
[0047] In some embodiments, the third neural network may be a target detection neural network (target detection neural network) .
[0048] Specifically, when the system captures enough defect samples, the system further identifies the defect category (category) and labels data required for a defect location within a two-dimensional bounding box. This part of data can then be used for training the target detection neural network, then the detection neural network is deployed to locate and detect a defective domain (defective domain) .
[0049] S403: Detect an anomaly location of the image by using the first neural network, the second neural network, and the third neural network.
[0050] Specifically, this can not only determine a defect sample with an unknown category, but also classify a defect with a known category, and can also implement a detection task for determining the defect location. Therefore, by using joint determining of the first neural network, the second neural network, and the third neural network, identification of a category and generation of location coordinates of the foregoing defect can be completed.
[0051] Differences between this application and other solutions, or the technical effects of this application are:
[0052] 1. Neural networks with different tasks are added in stages to form a quickly deployed vision system to achieve a defect detection goal for different task categories.
[0053] 2. The system is quickly deployed and the abnormal sample is identified without manual annotation by using the autoencoder neural network based on the memory mechanism and a partition scoring mechanism.
[0054] 3. Image data is acquired based on an identification result of the autoencoder neural network, manual classification is performed, a data annotation man-hour is reduced, and defect category is identified by training the convolutional neural network. By using two neural networks at a same time, a defect with an unknown category can be identified, and a defect with a known category may be classified.
[0055] 4. Classification data is labeled with the two-dimensional bounding box, and the target detection neural network is trained and deployed into a system, so that the system can not only identify a defect category but also locate some of defects.
[0056] System structure
[0057] FIG. 1 is a diagram of a topological structure and a data flow architecture according to an embodiment of this application. As shown in the figure, the system provided in this application includes two processes, a data acquisition process 101 and an artificial intelligence vision system performance enhancement process 102.
[0058] In the data acquisition process 101, there are three main processes. 103 is first performed, that is, an image that is a normal sample is acquired. In addition, manual annotation is not needed. Next, 105 is performed, that is, a system identifies a category of a result image and manual annotation is performed. Finally, 107 is performed, that is, classified defect data is labeled with a bounding box.
[0059] The artificial intelligence vision system performance enhancement process 102 is divided into three steps. When the system is first launched, as shown in 104, the system only has an autoencoder neural network, so the system can only identify whether the sample is abnormal. Then, as shown in 106, by combining two neural networks, that is, including the autoencoder neural network and a convolutional neural network, and making a joint decision, the system can identify defect category. Finally, as shown in 108, after combining three neural networks, that is, including the autoencoder neural network, the convolutional neural network, and a detection neural network, and making a decision, the system can find defects in some abnormal samples.
[0060] FIG. 2 shows a basic workflow of a system.
[0061] Step 201: In a first step of system deployment, a defect category is not completely determined yet, and a normal sample image is easy to acquire, so a large quantity of normal samples are acquired first.
[0062] Step 202: An acquired image data does not need to be manually annotated.
[0063] Step 203: Train and deploy a network by using the data in step 202. A neural network can recover an input image and determine whether an anomaly area exists in the image by comparing a difference between the input image and an image recovered by the network. If the input image is a normal sample as training data, the difference between the input image and the image recovered by the network is small. If the input image is an image with an anomaly area, the difference between the input image and the image recovered by the network becomes larger. Structure similarity index measure (Structure Similarity Index Measure, SSIM) of two images is calculated through segmentation to determine whether a segmented area is an anomaly area. The neural network may be a memory augmented deep autoencoder (memory augmented deep autoencoder) .
[0064] Step 204: At this stage, the system can only predict whether abnormal data exists, that is, a difference between Good or NOT Good.
[0065] Step 205: After the system runs for a period of time, acquire an image that is a normal sample and an image with an abnormal result.
[0066] Step 206: Classify and label the image based on the defect category.
[0067] Step 207: Train a convolutional neural network by using annotated image data, and then deploy the network into the system. During a prediction process, the image is input to the system and the convolutional neural network identifies the defect category. The neural network includes the memory-augmented deep autoencoder and a CNN-based classifier (CNN-based Classifier) .
[0068] Step 208: In this case, the system includes two joint decision-making neural networks, and the image is input to the two neural networks at a same time. When both an output result of the autoencoder neural network and a classification result of the convolutional neural network indicate that the image is normal, the image is determined to be a normal sample. If the image is abnormal, the image may be classified into Defect Category 1, Defect Category 2, ..., or Defect Category N.
[0069] Step 209: Acquire sample data at a location where a defect location may be determined. The defect category in this part of the data may be classified based on a criterion and located in the image.
[0070] Step 210: Use an annotation tool to annotate the defect sample, use a bounding box to frame a defect area, and specify a category.
[0071] Step 211: Train a CNN-based target detection neural network by using annotated data and deploy the CNN-based target detection neural network to an artificial intelligence decision-making system. The neural network includes the memory-augmented deep autoencoder, the CNN-based classifier (CNN-based Classifier) neural network, and a CNN-based Target Detector.
[0072] Step 212: At this time, the artificial intelligence system can classify the image, the abnormal data can be matched to the defect category (Defect Category) , and locate a defect image area. In some embodiments, a defect location coordinate (defect location coordinate) may be output.
[0073] FIG. 3 shows a structure of a neural network for anomaly detection.
[0074] Step 301: Use a convolutional neural layer to extract and encode a feature, in which a CNN-based feature encoder is used.
[0075] Step 302: Low-level information extracted by the encoder in the previous step is useless, and a more useful part is extracted by tightening processing (tightening processing) in this module.
[0076] Step 303: Use a convolution structure symmetrical with the encoder to decode the feature to obtain a recovered image.
[0077] Step 304: Use a sliding window (sliding window) to score a similarity between the input image and the recovered image within a window area. For example, structure similarity index measure (Structure Similarity Index Measure, SSIM) may be calculated by using a sliding window whose size is 7. If a quantity of windows with an abnormal value exceeds a threshold, the image may be determined to be an abnormal image.
[0078] Further, an application scenario of this application may be that the field of industrial defect detection is widely used, deployment can be speed up, deployment costs are reduced, and different category of defects can be effectively diagnosed. In addition, the application scenario of this application has the following advantages:
[0079] An advantage 1 is that an unsupervised learning neural network is used as a preliminary defect detection solution. Specifically, an advantage based on a technical feature is that training can be started quickly only with normal sample image data, and deployment can be accelerated to achieve preliminary identification of an abnormal sample. Further, a service impact brought by the advantage may be low deployment costs and flexible expansion.
[0080] An advantage 2 is that a supervised learning training image classification and a target detection neural network are used to improve detection accuracy and iteratively complete a construction of an AI decision-making system. Specifically, an advantage based on a technical feature is that recognition accuracy of the artificial intelligence decision-making system can be continuously improved with iteration of data annotation to enrich a task hierarchy. Further, a service impact brought by the advantage may be that an initial investment is reduced and a project investment risk is reduced. In addition, a system circulation is improved and has long-term value.
[0081] It may be understood that, although each step and example in FIG. 1 to FIG. 4 is displayed sequentially according to arrows, the steps are not necessarily performed according to a sequence indicated by the arrows. Unless otherwise explicitly specified in this application, execution of the steps is not strictly limited, and the steps may be performed in other sequences. In addition, at least some steps in FIG. 1 to FIG. 4 may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at a same time instant, and may be performed at different time instants. The sub-steps or stages are not necessarily performed in sequence, and the sub-steps or stages may be performed alternately with at least some of other steps, sub-steps or stages of other steps.
[0082] FIG. 5 provides an apparatus 500 for image anomaly detection. The apparatus 500 includes:
[0083] a first neural network module 501, configured to identify an image by using a first neural network, and train a second neural network based on an identified abnormal image;
[0084] a second neural network module 502, configured to use the first neural network and the second neural network jointly to identify a category of the abnormal image and label the abnormal image with a box, and training a third neural network by using the category of the image and the labeling box; and
[0085] a third neural network module 503, configured to detect an anomaly location of the image by using the first neural network, the second neural network, and the third neural network.
[0086] Further, the first neural network module 501 is further configured to identify an image by using an autoencoder neural network, and train a convolutional neural network based on an identified abnormal image.
[0087] Further, the first neural network module 501 is configured to identify an image by using the autoencoder neural network, and score based on a similarity.
[0088] The second neural network module 502 is configured to train a detection neural network by using the category of the image and the labeling box.
[0089] It should be noted that the apparatus may include more modules or fewer modules to implement the described functions.
[0090] For example, at least one module in FIG. 5 may be further divided into a plurality of different sub-modules, and each sub-module is configured to perform at least a part of the operations described herein in connection with the corresponding module. In addition, in some examples, the apparatus 500 may further include an additional module, configured to perform another operation already described in the specification. In addition, a person skilled in the art may understand that the example apparatus 500 may be implemented in software, hardware, firmware, or any combination thereof.
[0091] FIG. 6 provides a computer device. According to an embodiment, the computer device 600 may include a processor 602. The processor 602 executes a computer program stored in a memory 604. The foregoing method is implemented when the computer program is executed by the processor.
[0092] A person skilled in the art may understand that the structure shown in FIG. 6 is only a block diagram of a partial structure related to a solution in this application, and FIG. 6 not constitute a limitation to the computer device to which the solution in this application is applied. Specifically, the computer device may include more components or fewer components than those shown in the figure, or some components may be combined, or a different component deployment may be used.
[0093] A person of ordinary skill in the art may understand that all or some of procedures of the method in the foregoing implementations may be implemented by a computer program instructing relevant hardware. The computer program may be stored in a non-volatile computer-readable storage medium. When the computer program is executed, the procedures of the foregoing method implementations may be implemented. References to the memory, the storage, the database, or other medium used in the implementations provided in this application may all include at least one of a non-volatile memory or a volatile memory. The non-volatile memory may include a read-only memory (Read-Only Memory, ROM) , a magnetic tape, a floppy disk, a flash memory, an optical memory, or the like. The volatile memory may include a random access memory (Random Access Memory, RAM) or an external cache. As a description and not a limitation, the RAM may be in various forms, such as a static random access memory (Static Random Access Memory, SRAM) or a dynamic random access memory (Dynamic Random Access Memory, DRAM) .
[0094] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The foregoing step is implemented when the computer program is executed by a processor.
[0095] This application further provides a computer program product. The computer program product is tangibly stored on a computer-readable medium and includes computer-executable instructions. At least one processor is enabled to perform the foregoing method when the computer-executable instructions are executed.
[0096] Further, the computer program may be stored and run in a cloud to perform the method. Further, components of the program may be laid out on a plurality of devices and clouds. For example, a corresponding step may be laid out and run locally or on a local computer, or run on different cloud devices to transmit a signal through a communication connection. Alternatively, the components of the program may also be laid out and run locally or on a local computer. This application does not limit the manners or methods described. A corresponding technology may be flexibly laid out and deployed. A device and technology such as a cloud, big data, and a supercomputing capability may be fully utilized to perform and complete the methods.
[0097] Some implementations of the present disclosure may include a product. The product may include a storage medium for storing logic. Examples of the storage medium may include one or more types of computer-readable storage medium capable of storing electronic data, including a volatile memory or non-volatile memory, a removable memory or non-removable memory, an erasable memory or non-erasable memory, a writable memory or rewritable memory, and the like. Examples of the logic may include various software units, such as a software component, a program, an application, a computer program, an application program, a system program, a machine program, operating system software, middleware, firmware, a software module, a routine, a subroutine, a function, a method, a procedure, a software interface, an application programming interface (API) , an instruction set, computing code, computer code, a code segment, a computer code segment, a word, a value, a symbol, or any combination thereof. In some implementations, for example, the product may store executable computer program instructions. The processor is enabled to perform the methods and / or operations described in this specification when the computer program instructions are executed by the processor. The executable computer program instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, and dynamic code. The executable computer program instructions may be implemented based on a predefined computer language, manner, or syntax for commanding a computer to perform a specific function. The instructions may be implemented by using any suitable high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language.
[0098] The above includes examples of the disclosed architecture. Certainly, it is impossible to describe every conceivable combination of components and / or methods, but a person skilled in the art may understand that many other combinations and permutations are possible. Therefore, the novel architecture is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and scope of the appended claims.
Claims
1.A method for image anomaly detection, comprising:identifying an image by using a first neural network; training a second neural network based on an identified abnormal image;using the first neural network and the second neural network jointly to identify a category of the abnormal image and label the abnormal image with a box; training a third neural network by using the category of the image and the labeling box; anddetecting an anomaly location of the image by using the first neural network, the second neural network, and the third neural network.2.The method according to claim 1, wherein the identifying an image by using a first neural network; training a second neural network based on an identified abnormal image comprises:identifying an image by using an autoencoder neural network; and training a convolutional neural network based on an identified abnormal image.3.The method according to claim 2, wherein the identifying an image by using an autoencoder neural network comprises:identifying an image by using the autoencoder neural network and scoring based on a similarity.4.The method according to claim 1, the training a third neural network by using the category of the image and the labeling box comprises:training a detection neural network by using the category of the image and the labeling box.5.An apparatus (500) for image anomaly detection, comprising:a first neural network module (501) , configured to identify an image by using a first neural network, and training a second neural network based on an identified abnormal image;a second neural network module (502) , configured to use the first neural network and the second neural network jointly to identify a category of the abnormal image and label the abnormal image with a box, and configured to train a third neural network by using the category of the image and the labeling box; anda third neural network module (503) , configured to detect an anomaly location of the image by using the first neural network, the second neural network, and the third neural network.6.The apparatus (500) according to claim 5, whereinthe first neural network module (501) is further configured to identify an image by using an autoencoder neural network; and train a convolutional neural network based on an identified abnormal image.7.The apparatus (500) according to claim 6, whereinthe first neural network module (501) is configured to identify an image by using the autoencoder neural network, and score based on a similarity.8.The apparatus (500) according to claim 5, whereinthe second neural network module (502) is configured to train a detection neural network by using the category of the image and the labeling box.9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the steps of the method according to any one of claims 1 to 4 are implemented when the processor executes the computer program.10.A computer-readable storage medium, storing a computer program, wherein the steps of the method according to any one of claims 1 to 4 are implemented when the computer program is executed by a processor.11.A computer program product, wherein the computer program product is tangibly stored on a computer-readable medium and comprises computer-executable instructions, and at least one processor is enabled to perform the method according to any one of claims 1 to 4 when the computer-executable instructions are executed.
Citation Information
Patent Citations
Industrial quality inspection method and device and computer readable storage medium
CN112528975A
Abnormality detection network training and anomaly detection method, device and equipment
CN115810011A
Abnormal image detection method and device, model training method and device, equipment and medium
CN116958020A
Method and apparatus for analyzing a product, training method, system, computer program, and computer-readable storage medium
US20230022631A1
Cited By
Puncher plug anomaly detection method and device, electronic equipment and medium
CN121121409A
Multi-class anomaly detection method based on memory guidance and class decoupling
CN121600331A
Passenger car part anomaly detection method based on deep residual network
CN122223009A