A method for constructing an image detection model, an image detection method and an image detection device
Through the image detection model of cascade feedback network structure, the problem of inefficient detection of abnormal areas in industrial quality inspection is solved, and unsupervised learning and efficient abnormal area detection is achieved.
Patent Information
- Application Number
- CN202011495984.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-12-17
AI Technical Summary
The existing deep learning methods are insufficient in industrial quality inspection, and it is difficult to effectively detect abnormal areas on the surface of an object, especially when the range of changes in the shape, posture and color information of the object are large.
The image detection model with a cascade feedback network structure is adopted, and the network nodes formed by cascade feedback are configured through multiple shallow autoencoders, and the loss function is configured and multiple normal sample images are trained to obtain the image detection model.
The unsupervised learning method is realized, so that only multiple normal images are required to participate in training, and no abnormal sample images and pre-labeling are required, which improves the construction efficiency of image detection models, and simplifies the detection of abnormal areas through high-quality reconstruction images.
Smart Images

Figure CN112435258B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for constructing an image detection model, an image detection method and a device. Background Art
[0002] In recent years, deep learning has become the focus of attention in various fields at home and abroad. Deep learning includes supervised learning and unsupervised learning. In the field of computer vision, supervised learning refers to training neural networks through one-to-one correspondence between images and annotation information, so that they can complete tasks such as classification, target detection, and semantic segmentation; unsupervised learning refers to training neural networks only with unannotated image information, so that they can complete tasks such as clustering, anomaly detection, and image generation. In the field of industrial quality inspection, the most widely used methods are manual feature selection methods and supervised deep learning methods (hereinafter referred to as supervised learning methods).
[0003] There are still some limitations in the manual feature selection method: the shape, posture, color and other information of the object to be detected need to change within a certain range. When the shape, posture and color information of the object change too much, it is difficult to judge the pixel accuracy of abnormal areas (such as holes, cracks, cuts, printing, etc. on the surface of the object) and normal areas through manual standards. Whether between standard images and defective images, between standard images and standard images, or between defective images of the same type, it is often difficult to detect by manually selecting features when the shape and posture of the object surface change within a large range.
[0004] In recent years, supervised learning methods have become popular to solve the problem of manual feature selection methods being ineffective. By designing a convolutional neural network, images of objects to be detected (including a large number of normal images and abnormal images) are collected and labeled to form a data set, and then the convolutional neural network is trained using the data set to achieve automatic feature selection and judgment. Although supervised learning methods can still generate results with high accuracy and robustness when the shape, posture, color and other information of the object vary widely, there are also some obvious shortcomings. On the one hand, it is difficult to obtain a sufficient number and variety of abnormal samples. On the other hand, the labeling of a large number of images directly leads to time-consuming and high-cost problems. Summary of the invention
[0005] The main technical problem solved by the present invention is: how to overcome the shortcomings of existing deep learning methods in industrial quality inspection. In order to solve the above technical problems, the present application provides a method for constructing an image detection model, an image detection method and a device.
[0006] According to the first aspect, in an embodiment, a method for constructing an image detection model includes: establishing the structure of a cascade feedback network; the cascade feedback network includes multiple network nodes formed by multiple shallow autoencoders through cascade feedback; obtaining a corresponding loss function according to the structural configuration of the cascade feedback network; training the cascade feedback network using multiple normal sample images of objects to be detected, updating the network parameters of the cascade feedback network using the loss function, and obtaining an image detection model after the training is completed.
[0007] The structure of establishing a cascade feedback network includes: using convolutional neural units to form each shallow autoencoder; the shallow autoencoder includes an encoder composed of a convolutional layer and a downsampling layer, and a decoder composed of an upsampling layer and a convolutional layer; the encoder is used to receive the image input by the shallow autoencoder and convert it into semantic information, and the decoder is used to restore the semantic information and output a reconstructed image; a plurality of shallow autoencoders are sequentially sorted, the output of each shallow autoencoder is fed back to the input of the next shallow autoencoder, and each shallow autoencoder is used as a network node in the cascade feedback network; each network node is set as a node group, and the cascade feedback network is established by cascading each network node in the node group.
[0008] The corresponding loss function is obtained according to the structural configuration of the cascade feedback network, including: calculating the Euclidean distances of the images corresponding to the first and last two network nodes in the node grouping, thereby obtaining the first image reconstruction quality represented by the Euclidean distance of the image, and expressed by the formula:
[0009]
[0010] Wherein, x0 is the image input to the first network node in the cascade feedback network, is the reconstructed image output by the first network node, is the reconstructed image output by the Nth network node; the variance statistics of several network nodes in the node group are calculated to obtain the reconstruction quality of the second image represented by the variance statistics, and it is expressed by the formula:
[0011]
[0012] Where m is the number of network nodes, c is the number of channels of the reconstructed image output by the kth network node, is the value of the i-th channel on the reconstructed image output by the k-th network node, is the average value of the ith channel on the m reconstructed images, and satisfy
[0013]
[0014] The third image reconstruction quality is calculated according to the first image reconstruction quality and the second image reconstruction quality, and is expressed by the formula:
[0015] Loss d =Loss b +Loss var ;
[0016] The first image reconstruction quality expression formula Loss b , or the third image reconstruction quality expression formula Loss d The loss function configured as the corresponding one of the cascade feedback network.
[0017] The method uses multiple normal sample images of the object to be detected to train the cascade feedback network, updates the network parameters of the cascade feedback network through the loss function, and obtains the image detection model after the training is completed, including: obtaining multiple normal sample images of the object to be detected; the normal sample images do not contain the abnormal surface area of the object to be detected; using the normal sample images as the image input to the first network node in the cascade feedback network, and inputting each of the normal sample images into the cascade feedback network in turn for training; terminating the training when the difference between the before and after calculations of the loss function corresponding to the cascade feedback network is less than a preset threshold, or the corresponding loss function reaches a preset number of iterations; and obtaining the image detection model of the object to be detected using the cascade feedback network with updated network parameters at the end of the training.
[0018] According to the second aspect, an embodiment provides an image detection method based on cascade feedback, which includes: obtaining an image to be detected of an object to be detected; inputting the image to be detected into the image detection model constructed by the construction method described in the first aspect above, and detecting to obtain a reconstructed image output by any network node in the cascade feedback network; comparing the reconstructed image with the image to be detected to obtain an abnormal surface area of the object to be detected.
[0019] The method of comparing the reconstructed image with the image to be detected to obtain the abnormal surface area of the object to be detected includes: constructing an evaluation function of the abnormal surface area using the reconstructed image and the image to be detected; when the value of the evaluation function is greater than or equal to a predetermined value, determining that the reconstructed image contains the abnormal surface area of the object to be detected, and obtaining the abnormal surface area of the object to be detected by comparing the difference between the reconstructed image and the image to be detected; and outputting the image to be detected of the object to be detected and the abnormal surface area of the object to be detected.
[0020] The method of constructing an evaluation function of the surface abnormal area by using the reconstructed image and the image to be detected includes: obtaining the reconstructed image output by each network node in the cascade feedback network after the image to be detected is input into the image detection model, thereby constructing an evaluation function of the surface abnormal area, wherein the evaluation function is expressed by any of the following formulas:
[0021]
[0022] Among them, x0′ is the image to be detected, A reconstructed image output by any network node when the image to be detected is input into the cascade feedback network, is the output image of the intermediate network layer of any network node when the image to be detected is input into the cascade feedback network, M(x′0) is the output image of the intermediate network layer of the first network node when the image to be detected is input into the cascade feedback network, s is a set consisting of a number of network nodes, subscripts n, j, k are all serial numbers of network nodes; m is the number of a number of network nodes, c is the number of channels of the reconstructed image output by the kth network node, The value of the i-th channel on the reconstructed image output by the k-th network node when the image to be detected is input into the cascade feedback network, x i ' is the average value of the ith channel on the m reconstructed images when the image to be detected is input into the cascade feedback network.
[0023] According to the third aspect, an embodiment provides an image detection device, which includes: an image acquisition component, used to acquire an image to be detected of an object to be detected; a processor, connected to the image acquisition component, and used to construct an image detection model through the construction method described in the first aspect above, and / or, to obtain a surface abnormality area of the object to be detected in the image to be detected through the image detection method described in the second aspect above; a display, connected to the processor, and used to display the image to be detected and the surface abnormality area of the object to be detected.
[0024] The processor includes a model building module and an anomaly detection module; the model building module is used to train a pre-established cascade feedback model using one or more normal sample images, and obtain an image detection model by updating network parameters through a loss function; the cascade feedback network includes a plurality of network nodes formed by a plurality of shallow autoencoders through cascade feedback, and the loss function is configured according to the structure of the cascade feedback network; the anomaly testing module is connected to the model building module, and is used to input the image to be detected into the image detection model, and output the surface abnormal area of the object to be detected through detection processing.
[0025] According to the fourth aspect, an embodiment provides a computer-readable storage medium, which includes a program, and the program can be executed by a processor to implement the construction method described in the first aspect above, and / or, to implement the image detection method described in the second aspect above.
[0026] The beneficial effects of this application are:
[0027] A method for constructing an image detection model, an image detection method and an apparatus according to the above-mentioned embodiment. The construction method includes: establishing the structure of a cascade feedback network; obtaining a corresponding loss function according to the structural configuration of the cascade feedback network; training the cascade feedback network using multiple normal sample images of the object to be detected, updating the network parameters of the cascade feedback network through the loss function, and obtaining an image detection model after the training is completed. The image detection method includes: obtaining an image to be detected of the object to be detected; inputting the image to be detected into the constructed image detection model, detecting and obtaining a reconstructed image output by any network node in the cascade feedback network; comparing the reconstructed image with the image to be detected, and obtaining the surface abnormal area of the object to be detected. First, a cascade feedback network is used to train an image detection model. This unsupervised learning method only requires multiple normal images to participate in the training, and does not require abnormal sample images and prior annotations. It is not only easy to obtain a training set, but also does not require time and effort for annotation, which is conducive to improving the efficiency of building the image detection model. Second, a shallow autoencoder is used in the cascade feedback network to form network nodes. This structure, which is similar to a recurrent neural network, can keep the parameters of each network node completely consistent, and the number of parameters can be greatly reduced compared to existing methods, which has the advantages of being easy to transmit, store and deploy, thereby simplifying the network training process and accelerating the convergence of the loss function. Third, the image to be detected is input into the image detection model to facilitate detection and obtain the reconstructed image output by any network node in the cascade feedback network. The generated reconstructed image has the characteristics of high pixel accuracy and good feature reconstruction effect, which makes the reconstruction error of the abnormal area relatively large. It is only necessary to separate the abnormal area from the normal area through a simple standard to complete the detection of the surface abnormal area. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flowchart of a method for constructing an image detection model in Example 1 of the present application;
[0029] Figure 2 Flowchart for establishing a cascade feedback network;
[0030] Figure 3 Flowchart for configuring loss function;
[0031] Figure 4 Flowchart for training an image detection model;
[0032] Figure 5 is a schematic diagram of the structure of a shallow autoencoder, where Figure 5a The connection diagram of the encoder and decoder is shown in Figure 2. Figure 5b Schematic diagram of the connection between convolutional layer, downsampling layer and upsampling layer;
[0033] Figure 6 Schematic diagram of the principle of cascade feedback of multiple shallow autoencoders;
[0034] Figure 7 This is a flow chart of an image detection method based on cascade feedback in Embodiment 2 of the present application;
[0035] Figure 8 A flow chart for obtaining the abnormal surface area of the object to be detected;
[0036] FIG9 is the detection result of the object to be detected, where Figure 9a is the detection result of the object to be detected without abnormal area, Figure 9b is the detection result of the object to be detected with a hole area. Fig.9c is the detection result of the object to be detected with a bursting area, Figure 9d is the detection result of the object to be detected with the cut mark area. Fig.9e is the detection result of the object to be detected with a printed area;
[0037] Fig.10 This is a schematic diagram of the structure of the image detection device in the third embodiment of the present application;
[0038] Fig.11 is a schematic diagram of the structure of the processor;
[0039] Fig.12 This is a schematic diagram of the structure of the image detection device in Example 4 of the present application. DETAILED DESCRIPTION
[0040] The present invention is further described in detail below by specific embodiments in conjunction with the accompanying drawings. Wherein similar elements in different embodiments adopt associated similar element numbers. In the following embodiments, many detailed descriptions are for making the present application better understood. However, those skilled in the art can easily recognize that some features can be omitted in different situations, or can be replaced by other elements, materials, methods. In some cases, some operations related to the present application are not shown or described in the specification, this is to avoid the core part of the present application being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations according to the description in the specification and the general technical knowledge in the art.
[0041] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various implementations. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a required sequence, unless otherwise specified that a certain sequence must be followed.
[0042] The serial numbers of the components in this document, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning. The "connection" and "coupling" mentioned in this application, unless otherwise specified, include direct and indirect connections (couplings).
[0043] The main unsupervised image detection methods include Generative Adversarial Network (GAN), Auto Encoder (AE), and Variational Auto Encoder (VAE). Among them, the Generative Adversarial Network (GAN) consists of a generative network and a discriminative network. The Generative Adversarial Network randomly samples from the latent space as input, and its output needs to imitate the real samples of the training set as much as possible. Among them, the Auto Encoder (AE) consists of an encoder and a decoder. The image information is generated by the encoder to generate high-dimensional, low-resolution semantic information, and the semantic information is directly used as the latent variable. The decoder restores the latent variable to an image with the same format as the original image through upsampling and convolutional neural network. The output image needs to imitate the input image as much as possible to achieve the effect of image reconstruction. Among them, the variational autoencoder (VAE) is also composed of an encoder and a decoder. The encoder generates high-dimensional, low-resolution semantic information from image information. By calculating the mean, variance and other information of the semantic information generated by the encoder, latent variables are sampled in random distributions such as Gaussian noise. The decoder restores the latent variables to image information through upsampling and convolutional neural networks. Its output needs to imitate the input image as much as possible to achieve the effect of image reconstruction.
[0044] The above-mentioned generative adversarial network, autoencoder, and variational autoencoder are applied in the anomaly detection scenario. Their working principle is: first use normal images to train the neural network, and its output needs to imitate the input image as much as possible, that is, the neural network has a small reconstruction error in the normal area (the input original image I outputs the reconstructed image O through the corresponding neural network, and the reconstruction error refers to the difference between the reconstructed image O and the original image I). At the same time, since the abnormal image data is not used for training, the neural network often has a large reconstruction error in the abnormal area. By generating reconstruction errors, we judge that the area with a small reconstruction error is a normal area, and the area with a large reconstruction error is an abnormal area, so as to detect the abnormal area with pixel accuracy. However, there are still some application deficiencies in using any of these methods alone. The shortcomings of the generative adversarial network and the variational autoencoder are: since the information of the original image is not fully utilized, it is difficult to generate a reconstructed image with high pixel accuracy (that is, the generated image can only be roughly close to the input image, and the position accuracy and numerical accuracy of the pixels of the generated reconstructed image are poor), and it is difficult to formulate a judgment standard to distinguish between defective areas and normal areas. The shortcomings of the autoencoder are: due to the use of multi-layer downsampling, the pixel position accuracy is poor and the reconstruction effect of smaller features is poor. It is difficult to judge normal areas and abnormal areas by comparing the differences between the original image and the reconstructed image (i.e. generating reconstruction errors).
[0045] For image reconstruction, several methods are needed to solve the accuracy of feature reconstruction. One is to use a deeper network structure and ensure the position accuracy and feature accuracy of the output detection by fusing the bottom-level features with the high-level features; the other is to use a wider network structure and ensure the adaptability to objects of different sizes by fusing the features of different receptive field sizes. A large receptive field has a good effect on the detection of large objects, and a small receptive field meets the detection requirements of small objects. In the technical solution of this application, for the input image x, it is necessary to construct a function To better reflect the transformation and mapping relationship between the input image and the output image, where ω is the reconstruction and solution parameter.
[0046] When using an autoencoder to reconstruct an image, if a deep convolutional neural network structure (i.e., multiple convolutional layers and downsampling layers) is used, the reconstruction errors of normal and abnormal regions can be compared to distinguish the two. However, due to the presence of multiple convolutional layers and downsampling layers, the generated reconstructed image has disadvantages such as poor pixel position accuracy and poor reconstruction effect on smaller features. If a shallow convolutional neural network structure is used, the pixel position accuracy is higher and the reconstruction effect on smaller features is better. However, since the number of convolutional layers and downsampling layers of the shallow autoencoder is small, it means that the encoder uses semantic encoding generated by relatively low-level features. In the reconstruction process, the abnormal region is very likely to contain similar low-level features. Therefore, the reconstruction error of the abnormal image may be close to that of the normal image in terms of value, and it is difficult to distinguish the two by comparing the reconstruction errors of the normal and abnormal regions. Then, the generated reconstructed image should meet the following requirements: high pixel accuracy, good reconstruction effect of smaller features, and large reconstruction error of the abnormal region, so that the abnormal region and the normal region can be separated by a simple standard in the end.
[0047] The technical solution of the present application is to construct an image detection model based on the idea of cascading and feedback. The purpose of the cascade process is to build a deeper network structure while maintaining the position information and feature information unique to the shallow network; the purpose of the feedback is to maintain the normal structural features during the reconstruction process and gradually increase the distance between abnormal features and normal features. The present application adopts a combination of a recurrent convolutional neural network (RNN) and an autoencoder (Auto Encoder), generates a high-quality reconstruction of the original image through a shallow autoencoder, and inputs the output reconstructed image into the shallow autoencoder again, and so on, continuously using the reconstruction result of the previous iteration as the input of this iteration, and gradually amplifying the reconstruction error of the abnormal area after multiple iterations, while the reconstructed image of the normal area remains basically unchanged. Because the shallow autoencoder contains fewer continuous downsampling layers, the reconstructed image always maintains a high pixel position accuracy and a good reconstruction effect of smaller features during the reconstruction process.
[0048] The technical solution of the present application is described in detail below with reference to the embodiments.
[0049] Embodiment 1
[0050] Please refer to Figure 1 , this embodiment discloses a method for constructing an image detection model, which includes steps S110-S130, which are described below respectively.
[0051] Step S110, establishing a cascade feedback network structure, wherein the cascade feedback network includes a plurality of network nodes formed by a plurality of shallow autoencoders through cascade feedback.
[0052] The cascade feedback network is a network structure formed by cascade feedback of multiple shallow autoencoders. Multiple shallow autoencoders are arranged in sequence to form a hierarchical structure, and then the output of each shallow autoencoder is fed back to the input of the next shallow autoencoder. At this time, each shallow autoencoder is used as a network node in the cascade feedback network.
[0053] Shallow autoencoders are an autoencoder with fewer convolutional and downsampling layers. They are also artificial neural networks that achieve efficient representation of output data through unsupervised learning. This efficient representation of output data is called encoding, and the capacity of the encoded information is generally much smaller than the input data, making the autoencoder useful for dimensionality reduction. More importantly, autoencoders can be used as powerful feature detectors for pre-training of deep neural networks.
[0054] Step S120, obtaining a corresponding loss function according to the structural configuration of the cascade feedback network.
[0055] Since there are multiple network nodes in the cascade feedback network, and each network node is connected in the form of cascade feedback, the input and output of each network node can be represented, so that the image reconstruction quality can be obtained through image Euclidean distance calculation, pixel statistical analysis, etc., thereby configuring the loss function corresponding to the cascade feedback network.
[0056] Step S130, using multiple normal sample images of the object to be detected to train the cascade feedback network, updating the network parameters of the cascade feedback network through the loss function, and obtaining the image detection model after the training is completed.
[0057] In order to learn the surface features of the object to be detected, multiple normal sample images can be input into the cascade feedback network in sequence for training; the training is terminated when the difference between the before and after calculations of the loss function corresponding to the cascade feedback network is less than a preset threshold, or the corresponding loss function reaches a preset number of iterations, so that the trained cascade feedback network is used as the image detection model of the object to be detected.
[0058] It should be noted that the objects to be detected here can be products on an industrial assembly line, mechanical parts in an object box, tools on an operating table, etc. When capturing an image of the surface of the object to be detected, the surface features of the object will be displayed or presented on the corresponding pattern. If there are holes, cracks, cuts, printing, dust, flaws, dirt and other defects on the surface of the object to be detected, the captured image will be an abnormal sample image; if there are no such defects on the surface of the object to be detected, the captured image will be a normal sample image.
[0059] In this embodiment, see Figure 2 The above-mentioned step S110 mainly involves the process of establishing a cascade feedback network, which may specifically include steps S111-S113, which are described as follows.
[0060] Step S111, using convolutional neural units to form each shallow autoencoder.
[0061] See also Figure 5a and Figure 5b , the shallow autoencoder includes an encoder composed of a convolutional layer and a downsampling layer, and a decoder composed of an upsampling layer and a convolutional layer; the convolutional layer and downsampling layer contained in the encoder have a one-to-one correspondence with the convolutional layer and upsampling layer contained in the decoder. In addition, the encoder is used to receive the image input by the shallow autoencoder and convert it into semantic information, such as performing latent semantic encoding on the features A, B, C, and D in the input image; correspondingly, the decoder is used to restore the semantic information and output the reconstructed image, such as outputting the image features A′, B′, C′, and D′, where the features A′, B′, C′, and D′ usually only differ from the features A, B, C, and D in terms of information capacity.
[0062] For each shallow autoencoder, since fewer convolution operations and downsampling operations are used, it is easy to ensure that the pixel position accuracy of the reconstructed image is high, and the reconstruction effect of smaller features is better. The working principle of the shallow autoencoder can be expressed by the following formula:
[0063]
[0064] Where x is the input image, is the output feature, z is the latent semantic code, E is the encoder neural network, and D is the decoder neural network. The goal of the shallow autoencoder is to make the output Try to keep it consistent with input x.
[0065] It should be noted that shallow encoders work by simply learning to copy inputs to outputs, a task that can be made extremely difficult by adding constraints to the neural network in different ways. For example, the size of the internal representation can be limited, or noise can be added to the training data and the autoencoder can be trained to recover the original features. These constraints prevent the autoencoder from mechanically copying the input to the output and force it to learn an efficient representation of the data.
[0066] Since the shallow autoencoder has fewer convolutional layers and downsampling layers, it means that the encoder uses relatively low-level features to generate semantic codes. During the reconstruction process, the abnormal area is very likely to contain similar low-level features. Therefore, the reconstruction error of the abnormal image may be numerically close to that of the normal image.
[0067] Step S112, sorting the multiple shallow autoencoders in sequence, feeding back the output of each shallow autoencoder to the input of the next shallow autoencoder, and using each shallow autoencoder as a network node in the cascade feedback network.
[0068] See also Figure 6 , N shallow autoencoders (shallow autoencoder 1, shallow autoencoder 2, ..., shallow autoencoder N) are arranged in sequence, the normal sample image is input to shallow autoencoder 1, the image is reconstructed after passing through shallow autoencoder 1, the reconstructed image R1 is used as input to shallow autoencoder 2, the image reconstruction result is obtained, and the reconstructed image R2 is input to the next shallow autoencoder, and so on, the output of the previous iteration is continuously used as the input of the next iteration, after multiple cycles of iteration, the reconstruction error is gradually amplified. In the final image reconstruction result, that is, the reconstructed image R N In the reconstruction, the reconstruction error of the abnormal area will be significantly greater than that of the normal area, so the abnormal area can be distinguished from the normal area.
[0069] Some formulas can be used to represent the image iteration process of shallow autoencoder:
[0070]
[0071] Among them, φ is a shallow autoencoder, ω is the weight parameter of the neural network, x0 is the input image of the cascade feedback network (the input image in the training stage is a normal sample image), is the reconstructed image generated by x0 after passing through the first network node, yes The reconstructed image generated by the nth (n>0) node (this article uses ~ to distinguish the reconstruction result from the original image, and the one with ~ is the reconstruction result).
[0072] See also Figure 6 , N shallow autoencoders are structurally similar to a cascade feedback network, and the constituent units are each shallow autoencoder, which contains several shallow autoencoders with the same (or different) parameters, and each shallow autoencoder can be called a network node. It should be noted that these network nodes are trained using normal sample images, and the training goal is to make the output of the network node as consistent as possible with the input normal image, so that a trained neural network can be obtained.
[0073] Step S113, each network node is set as a node group, and a cascade feedback network is established by cascading each network node in the node group. Figure 6 The N network nodes in are set as a node group, so that the cascade feedback forms a cascade feedback network.
[0074] In this embodiment, see Figure 3 The above-mentioned step S120 mainly involves the process of configuring the loss function, which may specifically include steps S121-S124, which are described as follows.
[0075] Step S121, calculate the image Euclidean distances corresponding to the first and last two network nodes in the node group, so as to obtain the first image reconstruction quality represented by the image Euclidean distance, which is expressed by the formula:
[0076]
[0077] Among them, x0 is the image input to the first network node in the cascade feedback network, is the reconstructed image output by the first network node, It is the reconstructed image output by the Nth network node.
[0078] In a cascade feedback network, assume that there are N network nodes, and the input of the cascade feedback network is x0. Then the input of the first network node is x0, and the output is The second network node input is The output is By analogy, we can know that the input of the Nth network node is The output is First image reconstruction quality Loss b The configuration is mainly based on the difference between the reconstruction result of the cascade feedback network and the original image.
[0079] For the first image reconstruction quality, Loss b Represents the quality of the reconstruction result of the training image, which mainly consists of two parts: the Euclidean distance between the reconstruction result output by the first node and the training image, and the Euclidean distance between the reconstruction result output by the last node and the training image. It can be understood that using normal sample images to train the cascade feedback network can ensure that each network node of the cascade feedback network has a good reconstruction result for the normal area (smaller Euclidean distance from the original image), while for the abnormal area that has not been trained, the reconstruction result is poor (larger Euclidean distance from the original image).
[0080] Step S122, calculate the variance statistics of several network nodes in the node group, so as to obtain the second image reconstruction quality represented by the variance statistics, which can be expressed as follows:
[0081]
[0082] Where m is the number of network nodes, c is the number of channels of the reconstructed image output by the kth network node, is the value of the i-th channel on the reconstructed image output by the k-th network node, is the average value of the ith channel on the m reconstructed images, and satisfy
[0083]
[0084] In the cascade feedback network, the reconstructed image output by each network node can be used Indicates that nodes are distinguished by subscripts. For example, the reconstruction result output by the i-th node is Each network node outputs a w*h*c reconstructed image (w is the image width, h is the image height, and c is the number of image channels). l (2<l≤N) network nodes are selected from a total of N network nodes to obtain the output of each network node. Combined with the original image itself, a total of l+1 images are obtained, including 1 original normal sample image and m-1 reconstructed images of different network nodes. By analyzing each pixel point, w*h*c histograms can be obtained. Each histogram consists of l+1 groups of data. Statistical analysis of the histogram (such as calculating the variance, calculating the difference between any number of groups, etc.) can be performed to analyze the abnormal area. Second image reconstruction quality Loss var The configuration is mainly based on the statistical information of multiple network nodes of the cascade feedback network.
[0085] Step S123, a third image reconstruction quality is calculated according to the first image reconstruction quality and the second image reconstruction quality, and is expressed as:
[0086] Loss d =Loss b +Loss var .
[0087] It can be understood that the third image reconstruction quality Loss d is the reconstruction quality Loss of the first image b and the second image reconstruction quality Loss var Combine to form.
[0088] Step S124: Substitute the first image reconstruction quality expression formula Loss b , or the third image reconstruction quality expression formula Loss d Configured as the loss function corresponding to the cascade feedback network.
[0089] It can be understood that the key to training the cascade feedback network is to construct the corresponding loss function. The quality of the loss function reflects the ability to build the image detection model to a certain extent. d When configured as the loss function corresponding to the cascade feedback network, the difference between the reconstruction result of the cascade feedback network and the original image is taken into account, and the statistical information of multiple network nodes of the cascade feedback network is also taken into account.
[0090] In this embodiment, see Figure 4 The above-mentioned step S140 mainly involves the process of training to obtain the image detection model, which may specifically include steps S131-S133, which are described as follows.
[0091] Step S131, obtaining a plurality of normal sample images of the object to be detected. The normal sample images here do not contain abnormal surface areas of the object to be detected.
[0092] Step S132, using the normal sample image as the image input to the first network node in the cascade feedback network, and inputting each normal sample image into the cascade feedback network in sequence for training.
[0093] Step S133, the training is terminated when the difference between the before and after calculations of the loss function corresponding to the cascade feedback network is less than a preset threshold, or the corresponding loss function reaches a preset number of iterations; when the training is terminated, the image detection model of the object to be detected is obtained using the cascade feedback network with updated network parameters.
[0094] It can be understood that in this embodiment, a cascade feedback network is used to train the image detection model. This unsupervised learning method only requires multiple normal images to participate in the training. No abnormal sample images and prior annotations are required. It is not only easy to obtain a training set, but also does not require time and effort for annotation, which is conducive to improving the efficiency of building the image detection model.
[0095] It can be understood that the present embodiment uses a shallow autoencoder in a cascade feedback network to form network nodes. This structure, which is similar to a recurrent neural network, can keep the parameters of each network node completely consistent, and the number of parameters can be greatly reduced compared to existing methods, which has the advantage of being easy to transmit, store, and deploy, thereby simplifying the network training process and accelerating the convergence of the loss function.
[0096] Embodiment 2
[0097] Please refer to Figure 7 , the present application discloses a method for constructing an image detection model, which includes steps S210-S230, which are described below respectively.
[0098] Step S210, obtaining an image of the object to be detected. The image to be detected may contain a normal surface area and an abnormal surface area of the object to be detected, so it is necessary to obtain the abnormal surface area by image detection.
[0099] Step S220, inputting the image to be detected into the constructed image detection model, and detecting and obtaining the reconstructed image output by any network node in the cascade feedback network.
[0100] It should be noted that the image detection model here is an image detection model constructed by the construction method in Example 1, which is a trained cascade feedback network and includes multiple shallow autoencoders formed by cascade feedback, and each shallow autoencoder serves as a network node in the cascade feedback network.
[0101] For any network node, the reconstructed image output by it can be expressed as
[0102]
[0103] Among them, φ is a shallow autoencoder, ω is the weight parameter of the neural network, The input image of the nth network node when the image to be detected is input into the image detection model. The reconstructed image output by the nth network node when the image to be detected is input into the image detection model.
[0104] Step S230: Compare the reconstructed image with the image to be detected to obtain the abnormal surface area of the object to be detected.
[0105] In this embodiment, see Figure 8 The above-mentioned step S230 mainly involves the process of comparing and obtaining the abnormal surface area of the object to be detected, which may specifically include steps S231-S233, which are described as follows.
[0106] Step S231, constructing an evaluation function of the surface abnormal area using the reconstructed image and the image to be detected.
[0107] In a specific embodiment, after the image to be detected is input into the image detection model, the reconstructed image output by each network node in the cascade feedback network is obtained, so as to construct an evaluation function of the surface abnormal area. The evaluation function is expressed by any of the following formulas:
[0108]
[0109] Among them, x0′ is the image to be detected, The reconstructed image output by any network node when the image to be detected is input into the cascade feedback network. is the output image of the intermediate network layer of any network node when the image to be detected is input into the cascade feedback network, M(x′0) is the output image of the intermediate network layer of the first network node when the image to be detected is input into the cascade feedback network, s is a set consisting of a number of network nodes, subscripts n, j, k are all serial numbers of network nodes; m is the number of a number of network nodes, c is the number of channels of the reconstructed image output by the kth network node, The value of the i-th channel on the reconstructed image output by the k-th network node when the image to be detected is input into the cascade feedback network. The average value of the i-th channel on the m reconstructed images when the image to be detected is input into the cascade feedback network.
[0110] It should be noted that since M() represents the output result of the intermediate network layer, M can represent any of the convolutional layer, downsampling layer, and upsampling layer.
[0111] Step S232, when the value of the evaluation function is greater than or equal to a predetermined value, it is determined that the reconstructed image contains a surface abnormality region of the object to be detected, and the surface abnormality region of the object to be detected is obtained by comparing the difference between the reconstructed image and the image to be detected.
[0112] for If the function value of any evaluation function represented by is greater than or equal to a predetermined value (which can be preset by the user or generated by the system by default), it indicates that If there is a difference between the reconstructed image represented by and the image to be detected x′0, or between other reconstructed images, then the reconstructed image The area of the surface that contains the anomaly of the object to be detected.
[0113] It should be noted that when performing a difference comparison between the reconstructed image and the image to be detected, it is only necessary to calculate the grayscale difference between the reconstructed image and the image to be detected. The pixel points whose calculation results are greater than the preset threshold are the pixel points in the surface abnormality area; after counting the pixel points, the surface abnormality area of the object to be detected can be obtained.
[0114] Step S233: outputting the image to be detected of the object to be detected and the abnormal surface area of the object to be detected.
[0115] For the hazelnut object to be detected, its image to be detected and the abnormal surface area can be referred to Figure 9. Figure 9a The left image in the figure is the image to be detected without abnormalities on the hazelnut surface, and the right image is the difference comparison result between the reconstructed image output by any network node and the image to be detected. The results show that there is no abnormal surface area on the hazelnut surface. Figure 9b The left image in the figure is the image to be detected with holes on the surface of the hazelnut, and the right image is the difference comparison result between the reconstructed image output by any network node and the image to be detected. The result shows that there is a surface abnormal area with the same shape as the hole on the surface of the hazelnut. Fig.9c The left image in the figure is a detection image with cracks on the surface of hazelnuts, and the right image is the difference comparison result between the reconstructed image output by any network node and the detection image. The results show that there is a surface abnormal area with the same shape as the cracks on the surface of hazelnuts. Figure 9dThe left image in the figure is the image to be tested with a cut on the surface of the hazelnut, and the right image is the difference comparison result between the reconstructed image output by any network node and the image to be tested. The result shows that there is an abnormal surface area with the same shape as the cut on the surface of the hazelnut. Fig.9e The left image in the figure is the image to be detected with printing on the surface of hazelnut, and the right image is the difference comparison result between the reconstructed image output by any network node and the image to be detected. The results show that there is a surface abnormal area with the same shape as the printing on the surface of hazelnut.
[0116] In FIG9 , the abnormal area on the surface of the hazelnut has a larger gray value because the difference is large, while the normal area on the surface of the hazelnut has a smaller gray value because the difference is small.
[0117] It can be understood that in this embodiment, the image to be detected is input into the image detection model to facilitate the detection of the reconstructed image output by any network node in the cascade feedback network. The generated reconstructed image has the characteristics of high pixel accuracy and good feature reconstruction effect, which makes the reconstruction error of the abnormal area relatively large. It is only necessary to separate the abnormal area from the normal area through simple standards to complete the detection of the surface abnormal area.
[0118] Embodiment 3
[0119] Please refer to Fig.10 In this embodiment, an image detection device is disclosed. The image detection device 3 mainly includes an image acquisition component 31, a processor 32 and a display 33, which are described below respectively.
[0120] The image acquisition component 31 is used to acquire the image to be detected of the object to be detected.
[0121] It should be noted that a CCD camera, a CMOS camera, a 3D camera or a video camera, and other grayscale or color cameras can be used to complete the acquisition of the image to be detected. If the camera / video camera takes a color image, the color image needs to be converted into a grayscale image to form the image to be detected. Of course, the image acquisition component 31 can also collect normal sample images of the object to be detected, thereby providing samples for the training of the cascade feedback network.
[0122] It should be noted that the normal sample images of the object to be detected are used to participate in the training of the cascade feedback network, and the image to be detected of the object to be detected is used to input the image detection model to identify the surface abnormal area in the image. In addition, the object to be detected involved can be a product on an assembly line, a part on a tool table, or a person, animal, plant or other object, which is not specifically limited here.
[0123] The processor 32 is connected to the image acquisition component 31, and the processor 32 is used to construct an image detection model through the construction method disclosed in Example 1, and / or obtain the surface abnormal area of the object to be detected in the image to be detected through the image detection method disclosed in Example 2.
[0124] The display 33 is connected to the processor 32, and is used to display the image to be detected and the surface abnormal area of the object to be detected.
[0125] In this implementation, see Fig.11 The processor 32 may include a model building module 321 and an anomaly detection module 322, which are described below respectively.
[0126] The model building module 321 is used to train the pre-established cascade feedback model using one or more normal sample images, and obtain the image detection model by updating the network parameters through a loss function. The cascade feedback network here includes multiple network nodes formed by multiple shallow autoencoders through cascade feedback, and the loss function is configured according to the structure of the cascade feedback network. For the specific functions of the model building module 321, please refer to steps S110-S130 in Example 1, which will not be repeated here.
[0127] The abnormality detection module 322 is connected to the model building module 321, and is used to input the image to be detected into the image detection model, and output the surface abnormal area of the object to be detected through detection processing. The specific functions of the abnormality detection module 322 can refer to steps S210-S230 in Example 2, which will not be repeated here.
[0128] Embodiment 4:
[0129] Based on the construction method disclosed in the first embodiment and the image detection method disclosed in the second embodiment, an image detection device is disclosed in this embodiment.
[0130] Please refer to Fig.12 The image detection device 4 mainly includes a memory 41 and a processor 42. The memory 41 is a computer-readable storage medium for storing a program, which may be a program code corresponding to the construction method S110-S130 in the first embodiment, or a program code corresponding to the image detection method S210-S230 in the second embodiment.
[0131] The processor 42 is connected to the memory 41 and is used to execute the program stored in the memory 41 in a corresponding manner. The functions implemented by the processor 42 can refer to the processor 32 in the third embodiment, and will not be described in detail here.
[0132] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above-mentioned embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above-mentioned functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above-mentioned functions can be implemented. In addition, when all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and can be downloaded or copied and saved in the memory of the local device, or the system of the local device is updated, and when the program in the memory is executed by the processor, all or part of the functions in the above-mentioned embodiments can be implemented.
[0133] The above specific examples are used to illustrate the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art, according to the concept of the present invention, some simple deductions, modifications or substitutions can be made.
Claims
1. A method for constructing an image detection model, characterized in that: include: Establish the structure of the cascade feedback network; The cascade feedback network includes a plurality of network nodes formed by cascade feedback of a plurality of shallow autoencoders, wherein each of the shallow autoencoders is used as a network node in the cascade feedback network, each network node is set as a node group, and the cascade feedback network is established by cascading each network node in the node group; A corresponding loss function is obtained according to the structural configuration of the cascade feedback network; wherein a first image reconstruction quality representation formula or a third image reconstruction quality representation formula is configured as a loss function, the first image reconstruction quality representation formula is obtained by calculating the Euclidean distance of images corresponding to the first and last two network nodes in the node grouping, the third image reconstruction quality representation formula is obtained by using the first image reconstruction quality representation formula and the second image reconstruction quality representation formula, and the second image reconstruction quality representation formula is obtained by calculating the variance statistics of a number of network nodes in the node grouping; The cascade feedback network is trained using a plurality of normal sample images of the object to be detected, and the network parameters of the cascade feedback network are updated using the loss function, so that an image detection model is obtained after the training is completed.
2. The construction method according to claim 1, characterized in that: The structure of establishing a cascade feedback network includes: Each shallow autoencoder is composed of convolutional neural units; the shallow autoencoder includes an encoder composed of a convolutional layer and a downsampling layer, and a decoder composed of an upsampling layer and a convolutional layer; the encoder is used to receive the image input by the shallow autoencoder and convert it into semantic information, and the decoder is used to restore the semantic information and output a reconstructed image; The plurality of shallow autoencoders are sequentially sorted, and the output of each shallow autoencoder is fed back to the input of the next shallow autoencoder, and each shallow autoencoder is used as a network node in the cascade feedback network; Each network node is set as a node group, and the cascade feedback network is established by cascading each network node in the node group.
3. The construction method according to claim 2, characterized in that: The obtaining of a corresponding loss function according to the structural configuration of the cascade feedback network includes: The Euclidean distances of the images corresponding to the first and last two network nodes in the node group are calculated, so as to obtain the first image reconstruction quality represented by the Euclidean distance of the image, which can be expressed as follows: Wherein, x0 is the image input to the first network node in the cascade feedback network, is the reconstructed image output by the first network node, The reconstructed image output by the Nth network node; The variance statistics of several network nodes in the node group are calculated to obtain the second image reconstruction quality represented by the variance statistics, which can be expressed as follows: Where m is the number of network nodes, c is the number of channels of the reconstructed image output by the kth network node, is the value of the i-th channel on the reconstructed image output by the k-th network node, is the average value of the ith channel on the m reconstructed images, and satisfy The third image reconstruction quality is calculated according to the first image reconstruction quality and the second image reconstruction quality, and is expressed by the formula: Loss d =Loss b +Loss var ; The first image reconstruction quality expression formula Loss b , or the third image reconstruction quality expression formula Loss d The loss function configured as the corresponding one of the cascade feedback network.
4. The construction method according to claim 3, characterized in that: The method of training the cascade feedback network using a plurality of normal sample images of the object to be detected, updating the network parameters of the cascade feedback network using the loss function, and obtaining the image detection model after the training is completed includes: Acquire a plurality of normal sample images of the object to be detected; wherein the normal sample images do not contain abnormal surface areas of the object to be detected; Using the normal sample image as the image input to the first network node in the cascade feedback network, and sequentially inputting each of the normal sample images into the cascade feedback network for training; The training is terminated when the difference between the before and after calculations of the loss function corresponding to the cascade feedback network is less than a preset threshold, or the corresponding loss function reaches a preset number of iterations; When the training is finished, the image detection model of the object to be detected is obtained by using the cascade feedback network with updated network parameters.
5. An image detection method based on cascade feedback, characterized in that: include: Acquire an image of the object to be detected; Inputting the image to be detected into the image detection model constructed by the construction method according to any one of claims 1 to 4, and detecting to obtain a reconstructed image output by any network node in the cascade feedback network; The reconstructed image is compared with the image to be detected to obtain the abnormal surface area of the object to be detected.
6. The image detection method according to claim 5, characterized in that: The step of comparing the reconstructed image with the image to be detected to obtain the abnormal surface area of the object to be detected includes: constructing an evaluation function of a surface abnormality region using the reconstructed image and the image to be detected; When the value of the evaluation function is greater than or equal to a predetermined value, determining that the reconstructed image contains the abnormal surface area of the object to be detected, and obtaining the abnormal surface area of the object to be detected by comparing the difference between the reconstructed image and the image to be detected; Output the image to be detected of the object to be detected and the abnormal surface area of the object to be detected.
7. The image detection method according to claim 6, characterized in that: The step of constructing an evaluation function of a surface abnormality region by using the reconstructed image and the image to be detected comprises: Obtain the reconstructed image output by each network node in the cascade feedback network after the image to be detected is input into the image detection model, so as to construct an evaluation function of the surface abnormal area, and the evaluation function is expressed by any of the following formulas: Among them, x0′ is the image to be detected, A reconstructed image output by any network node when the image to be detected is input into the cascade feedback network, is the output image of the intermediate network layer of any network node when the image to be detected is input into the cascade feedback network, M(x′0) is the output image of the intermediate network layer of the first network node when the image to be detected is input into the cascade feedback network, s is a set consisting of a number of network nodes, subscripts n, j, k are all serial numbers of network nodes; m is the number of a number of network nodes, c is the number of channels of the reconstructed image output by the kth network node, The value of the i-th channel on the reconstructed image output by the k-th network node when the image to be detected is input into the cascade feedback network, The average value of the ith channel on the m reconstructed images when the image to be detected is input into the cascade feedback network.
8. An image detection device, characterized in that: include: An image acquisition component, used for acquiring an image to be detected of an object to be detected; A processor connected to the image acquisition component, and configured to construct an image detection model by the construction method described in any one of claims 1 to 4, and / or to obtain a surface abnormal area of the object to be detected in the image to be detected by the image detection method described in any one of claims 5 to 7; A display is connected to the processor and is used to display the image to be detected and the surface abnormal area of the object to be detected.
9. The image detection device as claimed in claim 8, characterized in that: The processor includes a model building module and an anomaly detection module; The model building module is used to train a pre-established cascade feedback model using one or more normal sample images, and to obtain an image detection model by updating network parameters through a loss function; the cascade feedback network includes a plurality of network nodes formed by a plurality of shallow autoencoders through cascade feedback, and the loss function is configured according to the structure of the cascade feedback network; The abnormality detection module is connected to the model building module, and is used to input the image to be detected into the image detection model, and output the surface abnormal area of the object to be detected through detection processing.
10. A computer-readable storage medium, characterized in that: The method comprises a program which can be executed by a processor to implement the construction method according to any one of claims 1 to 4, and / or to implement the image detection method according to any one of claims 5 to 7.