Neural network training method and device, defect detection method and device and storage medium
By generating defect feature representations containing logical defects and training defect images, training neural networks solves the problem of difficulty in detecting logical defects in the prior art, accurately detecting and positioning of logical defects, and reducing production costs.
Patent Information
- Application Number
- CN202311608431.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
The existing defect detection methods are difficult to effectively simulate and detect logical defects due to the number of errors, types of errors, incorrect arrangement and combination of objects in the image, resulting in inaccurate detection and positioning of images with logical defects.
By inputting training images, obtaining their feature representations, and generating defect feature representations containing logical defects, further generating training defect images, using these images to train the neural network to adjust its parameters to detect logical defects.
It realizes accurate detection and positioning of logical defects, avoids the collection and manual labeling of large amounts of defect data, greatly reduces production costs and improves user experience.
Smart Images

Figure CN120047377A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a neural network training method, as well as a method, device, and computer-readable storage medium for defect detection using a neural network. Background Art
[0002] Defect detection is an important factor in product quality control. Since various defects will continuously occur during production, it is usually impossible to exhaustively list all the situations where various defects occur. How to robustly and automatically detect these defects is an urgent problem to be solved.
[0003] In current defect detection methods, a neural network model for defect detection is usually trained by a large number of simulated defect samples. However, since defect simulation is generally limited to structural defects such as surface scratches and depressions, it is unable to effectively simulate logical defects caused by, for example, the wrong number, wrong type, and wrong permutation and combination of certain object images, making it difficult to accurately detect and locate images with logical defects.
[0004] Therefore, an improved neural network training method and a defect detection method for accurately detecting logical defects in images are needed. Summary of the Invention
[0005] To solve the above technical problems, according to one aspect of the present invention, there is provided a neural network training method, including: inputting a training image and obtaining a feature representation of the training image; generating a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generating a training defect image at least according to the training image and the defect feature representation; and performing defect detection according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust parameters of the neural network.
[0006] According to another aspect of the present invention, there is provided a neural network training device, including: an input unit configured to input a training image and obtain a feature representation of the training image; a defect generation unit configured to generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; an image generation unit configured to generate a training defect image at least according to the training image and the defect feature representation; and a training unit configured to perform defect detection according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust parameters of the neural network.
[0007] According to another aspect of the present invention, there is provided a neural network training device, comprising: a processor; and a memory in which computer program instructions are stored, wherein when the computer program instructions are run by the processor, the processor is caused to perform the following steps: input a training image and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation, so as to train the neural network and adjust parameters of the neural network.
[0008] According to another aspect of the present invention, there is provided a computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the following steps are implemented: input a training image and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation, so as to train the neural network and adjust parameters of the neural network.
[0009] According to another aspect of the present invention, there is provided a defect detection method, comprising: input a to-be-detected image, use a neural network to reconstruct the to-be-detected image, obtain a reconstructed image, and obtain a feature representation of the reconstructed image; use the neural network to perform defect detection on the to-be-detected image according to the reconstructed image and the feature representation of the reconstructed image; wherein the neural network is trained by the following method: input a training image and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation, so as to train the neural network and adjust parameters of the neural network.
[0010] According to another aspect of the present invention, there is provided an apparatus for defect detection using a neural network, including: an input unit configured to input an image to be detected, reconstruct the image to be detected using a trained neural network, obtain a reconstructed image, and obtain a feature representation of the reconstructed image; a detection unit configured to perform defect detection on the image to be detected using the neural network according to the reconstructed image and the feature representation of the reconstructed image, wherein the neural network is trained in the following manner: input a training image and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust parameters of the neural network.
[0011] According to another aspect of the present invention, there is provided an apparatus for defect detection using a neural network, including: a processor; and a memory in which computer program instructions are stored, wherein when the computer program instructions are run by the processor, the processor is caused to perform the following steps: input an image to be detected, reconstruct the image to be detected using a neural network, obtain a reconstructed image, and obtain a feature representation of the reconstructed image; perform defect detection on the image to be detected using the neural network according to the reconstructed image and the feature representation of the reconstructed image; wherein the neural network is trained by the following method: input a training image and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust parameters of the neural network.
[0012] According to another aspect of the present invention, there is provided a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the following steps are implemented: input an image to be detected, reconstruct the image to be detected using a neural network, obtain a reconstructed image, and obtain a feature representation of the reconstructed image; perform defect detection on the image to be detected using the neural network according to the reconstructed image and the feature representation of the reconstructed image; wherein the neural network is trained by the following method: input a training image, and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation to train the neural network and adjust parameters of the neural network.
[0013] According to the above neural network training method, device and computer-readable storage medium of the present invention, and the method, device and computer-readable storage medium for defect detection using a neural network, it is possible to simulate a training defect image with logical defects by using a defect feature representation generated from a feature representation of a training image, and use the simulated training defect image to train a neural network. The present invention proposes a robust defect detection system, which realizes precise generation and robust detection of logical defects of different types of products or objects, avoids the collection of a large amount of defect data and the manual annotation process, greatly reduces the production cost, and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] By describing embodiments of the present invention in detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become clearer.
[0015] Figure 1 A flowchart showing a neural network training method according to an embodiment of the present invention;
[0016] Figure 2 A flowchart showing a method for defect detection using a neural network according to an embodiment of the present invention;
[0017] Figure 3 A schematic diagram showing a training image according to an example of an embodiment of the present invention;
[0018] Figure 4 A schematic diagram showing a feature representation of a training image according to an example of an embodiment of the present invention;
[0019] Figure 5Schematic diagram showing a defect contour generated after processing the contour of a training image according to an example of an embodiment of the present invention;
[0020] Figure 6 Shows an example according to an embodiment of the present invention, for Figure 3 Schematic diagram of the bending enhancement operation performed on the shown training image;
[0021] Figure 7 Shows an example according to an embodiment of the present invention, for Figure 4 The contour of the shown training image is subjected to a bending enhancement operation corresponding to Figure 6 Schematic diagram;
[0022] Figure 8 Shows an example according to an embodiment of the present invention, according to Figure 5 Example of a training defect image generated using an image generation network based on the shown defect contour;
[0023] Figure 9 Block diagram showing a neural network training device according to an embodiment of the present invention;
[0024] Figure 10 Block diagram showing a neural network training device according to an embodiment of the present invention;
[0025] Figure 11 Block diagram showing a device for defect detection using a neural network according to an embodiment of the present invention;
[0026] Figure 12 Block diagram showing a device for defect detection using a neural network according to an embodiment of the present invention. Detailed implementation manners
[0027] The following will describe a neural network training method, device, and computer-readable storage medium according to an embodiment of the present invention, as well as a method, device, and computer-readable storage medium for defect detection using a neural network, with reference to the accompanying drawings. In the drawings, the same reference numerals denote the same elements throughout. It should be understood that the embodiments described herein are merely illustrative and should not be construed as limiting the scope of the present invention.
[0028] In the defect detection method using a neural network, when training a neural network for defect detection, a large number of training images containing various defects are often introduced as samples to adjust the parameters of the neural network. However, in the current defect detection methods, a large number of simulated defect samples are usually used to train a neural network model for defect detection. However, during the assembly or processing of the product to be detected, some areas may have non-exhaustive anomalies such as the wrong number or type of objects, and the existing defect simulations are generally limited to structural defects such as surface scratches and depressions, but cannot effectively simulate logical defects caused by the wrong number, wrong type, wrong permutation and combination of objects in the product image, making it difficult to accurately detect and locate images with logical defects.
[0029] In an embodiment of the present invention, a robust neural network training and defect detection method, device, and storage medium are proposed. The neural network training and defect detection method of the embodiment of the present invention does not require manual annotation and can effectively simulate logical defects, so that the trained neural network can accurately detect the logical defects in the image. The method, device, and medium of the embodiment of the present invention can be applied to the detection and analysis scenarios of products such as breakfast boxes, screw tool kits, and cable plugs, and can also be applied to other computer vision tasks, such as defect detection on road surfaces, chip surfaces, or any other object surfaces. During the application process of the embodiment of the present invention, a convolutional neural network (CNN) that can extract more complex features at multiple semantic levels instead of limited low-level features can be used.
[0030] Figure 1 The flowchart of a neural network training method 100 according to an embodiment of the present invention is shown. The following refers to Figure 1 Describe the neural network training method according to an embodiment of the present invention.
[0031] In step S101, input a training image and obtain a feature representation of the training image.
[0032] In an embodiment of the present invention, in order to train the neural network's ability to recognize logical defects, optionally, the feature representation of the acquired training image may include a feature representation that can reflect the logical defects of the image. For example, the feature representation of the training image may include at least one of the contour of the training image, the image segmentation information of the training image, and the multi-modal information of the training image. In one example, the feature representation of the training image may include the contour of the training image. The contour of the training image can generally show the contours of the objects existing in the training image and reflect the position, quantity, and sorting of the objects to reflect the possible logical defects in the training image. In another example, the feature representation of the training image may include the image segmentation information of the training image. The image segmentation information of the training image can show the positions, quantities, sorting, etc. of the objects existing in the training image through an object segmentation process based on thresholds, graph theory, clustering, etc. to reflect the possible logical defects in the training image. In yet another example, the feature representation of the training image may include the multi-modal information of the training image. The multi-modal information of the training image may be information of multiple modalities of the training image. For example, it may include various perspectives of information such as the text description information and image description information of the objects in the training image obtained by using a multi-modal large model, etc., to reflect the positions, quantities, sorting, etc. of the objects existing in the training image and show the possible logical defects in the training image. In still another example, according to the specific application scenario of the embodiment of the present invention, the feature representation of the training image may further include various relevant information such as the color information, shape information, classification information, etc. of the training image that can reflect whether there are logical defects in the objects in the training image, and no limitations are made here.
[0033] In an embodiment of the present invention, optionally, training images without corresponding logical defects can be used to train a feature detection network applicable to the embodiment of the present invention to extract features from the training images and output the feature extraction results corresponding to the training images. For example, when the feature representation of the training image is the contour of the training image, a lightweight edge detection network with fewer parameters can be trained by using the training image in combination with a pre-trained edge detection network with relatively more parameters, so that the lightweight edge detection network obtains edge detection capabilities similar to those of the pre-trained edge detection network with more parameters by learning the output of the pre-trained edge detection network with more parameters in an unsupervised training manner. More specifically, for the same input training image, it is desired that the output of the lightweight edge detection network be as close as possible to the output of the pre-trained edge detection network after training. Therefore, the goal of the above training process is to minimize the output difference between the lightweight edge detection network and the pre-trained edge detection network. After training the lightweight edge detection network, the trained lightweight edge detection network can be used to perform edge detection on the input training image and output the contour corresponding to the training image. The above describes an example where the feature representation of the training image is the contour of the training image. In practical applications, different feature representations of the training images can be used to train and apply corresponding feature detection networks respectively to obtain the feature representation detection results corresponding to the training images.
[0034] In step S102, a defect feature representation of the training image is generated according to the feature representation of the training image, and the defect feature representation includes at least one logical defect of the training image.
[0035] In an embodiment of the present invention, generating a defect feature representation of the training image according to the feature representation of the training image may include: obtaining the feature representation of at least part of the region in the training image, modifying the logic of the feature representation of the at least part of the region, and obtaining the modified feature representation of the at least part of the region; fusing the modified feature representation of the at least part of the region with the feature representation of the training image to generate the defect feature representation. Specifically, on the basis of obtaining the feature representation of the training image, the feature representation of at least part of the region of the training image can be processed and modified accordingly, so that the modified feature representation has logical defects.
[0036] For example, when the feature representation of the training image includes the contour of the training image, image enhancement processing (such as various image operations like scaling, rotation, etc.) can be performed on the extracted contour of the training image to modify the position, quantity, order, etc. of the objects in at least part of the shown area, generating a modified feature representation of at least part of the area. For another example, when the feature representation of the training image includes the image segmentation information of the training image, the position, quantity, order, etc. of the objects shown in the image segmentation information extracted from at least part of the area can be modified to generate a modified feature representation of at least part of the area. For yet another example, when the feature representation of the training image includes the multi-modal information of the training image, the multi-modal information of the objects represented in at least part of the area can be adjusted, and the text or image descriptions of the position, quantity, order, etc. of the objects can be modified to generate a modified feature representation of at least part of the area. The above specific operations for modifying the feature representation of the training image are only examples, and different modifications reflecting logical defects can be made for different feature representations, which are not limited here.
[0037] In an embodiment of the present invention, optionally, the feature representation of the training image can be processed and modified multiple times to obtain multiple modified feature representations of at least part of the area. For example, multiple processes can be performed on the same type of feature representation of the training image, or multiple different types of feature representations of the training image can be processed separately, which are not limited here.
[0038] After obtaining one or more modified feature representations of at least part of the area, the one or more modified feature representations of at least part of the area can be fused with the feature representations of the training images without logical defects obtained previously. For example, these feature representations can be superimposed to obtain a defective feature representation containing logical defects; for another example, a part of these feature representations can be selected and spliced to form a defective feature representation; for yet another example, multiple modified feature representations can be overlapped, spliced, or a part of each can be selected, and part of the area of the corresponding feature representation of the training image can be replaced to obtain a fused defective feature representation.
[0039] The above describes an example of the generation process for the defect feature representation with logical defects. According to another embodiment of the present invention, the defect feature representation of the training image may include not only logical defects but also at least one structural defect of the training image, such as depressions, protrusions, distortions, etc. Optionally, the processing and generation process for the defect feature representation of the structural defect of the training image is similar to the above, and it can also be obtained by generating a modified feature representation with structural defects and superimposing it on the feature representation of the training image. In one example, a defect feature representation with logical defects can also be generated by replacing one or more relatively large foreground regions of the feature representation of the training image with a modified feature representation; while a defect feature representation with logical defects can be generated by replacing one or more relatively small foreground regions of the feature representation of the training image with a modified feature representation. Among them, the replacement regions for logical defects can be, for example, one or more relatively large rectangles, and the replacement regions for logical defects can be, for example, one or more relatively small polygons, etc., which are not limited here.
[0040] In step S103, a training defect image is generated based on at least the training image and the defect feature representation.
[0041] According to an embodiment of the present invention, the image generation network can be trained first based on the training image and the feature representation of the training image; subsequently, the trained image generation network is used to generate a training defect image based on the defect feature representation.
[0042] Optionally, during the process of training the image generation network, multiple samples can be used for training. For example, the image generation network can be trained by receiving the input training image and its corresponding feature representation of the training image, and further, by performing various enhancement operations on the training image and correspondingly generating the enhanced feature representation as a sample to train the image generation network. Optionally, the enhancement operation on the training image can include a bending enhancement operation on the training image. For example, the training image can be divided into grids, several nodes in the grid are randomly selected, and the selected nodes are randomly translated horizontally or vertically to be used as the input image for the image generation network. In addition, optionally, when the feature representation of the training image is an image, the same bending enhancement operation can also be performed on the feature representation of the training image according to the same horizontal or vertical translation offset as that for the training image, so as to correspondingly generate the enhanced feature representation. The image generation network can learn the correspondence between each image and the feature representation of the image through training, so as to be able to generate the corresponding image using the feature representation.
[0043] After obtaining the trained image generation network, the trained image generation network can be used to generate corresponding training defect images according to the defect feature representation.
[0044] In step S104, defect detection is performed according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation, so as to train the neural network and adjust the parameters of the neural network.
[0045] According to the embodiments of the present invention, the neural network can be trained and the parameters of the neural network can be adjusted by receiving the input training image, the feature representation of the training image, the training defect image generated by the foregoing method, and the corresponding defect feature representation, so that the loss function of the neural network converges as much as possible.
[0046] Optionally, the input training image, the feature representation of the training image, the training defect image, and the defect feature can be used as samples to train the anomaly localization network for defect detection in the neural network and adjust the parameters of the anomaly localization network. In one example, the anomaly localization network may include a reconstruction sub-network and a localization sub-network. The reconstruction sub-network is used to convert the input training defect image containing defects into a reconstructed image and corresponding feature representation without defects, and the localization sub-network locates the defects by calculating the difference between the reconstructed image and the corresponding feature representation converted by the reconstruction sub-network and the original input image.
[0047] More specifically, the reconstruction sub-network can be composed of an encoder and a decoder. The encoder is used to extract the feature map of the input image, and the decoder is used to reconstruct the feature map back to the original resolution. Among them, when the encoder extracts the feature map of the image, different levels of feature maps can be extracted by using convolution, normalization, and pooling processes with different parameters. Specifically, the input image can be first convolved with a convolutional kernel to obtain a convolutional map; then, the convolutional map can be normalized by using a traditional rectified linear unit and batch normalization method to obtain a normalized convolutional map; finally, a maximum or average pooling process can be applied to the normalized convolutional map. In order to obtain rich multi-scale features, the reconstruction sub-network can adjust the relevant parameters and repeat the above process multiple times to extract multi-scale feature maps through multiple downsampling processes. The decoder can restore the resolution of the feature map of the image through corresponding convolution, normalization, and upsampling processes. The reconstruction sub-network can combine the feature maps with the same resolution of the encoder and decoder and uses multiple sets of convolutional processes. When the similarity between the reconstructed image converted by the reconstruction sub-network and the original input training image is higher, the probability that the training image does not contain defects is also higher, and vice versa. The structure of the localization sub-network of the anomaly localization network is similar to the above-mentioned reconstruction sub-network, but can further have cross-layer connection operations for fusing the feature maps of the same scale of the encoder and decoder in the localization sub-network.
[0048] During the defect detection process, the anomaly localization network can not only detect the logical defects of the training defect images modified and generated in the embodiments of the present invention, but also detect structural defects at the same time. Optionally, after the trained anomaly localization network performs defect detection on an image, it can output the specific location of the defect and can also output an estimated value for the degree of the defect at the same time.
[0049] According to the above neural network training method of the embodiments of the present invention, it is possible to simulate the training defect images with logical defects by using the defect feature representations generated from the feature representations of the training images, and use the simulated training defect images to train the neural network. The present invention proposes a robust defect detection system, which realizes the precise generation and robust detection of logical defects for different types of products or objects, avoids the collection of a large amount of defect data and the manual annotation process, greatly reduces the production cost, and improves the user experience.
[0050] Figure 2 FIG. 200 shows a flowchart of a method for defect detection using a neural network according to an embodiment of the present invention. In this embodiment, a neural network trained through the Figure 1 process shown can be used for defect detection. The following will refer to Figure 2 to describe the method for defect detection using a neural network according to an embodiment of the present invention.
[0051] In step S201, an image to be detected is input, and the neural network is used to reconstruct the image to be detected, obtain a reconstructed image, and obtain a feature representation of the reconstructed image.
[0052] In an embodiment of the present invention, the image to be detected can be input into the reconstruction sub-network of the anomaly localization network included in the neural network trained through the Figure 1 process shown for reconstruction. Optionally, the image to be detected can be an image without logical defects and / or structural defects, or an image with one or more logical defects and / or structural defects. During the process of reconstructing the image to be detected, a reconstructed image without defects can be obtained, and at the same time, a corresponding feature representation of the reconstructed image can be obtained for subsequent defect detection and localization processes.
[0053] As described above, the reconstruction sub-network in the anomaly localization network can be used to convert the input image to be detected into a reconstructed image and a corresponding feature representation. More specifically, the reconstruction sub-network can be composed of an encoder and a decoder. The encoder is used to extract the feature map of the input image to be detected, and the decoder is used to reconstruct the feature map back to the original resolution. Among them, when the encoder extracts the feature map of the image to be detected, different levels of feature maps can be extracted by using convolution, normalization, and pooling processes with different parameters. Specifically, the image to be detected can be first convolved with a convolution kernel to obtain a convolution map; then, the convolution map can be normalized by using a traditional linear correction unit and batch normalization method to obtain a normalized convolution map; finally, a maximum or average pooling process can be applied to the normalized convolution map. In order to obtain rich multi-scale features, the reconstruction sub-network can adjust relevant parameters and repeat the above process multiple times to extract multi-scale feature maps through multiple downsampling processes. The decoder can restore the resolution of the feature map of the image through corresponding convolution, normalization, and upsampling processes. This reconstruction sub-network can combine the feature maps with the same resolution of the encoder and the decoder and uses multiple sets of convolution processes.
[0054] In step S202, based on the reconstructed image and the feature representation of the reconstructed image, the neural network is used to detect defects in the image to be detected.
[0055] According to an embodiment of the present invention, defects can be located by using the localization sub-network in the anomaly localization network based on the reconstructed image and the feature representation of the reconstructed image. The structure of the localization sub-network of the anomaly localization network is similar to the above-mentioned reconstruction sub-network, but can further have a cross-layer connection operation for fusing the feature maps with the same scale of the encoder and the decoder in the localization sub-network.
[0056] During the defect detection process, the anomaly localization network can not only detect the logical defects of the image to be detected, but also detect structural defects simultaneously. Optionally, after the anomaly localization network performs defect detection on the image to be detected, it can output the specific location of the defect and can also output an estimated value for the degree of the defect for the user's reference.
[0057] According to an embodiment of the present invention, the neural network in the defect detection method is trained by using the steps as Figure 1 shown. The training method of the neural network may include: inputting a training image and obtaining a feature representation of the training image; generating a defect feature representation of the training image according to the feature representation of the training image, where the defect feature representation includes at least one logical defect of the training image; generating a training defect image at least according to the training image and the defect feature representation; performing defect detection according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust the parameters of the neural network. The specific training process has been described in detail in the steps as Figure 1 shown and will not be elaborated here.
[0058] The defect detection method according to an embodiment of the present invention can perform defect detection by using a trained neural network. During the training process of the neural network, a defect feature representation generated by using the feature representation of the training image can be used to simulate a training defect image with logical defects, and the simulated training defect image can be used to train the neural network. The present invention proposes a robust defect detection system, which realizes the precise generation and robust detection of logical defects of different types of products or objects, avoids the collection of a large amount of defect data and the manual annotation process, greatly reduces the production cost, and improves the user experience.
[0059] The following shows the specific implementation processes of a neural network training method and a defect detection method according to an example of an embodiment of the present invention.
[0060] In this example of the embodiment of the present invention, it is applied to the detection and analysis scenario of objects in a breakfast box. In this example, first, a training image is input, and a feature representation of the training image is obtained. Figure 3 shows a schematic diagram of a training image according to an example of an embodiment of the present invention. As Figure 3 shown, the training image adopted in this example can be an image without logical defects, that is, the positions, quantities, permutations, etc. of the objects in the training image are all correct. After obtaining the training image, a corresponding feature representation of the training image can also be obtained. Optionally, in this example, the feature representation of the training image can be the contour of the training image. Figure 4Schematic diagram of the feature representation of a training image showing an example according to an embodiment of the present invention. As Figure 4 shown, the contour of the training image in Figure 3 can be detected and extracted by a trained edge detection network. In this example, as described above, a pre-trained edge detection network with relatively many parameters can be combined with a training image to train a lightweight edge detection network with fewer parameters, so that the lightweight edge detection network obtains an edge detection ability similar to that of the pre-trained edge detection network and extracts Figure 3 the contour of the training image in
[0061] In this example, the feature representation being the contour of the training image is only an example. In addition, when the feature representation of the training image is image segmentation information, the feature representation of the training image can be obtained according to the result of image segmentation of the training image in Figure 3 ; and when the feature representation of the training image is multi-modal information, for example, text description information can be extracted from the training image in Figure 3 as the feature representation of the training image. In the text description information of the training image, relevant descriptions such as the category, number, and positional relationship of the items in the training image can be included. For example, the training image in Figure 3 can be described as "In a breakfast box, on the upper part of the left partition is an orange, in the middle is an orange, and on the lower part is a peach; on the upper part of the right partition are some grains, and on the lower part are some nuts and dried fruits".
[0062] Subsequently, a defect feature representation of the training image is generated according to the feature representation of the training image, and the defect feature representation includes at least one logical defect of the training image. Specifically, in this example, the feature representation of at least part of the area in the training image, such as the contour of at least part of the area, can be obtained, and the logic of the feature representation of the at least part of the area is modified to obtain the modified feature representation of the at least part of the area; the modified feature representation of the at least part of the area is fused with the feature representation of the training image to generate the defect feature representation.
[0063] For example, image enhancement processing (such as various image operations like scaling, rotation, etc.) can be performed on the contours of the extracted training images to modify the positions, quantities, orders, etc. of the objects in at least some of the shown regions, generating a feature representation of at least some of the modified regions. Additionally, optionally, the feature representation of the training image can also be processed and modified multiple times to obtain multiple feature representations of at least some of the modified regions. After obtaining one or more feature representations of at least some of the modified regions, the one or more feature representations of at least some of the modified regions can be fused with the feature representations of the previously obtained training images without logical defects. For example, these feature representations can be superimposed to obtain a defective feature representation containing logical defects; for another example, a part of these feature representations can be selected and spliced together to form a defective feature representation; for yet another example, multiple modified feature representations can be overlapped, spliced, or a part of each can be selected, and a partial region of the feature representation of the corresponding training image can be replaced to obtain a fused defective feature representation.
[0064] Figure 5 Figure 4 shows a schematic diagram of a defective contour generated after processing the contour of a training image according to an example of an embodiment of the present invention. Thus, Figure 5 At multiple positions of the breakfast box contour in [Figure 4], such as the upper left corner, lower left corner, and lower right corner regions, are all Figure 4 different, and are different in terms of the type, quantity, position, arrangement and combination order, etc. of the objects, thus Figure 4 compared with the contour of the training image in [Figure 4], multiple logical defects are generated.
[0065] Of course, Figure 5 the defective contour in [Figure 4] can also include structural defects caused by depressions, protrusions, distortions, etc. in the image. Figure 5 The defective contour in [Figure 4] can be a feature representation with both logical defects and structural defects, which is not limited herein.
[0066] Subsequently, at least a training defective image is generated based on the training image and the defective feature representation.
[0067] According to an example of an embodiment of the present invention, the image generation network can be trained first according to the training image and the feature representation of the training image; subsequently, using the trained image generation network, a training defect image can be generated according to the defect feature representation. During the process of training the image generation network, various samples can be used for training. For example, the image generation network can be trained by receiving the input training image and its corresponding feature representation of the training image, and further, by performing various enhancement operations on the training image and correspondingly generating the feature representation after the enhancement operation as a sample to train the image generation network. Optionally, the enhancement operation on the training image can include a bending enhancement operation on the training image. For example, the training image can be divided into grids, several nodes in the grid can be randomly selected, and the selected nodes can be randomly translated horizontally or vertically to be used as the input image for the image generation network. In addition, optionally, when the feature representation of the training image is an image, the same bending enhancement operation can also be performed on the feature representation of the training image according to the same horizontal or vertical translation offset amount for the training image, so as to correspondingly generate the feature representation after the enhancement operation. Figure 6 shows an example according to an embodiment of the present invention, for Figure 3 the schematic diagram of the bending enhancement operation performed on the training image shown. Figure 7 shows an example according to an embodiment of the present invention, for Figure 4 the outline of the training image shown in Figure 6 the corresponding bending enhancement operation schematic diagram. As Figure 6 , Figure 7 shown, by performing the same bending enhancement operation on the training image and the corresponding outline, the image and outline after the bending enhancement operation are generated. During the process of training the image generation network, the training image, the outline of the training image, the image after the bending enhancement operation, and the outline after the bending enhancement operation can all be used as samples to obtain the trained image generation network.
[0068] After obtaining the trained image generation network, the trained image generation network can be used to generate the corresponding training defect image according to the defect feature representation. Figure 8 shows an example according to an embodiment of the present invention, according to Figure 5 the example of the training defect image generated by the image generation network using the defect outline shown.
[0069] Finally, defect detection is performed according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust the parameters of the neural network.
[0070] According to an embodiment of the present invention, by receiving the input such as Figure 3The training images shown, Figure 4 the contours of the training images shown, and the generated Figure 8 the training defect images shown and the corresponding Figure 5 defect contours shown are used to train a neural network and adjust the parameters of the neural network so that the loss function of the neural network converges as much as possible.
[0071] Optionally, the above input images can be used as samples to train the anomaly localization network for defect detection in the neural network and adjust the parameters of the anomaly localization network. In one example, the anomaly localization network can include a reconstruction sub-network and a localization sub-network. The reconstruction sub-network is used to convert the input training defect image containing defects into a reconstructed image without defects and the corresponding feature representation. The localization sub-network locates the defects by calculating the differences between the reconstructed image converted by the reconstruction sub-network, the corresponding feature representation and the original input image.
[0072] During the defect detection process, the anomaly localization network can not only detect the logical defects of the training defect images modified and generated in the embodiments of the present invention, but also detect the structural defects at the same time. Optionally, after the trained anomaly localization network performs defect detection on an image, it can output the specific location of the defect and can also output an estimated value for the degree of the defect at the same time.
[0073] In another example of the embodiments of the present invention, after a neural network for defect detection has been trained, the neural network can be used to perform defect detection on the input image to be detected. Specifically, the image to be detected can be input first, and the neural network is used to reconstruct the image to be detected to obtain a reconstructed image, and the feature representation of the reconstructed image is obtained, such as the contour of the reconstructed image.
[0074] Subsequently, defect detection can be performed on the image to be detected according to the reconstructed image and the feature representation of the reconstructed image. During the defect detection process, the anomaly localization network in the neural network can not only detect the logical defects of the image to be detected, but also detect the structural defects at the same time. Optionally, after the anomaly localization network performs defect detection on the image to be detected, it can output the specific location of the defect and can also output an estimated value for the degree of the defect at the same time for the user to refer to.
[0075] Next, with reference to Figure 9 a neural network training device according to an embodiment of the present invention will be described. Figure 9 The block diagram of a neural network training device 900 according to an embodiment of the present invention is shown. As Figure 9As shown, the neural network training device 900 includes an input unit 910, a defect generation unit 920, an image generation unit 930, and a training unit 940. In addition to these units, the neural network training device 900 may further include other components. However, since these components are not related to the content of the embodiments of the present invention, their illustrations and descriptions are omitted here. In addition, since the specific details of the following operations performed by the neural network training device 900 according to the embodiments of the present invention are the same as the details described above with reference to Figure 1 the description is the same, the repeated description of the same details is omitted here to avoid repetition.
[0076] Figure 9 The input unit 910 of the neural network training device 900 in []] inputs the training image and obtains the feature representation of the training image.
[0077] In the embodiments of the present invention, in order to train the neural network's ability to recognize logical defects, optionally, the feature representation of the training image obtained by the input unit 910 may include a feature representation that can reflect the logical defects of the image. For example, the feature representation of the training image may include at least one of the contour of the training image, the image segmentation information of the training image, and the multi-modal information of the training image. In one example, the feature representation of the training image may include the contour of the training image. The contour of the training image can generally show the contours of the objects existing in the training image and reflect the position, quantity, and sorting of the objects to reflect the possible logical defects of the training image. In another example, the feature representation of the training image may include the image segmentation information of the training image. The image segmentation information of the training image can show the positions, quantities, sorting, etc. of the objects existing in the training image through an object segmentation process based on thresholds, graph theory, clustering, etc., to reflect the possible logical defects of the training image. In still another example, the feature representation of the training image may include the multi-modal information of the training image. The multi-modal information of the training image may be information of multiple modalities of the training image. For example, it may include various angle information such as the text description information and image description information of the objects in the training image obtained by using a multi-modal large model, so as to reflect the positions, quantities, sorting, etc. of the objects existing in the training image and show the possible logical defects of the training image. In yet another example, according to the specific application scenario of the embodiments of the present invention, the feature representation of the training image may further include various relevant information such as the color information, shape information, classification information, etc. of the training image that can reflect whether there are logical defects in the objects in the training image, and no limitations are made here.
[0078] In an embodiment of the present invention, optionally, the input unit 910 may use training images for training that do not have corresponding logic defects to train a feature detection network applicable to the embodiment of the present invention, so as to extract features from the training images and output the feature extraction results corresponding to the training images. For example, when the feature representation of the training image is the contour of the training image, a pre-trained edge detection network with relatively more parameters can be combined with the training image to train a lightweight edge detection network with fewer parameters, so that the lightweight edge detection network obtains edge detection capabilities similar to those of the pre-trained edge detection network with more parameters by learning the output of the pre-trained edge detection network with more parameters in an unsupervised training manner. More specifically, for the same input training image, it is desired that the output of the lightweight edge detection network is as close as possible to the output of the pre-trained edge detection network after training. Therefore, the objective of the above training process is to minimize the output difference between the lightweight edge detection network and the pre-trained edge detection network. After training the lightweight edge detection network, the trained lightweight edge detection network can be used to perform edge detection on the input training image and output the contour corresponding to the training image. The above describes an example where the feature representation of the training image is the contour of the training image. In practical applications, different feature representations of training images can be used to train and apply corresponding feature detection networks respectively to obtain the detection results of the feature representations of the corresponding training images.
[0079] The defect generation unit 920 generates a defect feature representation of the training image according to the feature representation of the training image, and the defect feature representation includes at least one logic defect of the training image.
[0080] In an embodiment of the present invention, the defect generation unit 920 obtains the feature representation of at least part of the regions in the training image, modifies the logic of the feature representation of the at least part of the regions, and obtains the modified feature representation of the at least part of the regions; fuses the modified feature representation of the at least part of the regions with the feature representation of the training image to generate the defect feature representation. Specifically, on the basis of obtaining the feature representation of the training image, corresponding processing and modification can be performed on the feature representation of at least part of the regions of the training image, so that the modified feature representation has logic defects.
[0081] For example, when the feature representation of the training image includes the contour of the training image, the defect generation unit 920 may perform image enhancement processing on the extracted contour of the training image (such as performing various image operations such as scaling, rotation, etc.) to modify the position, quantity, order, etc. of the objects in at least some of the shown regions, and generate a feature representation of at least some of the modified regions. For another example, when the feature representation of the training image includes the image segmentation information of the training image, the position, quantity, order, etc. of the objects shown in the image segmentation information extracted from at least some of the regions may be modified to generate a feature representation of at least some of the modified regions. For yet another example, when the feature representation of the training image includes the multi-modal information of the training image, the multi-modal information of the objects represented in at least some of the regions may be adjusted, and the text or image descriptions of the position, quantity, order, etc. of the objects may be modified to generate a feature representation of at least some of the modified regions. The above specific operations for modifying the feature representation of the training image are only examples, and different modifications reflecting logical defects may be performed for different feature representations, which are not limited herein.
[0082] In an embodiment of the present invention, optionally, the feature representation of the training image may be processed and modified multiple times to obtain multiple feature representations of at least some of the modified regions. For example, multiple processes may be performed on the same type of feature representation of the training image, or multiple different types of feature representations of the training image may be processed separately, which are not limited herein.
[0083] After obtaining one or more feature representations of at least some of the modified regions, the one or more feature representations of at least some of the modified regions may be fused with the feature representations of the training images without logical defects obtained previously. For example, these feature representations may be superimposed to obtain a defect feature representation containing logical defects; for another example, a part of these feature representations may be selected and spliced into a defect feature representation; for yet another example, multiple modified feature representations may be overlapped, spliced, or a part of each of them may be selected, and a partial region of the feature representation of the corresponding training image may be replaced to obtain a fused defect feature representation.
[0084] The above describes an example of the generation process for the defect feature representation with logical defects. According to another embodiment of the present invention, the defect feature representation of the training image may include not only logical defects but also at least one structural defect such as indentation, protrusion, distortion, etc. of the training image. Optionally, the processing and generation process for the defect feature representation of the structural defect of the training image is similar to the above, and it can also be obtained by generating a modified feature representation with a structural defect and superimposing it on the feature representation of the training image. In one example, a defect feature representation with logical defects can also be generated by replacing one or more relatively large foreground regions of the feature representation of the training image with the modified feature representation; and a defect feature representation with logical defects can be generated by replacing one or more relatively small foreground regions of the feature representation of the training image with the modified feature representation. Among them, the replacement region for logical defects can be, for example, one or more relatively large rectangles, and the replacement region for logical defects can be, for example, one or more relatively small polygons, etc., which are not limited herein.
[0085] The image generation unit 930 generates a training defect image at least based on the training image and the defect feature representation.
[0086] According to an embodiment of the present invention, the image generation unit 930 may first train an image generation network based on the training image and the feature representation of the training image; then, using the trained image generation network, generate a training defect image based on the defect feature representation.
[0087] Optionally, during the process of training the image generation network, multiple samples can be used for training. For example, the image generation network can be trained by receiving the input training image and its corresponding feature representation of the training image, and further, by performing various enhancement operations on the training image and correspondingly generating the enhanced feature representation as a sample to train the image generation network. Optionally, the enhancement operation on the training image may include a bending enhancement operation on the training image. For example, the training image can be divided into grids, several nodes in the grids are randomly selected, and the selected nodes are randomly translated horizontally or vertically to be used as the input image for the image generation network. In addition, optionally, when the feature representation of the training image is an image, the same bending enhancement operation can also be performed on the feature representation of the training image according to the same horizontal or vertical translation offset as that for the training image to correspondingly generate the enhanced feature representation. The image generation network can learn the correspondence between each image and the feature representation of the image through training so as to be able to generate the corresponding image using the feature representation.
[0088] After obtaining the trained image generation network, the trained image generation network can be utilized to generate corresponding training defect images according to the defect feature representation.
[0089] The training unit 940 performs defect detection based on the training images, the feature representations of the training images, the training defect images, and the defect feature representations, so as to train the neural network and adjust the parameters of the neural network.
[0090] According to an embodiment of the present invention, the training unit 940 can train the neural network and adjust the parameters of the neural network by receiving the input training images, the feature representations of the training images, the training defect images generated by the foregoing method, and the corresponding defect feature representations, such that the loss function of the neural network converges as much as possible.
[0091] Optionally, the training unit 940 can use the input training images, the feature representations of the training images, the training defect images, and the defect features as samples to train an anomaly localization network for defect detection in the neural network and adjust the parameters of the anomaly localization network. In one example, the anomaly localization network can include a reconstruction sub-network and a localization sub-network. The reconstruction sub-network is used to convert the input training defect images containing defects into reconstructed images without defects and corresponding feature representations, and the localization sub-network locates the defects by calculating the differences between the reconstructed images and the corresponding feature representations converted by the reconstruction sub-network and the original input images.
[0092] More specifically, the reconstruction sub-network can be composed of an encoder and a decoder. The encoder is used to extract the feature map of the input image, and the decoder is used to reconstruct the feature map back to the original resolution. Among them, when the encoder extracts the feature map of the image, different levels of feature maps can be extracted by using convolution, normalization, and pooling processes with different parameters. Specifically, the input image can be first convolved with a convolutional kernel to obtain a convolutional map; then, the convolutional map can be normalized by using a traditional rectified linear unit and batch normalization method to obtain a normalized convolutional map; finally, a maximum or average pooling process can be applied to the normalized convolutional map. In order to obtain rich multi-scale features, the reconstruction sub-network can adjust relevant parameters and repeat the above process multiple times to extract multi-scale feature maps through multiple downsampling processes. The decoder can restore the resolution of the feature map of the image through corresponding convolution, normalization, and upsampling processes. This reconstruction sub-network can combine the feature maps with the same resolution of the encoder and decoder and uses multiple groups of convolution processes. When the similarity between the reconstructed image converted by the reconstruction sub-network and the original input training image is higher, the probability that the training image does not contain defects is also higher, and vice versa. The structure of the localization sub-network of the anomaly localization network is similar to the above-mentioned reconstruction sub-network, but it can further have cross-layer connection operations to fuse the feature maps with the same scale of the encoder and decoder in the localization sub-network.
[0093] During the defect detection process, the anomaly localization network can not only detect the logical defects of the training defect images modified and generated in the embodiments of the present invention, but also detect structural defects at the same time. Optionally, after the trained anomaly localization network performs defect detection on an image, it can output the specific location of the defect and can also output an estimated value for the degree of the defect at the same time.
[0094] According to the above neural network training device of the embodiments of the present invention, it is possible to simulate training defect images with logical defects by using the defect feature representation generated from the feature representation of the training images, and use the simulated training defect images to train the neural network. The present invention proposes a robust defect detection system, which realizes the precise generation and robust detection of logical defects for different types of products or objects, avoids the collection of a large amount of defect data and the manual annotation process, greatly reduces the production cost, and improves the user experience.
[0095] Next, refer to Figure 10 to describe the neural network training device according to the embodiments of the present invention. Figure 10 FIG. shows a block diagram of a neural network training device 1000 according to an embodiment of the present invention. As Figure 10 shown, the device 1000 can be a computer or a server.
[0096] As Figure 10As shown, the neural network training device 1000 includes one or more processors 1010 and a memory 1020. Of course, in addition, the neural network training device 1000 may also include an input device, an output device (not shown), etc. These components can be interconnected through a bus system and / or other forms of connection mechanisms. It should be noted that Figure 10 The components and structure of the neural network training device 1000 shown are exemplary, not restrictive. According to needs, the neural network training device 1000 may also have other components and structures.
[0097] The processor 1010 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can utilize the computer program instructions stored in the memory 820 to execute the desired functions, which may include: inputting training images and obtaining feature representations of the training images; generating defect feature representations of the training images according to the feature representations of the training images, where the defect feature representations include at least one logical defect of the training images; generating training defect images at least according to the training images and the defect feature representations; performing defect detection according to the training images, the feature representations of the training images, the training defect images, and the defect feature representations to train the neural network and adjust the parameters of the neural network.
[0098] The memory 1020 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1010 can run the program instructions to implement the functions of the neural network training device in the embodiments of the present invention described above and / or other desired functions, and / or can execute the neural network training method according to the embodiments of the present invention. Various application programs and various data can also be stored in the computer-readable storage media.
[0099] Next, a computer-readable storage medium according to an embodiment of the present invention is described, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the following steps are implemented: inputting training images and obtaining feature representations of the training images; generating defect feature representations of the training images according to the feature representations of the training images, where the defect feature representations include at least one logical defect of the training images; generating training defect images at least according to the training images and the defect feature representations; performing defect detection according to the training images, the feature representations of the training images, the training defect images, and the defect feature representations to train the neural network and adjust the parameters of the neural network.
[0100] Next, with reference to Figure 11 the apparatus for defect detection using a neural network according to an embodiment of the present invention will be described. Figure 11 The block diagram of the apparatus 1100 for defect detection using a neural network according to an embodiment of the present invention is shown. As Figure 11 shown, the apparatus 1100 for defect detection using a neural network includes an input unit 1110 and a detection unit 1120. In addition to these units, the apparatus 1100 may further include other components. However, since these components are not related to the content of the embodiment of the present invention, their illustrations and descriptions are omitted here. In addition, since the specific details of the following operations performed by the apparatus 1100 according to the embodiment of the present invention are the same as the details described with reference to Figure 2 the description above, the repeated description of the same details is omitted here to avoid repetition.
[0101] Figure 11 The input unit 1110 of the apparatus 1100 for defect detection using a neural network in
[0102] inputs the image to be detected, and uses the neural network to reconstruct the image to be detected, obtains the reconstructed image, and obtains the feature representation of the reconstructed image. Figure 1 In the embodiment of the present invention, the input unit 1110 may input the image to be detected into the reconstruction sub-network of the anomaly localization network included in the neural network trained through the process shown in
[0103] As described above, the reconstruction sub-network in the anomaly localization network can be used to convert the input image to be detected into a reconstructed image and corresponding feature representations. More specifically, the reconstruction sub-network can be composed of an encoder and a decoder. The encoder is used to extract the feature map of the input image to be detected, and the decoder is used to reconstruct the feature map back to the original resolution. Among them, when the encoder extracts the feature map of the image to be detected, different levels of feature maps can be extracted by using convolution, normalization, and pooling processes with different parameters. Specifically, the convolution kernel can be first used to perform convolution on the image to be detected to obtain a convolution map; then, the traditional linear rectification unit and batch normalization method can be used to normalize the convolution map to obtain a normalized convolution map; finally, the maximum or average pooling process can be applied to the normalized convolution map. To obtain rich multi-scale features, the reconstruction sub-network can adjust the relevant parameters and repeat the above process multiple times to extract multi-scale feature maps through multiple downsampling processes. The decoder can restore the resolution of the feature map of the image through corresponding convolution, normalization, and upsampling processes. The reconstruction sub-network can combine the feature maps with the same resolution of the encoder and decoder and uses multiple sets of convolution processes.
[0104] The detection unit 1120 uses the neural network to perform defect detection on the image to be detected according to the reconstructed image and the feature representation of the reconstructed image.
[0105] According to an embodiment of the present invention, the defect can be located by using the localization sub-network in the anomaly localization network according to the reconstructed image and the feature representation of the reconstructed image. The structure of the localization sub-network of the anomaly localization network is similar to the above-mentioned reconstruction sub-network, but can further have a cross-layer connection operation for fusing the feature maps of the same scale of the encoder and decoder in the localization sub-network.
[0106] During the defect detection process, the anomaly localization network can not only detect the logical defects of the image to be detected, but also detect the structural defects at the same time. Optionally, after the anomaly localization network performs defect detection on the image to be detected, it can output the specific location of the defect and can also output an estimated value for the degree of the defect at the same time for the user to refer to.
[0107] According to an embodiment of the present invention, the neural network in the defect detection device is trained by using Figure 1trained according to the steps shown. The training method of the neural network may include: inputting training images and obtaining feature representations of the training images; generating defect feature representations of the training images according to the feature representations of the training images, where the defect feature representations include at least one logical defect of the training images; generating training defect images based on at least the training images and the defect feature representations; performing defect detection according to the training images, the feature representations of the training images, the training defect images, and the defect feature representations to train the neural network and adjust parameters of the neural network. The specific training process has been described in detail in Figure 1 the steps shown and will not be elaborated here.
[0108] According to the defect detection device of an embodiment of the present invention, defect detection can be performed using a trained neural network. During the training process of the neural network, defect feature representations generated by using the feature representations of training images can be used to simulate training defect images with logical defects, and the simulated training defect images can be used to train the neural network. The present invention proposes a robust defect detection system, which realizes precise generation and robust detection of logical defects of different types of products or objects, avoids the collection of a large amount of defect data and the manual annotation process, greatly reduces the production cost, and improves the user experience.
[0109] Next, with reference to Figure 12 describe a device for defect detection using a neural network according to an embodiment of the present invention. Figure 12 FIG. shows a block diagram of a device 1200 for defect detection using a neural network according to an embodiment of the present invention. As Figure 12 shown, the device 1200 may be a computer or a server.
[0110] As Figure 12 shown, the device 1200 includes one or more processors 1210 and a memory 1220. Of course, in addition, the device 1200 may also include an input device, an output device (not shown), etc., and these components may be interconnected through a bus system and / or other forms of connection mechanisms. It should be noted that Figure 12 the components and structure of the device 1200 shown are only exemplary and not restrictive. According to needs, the device 1200 may also have other components and structures.
[0111] The processor 1210 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may utilize the computer program instructions stored in the memory 1020 to execute desired functions, which may include: inputting an image to be detected, reconstructing the image to be detected using a neural network, obtaining a reconstructed image, and obtaining a feature representation of the reconstructed image; performing defect detection on the image to be detected using the neural network according to the reconstructed image and the feature representation of the reconstructed image; wherein, the neural network is trained by the following method: inputting a training image, and obtaining a feature representation of the training image; generating a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; generating a training defect image at least according to the training image and the defect feature representation; performing defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation to train the neural network and adjust the parameters of the neural network.
[0112] The memory 1220 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 1210 may run the program instructions to implement the functions of the device for defect detection using a neural network in the embodiments of the present invention described above and / or other desired functions, and / or may execute the method for defect detection using a neural network according to the embodiments of the present invention. Various application programs and various data may also be stored in the computer-readable storage media.
[0113] Next, a computer-readable storage medium according to an embodiment of the present invention is described, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the following steps are implemented: input an image to be detected, reconstruct the image to be detected using a neural network, obtain a reconstructed image, and obtain a feature representation of the reconstructed image; perform defect detection on the image to be detected using the neural network according to the reconstructed image and the feature representation of the reconstructed image; wherein, the neural network is trained by the following method: input a training image, and obtain a feature representation of the training image; generate a defect feature representation of the training image according to the feature representation of the training image, and the defect feature representation includes at least one logical defect of the training image; generate a training defect image at least according to the training image and the defect feature representation; perform defect detection according to the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust parameters of the neural network.
[0114] Of course, the above specific embodiments are merely examples and not limitations, and those skilled in the art can combine and combine some steps and devices from the above separately described embodiments according to the concept of the present invention to achieve the effects of the present invention. Such combined embodiments are also included in the present invention, and such combinations are not described one by one here.
[0115] Note that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. In addition, the above specific details of the invention are only for the purpose of illustration and easy understanding, and not limitations. The above details do not limit the present invention to necessarily adopt the above specific details to be implemented.
[0116] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present invention are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used here refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used here refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0117] The step flowcharts and the above method descriptions in the present invention are only illustrative examples and do not intend to require or imply that the steps of each embodiment must be carried out in the given order. As those skilled in the art will recognize, the steps in the above embodiments can be carried out in any order. Words such as "subsequently", "then", "next", etc. do not intend to limit the order of the steps; these words are only used to guide the reader through the description of the method. In addition, any reference to a singular element using, for example, the articles "a", "an", or "the" is not to be construed as limiting that element to the singular.
[0118] In addition, the steps and devices in each of the embodiments herein are not limited to being implemented in a particular embodiment. In fact, according to the concept of the present invention, relevant partial steps and partial devices in each of the embodiments herein can be combined to conceive new embodiments, and these new embodiments are also included within the scope of the present invention.
[0119] Each operation of the above-described method can be carried out by any suitable means capable of performing the corresponding function. The means can include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors.
[0120] The various illustrated logic blocks, modules, and circuits can be implemented or carried out using a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array signal (FPGA), or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any commercially available processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0121] The steps of the method or algorithm described in connection with the present invention can be directly embedded in hardware, in a software module executed by a processor, or in a combination of the two. The software module can exist in any form of tangible storage medium. Some examples of storage media that can be used include random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, etc. The storage medium can be coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium can be integral with the processor. The software module can be a single instruction or many instructions and can be distributed over several different code segments, different programs, and across multiple storage media.
[0122] The method of this invention includes one or more acts for implementing the described method. The method and / or acts can be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of the acts is specified, the order and / or use of the specific acts can be modified without departing from the scope of the claims.
[0123] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions on a tangible computer-readable medium. The storage medium can be any available tangible medium accessible by a computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other tangible medium that can be used to carry or store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, a disc includes a compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc.
[0124] Accordingly, a computer program product can perform the operations given herein. For example, such a computer program product can be a tangible computer-readable medium having instructions tangibly stored (and / or encoded) thereon that are executable by one or more processors to perform the operations described herein. The computer program product can include packaging materials.
[0125] Software or instructions can also be transmitted over a transmission medium. For example, a transmission medium such as coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, or microwave can be used to transmit software from a website, server, or other remote source.
[0126] In addition, the modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by the user terminal and / or the base station as appropriate. For example, such a device can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage component (such as RAM, ROM, a physical storage medium such as a CD or a floppy disk), so that the user terminal and / or the base station can obtain the various methods when coupled to the device or provided with the storage component. In addition, any other suitable techniques for providing the methods and techniques described herein to the device can be utilized.
[0127] Other examples and implementations are within the scope and spirit of the present invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination thereof. The features implementing the functions can also be physically located at various positions, including being distributed so that portions of the functions are implemented at different physical locations. Also, as used herein, including in the claims, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). In addition, the phrase "exemplary" does not mean that the examples described are preferred or better than other examples.
[0128] Various changes, substitutions, and alterations to the techniques described herein can be made without departing from the teachings of the technology defined by the appended claims. In addition, the scope of the claims of the present invention is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that currently exist or will later be developed and that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0129] The above description of the aspects of the invention is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Thus, the invention is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0130] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the invention to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those of skill in the art will recognize some variations, modifications, alterations, additions, and subcombinations thereof.
Claims
1. A neural network training method, comprising: Inputting a training image and obtaining a feature representation of the training image; Generating a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; Generating a training defect image based on at least the training image and the defect feature representation; Performing defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation to train the neural network and adjust parameters of the neural network.
2. The method according to claim 1, wherein, the feature representation of the training image includes: at least one of a contour of the training image, image segmentation information of the training image, and multi-modal information of the training image.
3. The method according to claim 1, wherein, generating the defect feature representation of the training image according to the feature representation of the training image includes: Obtaining a feature representation of at least a part of the training image, and modifying the logic of the feature representation of the at least a part of the region to obtain a modified feature representation of the at least a part of the region; Fusing the modified feature representation of the at least a part of the region with the feature representation of the training image to generate the defect feature representation.
4. The method according to claim 1, wherein, the defect feature representation of the training image further includes at least one structural defect of the training image.
5. The method according to claim 1, wherein, generating the training defect image based on at least the training image and the defect feature representation of the training image includes: Training an image generation network according to the training image and the feature representation of the training image; Using the trained image generation network to generate a training defect image according to the defect feature representation.
6. The method according to claim 1, wherein, performing defect detection according to the training image, the feature representation of the training image, the training defect image and the defect feature representation to train the neural network and adjust parameters of the neural network includes: Training the neural network by converting the training defect image and the defect feature representation into the training image and the feature representation of the training image respectively, and adjusting parameters of the neural network.
7. A defect detection method, comprising: Inputting an image to be detected, reconstructing the image to be detected by using a neural network, obtaining a reconstructed image, and obtaining a feature representation of the reconstructed image; Performing defect detection on the image to be detected by using the neural network according to the reconstructed image and the feature representation of the reconstructed image; wherein the neural network is trained by the following method: Inputting a training image and obtaining a feature representation of the training image; Generating a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; Generating a training defect image based on at least the training image and the defect feature representation; Perform defect detection based on the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust the parameters of the neural network.
8. A neural network training device, comprising: An input unit configured to input a training image and obtain a feature representation of the training image; A defect generation unit configured to generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; An image generation unit configured to generate a training defect image based at least on the training image and the defect feature representation; A training unit configured to perform defect detection based on the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust the parameters of the neural network.
9. A neural network training device, comprising: A processor; And a memory in which computer program instructions are stored, wherein, when the computer program instructions are run by the processor, the processor is caused to perform the following steps: Input a training image and obtain a feature representation of the training image; Generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; Generate a training defect image based at least on the training image and the defect feature representation; Perform defect detection based on the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust the parameters of the neural network.
10. A computer-readable storage medium having computer program instructions stored thereon, wherein, when the computer program instructions are executed by a processor, the following steps are implemented: Input a training image and obtain a feature representation of the training image; Generate a defect feature representation of the training image according to the feature representation of the training image, the defect feature representation including at least one logical defect of the training image; Generate a training defect image based at least on the training image and the defect feature representation; Perform defect detection based on the training image, the feature representation of the training image, the training defect image, and the defect feature representation to train the neural network and adjust the parameters of the neural network.