Neural network training method and defect detection method, device, and computer program
The neural network training method generates training defect images using perceptual maps and reconstruction networks to improve defect detection accuracy across diverse products, eliminating the need for extensive data and manual labeling.
Patent Information
- Application Number
- JP2024084087
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-05-23
- Filing Date
- 2024-05-23
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2044-05-23
AI Technical Summary
Current neural network training methods for defect detection are inadequate in accurately and robustly detecting random defects across different types of products or objects, requiring large amounts of mark data and separate training for each type, which is inefficient and inaccurate.
A neural network training method that involves generating training defect images using perceptual maps, performing reconstruction to remove defects, and adjusting network parameters based on these images, utilizing a reconstruction network with sub-modules for feature screening and spatial relationship extraction.
This approach enhances defect detection accuracy by generating realistic training images, reducing the need for manual labeling and large data collection, and enabling robust detection across various products or objects.
Smart Images

Figure 0007764915000001 
Figure 0007764915000002 
Figure 0007764915000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of image processing, and more particularly to a neural network training method, and a method, apparatus, and computer-readable storage medium for defect detection using a neural network. [Background technology]
[0002] Defect detection is an important factor in product quality control. Various defects occur constantly during production, and it is usually impossible to cover all the various defect occurrence situations. Therefore, how to detect these defects robustly and automatically is an urgent problem that needs to be solved.
[0003] In current defect detection methods, neural networks that perform defect detection are typically trained using a large amount of mark data, and training and defect detection are performed separately for each type of product or object. However, such neural network training methods and defect detection methods using such neural networks cannot achieve accurate generation and robust detection of random defects in different types of products or objects.
[0004] Therefore, there is a need for improved neural network training methods and defect detection methods that accurately detect image defects. Summary of the Invention
[0005] In order to solve the above technical problem, a neural network training method according to one aspect of the present invention includes the steps of inputting a training image and calculating a perceptual map of the training image; generating a training defect image based on at least the training image and the perceptual map of the training image, where the training defect image includes one or more defects generated based on the perceptual map of the training image; performing reconstruction on the training defect image using a reconstruction network including a first sub-module to obtain a training reconstructed image from which the reconstructed defects have been removed and a perceptual map of the training reconstructed image, where the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; and performing defect detection based on the training image, the training reconstructed image, and the perceptual map of the training reconstructed image to train a neural network and adjust parameters of the neural network.
[0006] A defect detection method according to another aspect of the present invention includes the steps of inputting a detected image, performing reconstruction on the detected image using a reconstruction network including a first submodule, and obtaining a reconstructed image from which defects have been removed after reconstruction and a perceptual map of the reconstructed image, wherein the first submodule is used to screen and process features by screening feature channels and extracting spatial relationships between each channel, and performing defect segmentation based on the detected image, the reconstructed image, and the perceptual map of the reconstructed image to obtain defect detection results.
[0007] According to another aspect of the present invention, a neural network training apparatus includes an input unit that inputs training images and calculates a perceptual map of the training images; a generation unit that generates training defect images based on at least the training images and the perceptual map of the training images, where the training defect images include one or more defects generated based on the perceptual map of the training images; a reconstruction unit that performs reconstruction on the training defect images using a reconstruction network including a first submodule, and obtains training reconstructed images from which the reconstructed defects have been removed and a perceptual map of the training reconstructed images, where the first submodule is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; and a training unit that performs defect detection based on the training images, the training reconstructed images, and the perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network.
[0008] According to another aspect of the present invention, an apparatus for training a neural network includes a processor and a memory in which computer program commands are stored. When the computer program commands are executed by the processor, the apparatus causes the processor to perform the following steps: inputting training images and calculating a perceptual map of the training images; generating training defect images based on at least the training images and the perceptual map of the training images, where the training defect images include one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect images using a reconstruction network including a first submodule to obtain training reconstructed images from which the reconstructed defects have been removed and a perceptual map of the training reconstructed images, where the first submodule is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; and performing defect detection based on the training images, the training reconstructed images, and the perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided having stored thereon computer program commands, which, when executed by a processor, cause the computer to perform the following steps: inputting training images and calculating a perceptual map of the training images; generating training defect images based on at least the training images and the perceptual map of the training images, wherein the training defect images include one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect images using a reconstruction network including a first sub-module to obtain training reconstructed images from which the reconstructed defects have been removed and a perceptual map of the training reconstructed images, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; and performing defect detection based on the training images, the training reconstructed images, and the perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network.
[0010] According to another aspect of the present invention, there is provided an apparatus for detecting defects using a neural network, comprising: an input unit for inputting a detected image, performing reconstruction on the detected image using a reconstruction network including a first submodule, and obtaining a reconstructed image from which defects have been removed after reconstruction and a perceptual map of the reconstructed image, the first submodule being used for screening and processing features by screening feature channels and extracting spatial relationships between each channel; and a detection unit for performing defect segmentation based on the to-be-detected image, the reconstructed image, and a perceptual map of the reconstructed image to obtain a defect detection result.
[0011] According to another aspect of the present invention, an apparatus for detecting defects using a neural network includes a processor and a memory in which computer program commands are stored, and when the computer program commands are executed by the processor, the apparatus causes the processor to perform the following steps: inputting a detected image, performing reconstruction on the detected image using a reconstruction network including a first submodule, and obtaining a reconstructed image from which the post-reconstruction defects have been removed and a perceptual map of the reconstructed image, wherein the first submodule is used to screen and process features by screening feature channels and extracting spatial relationships between each channel; and performing defect segmentation based on the detected image, the reconstructed image, and the perceptual map of the reconstructed image to obtain defect detection results.
[0012] According to another aspect of the present invention, a computer-readable storage medium storing computer program commands causes a processor to execute the following steps when the computer program commands are executed: inputting a detected image, performing reconstruction on the detected image using a reconstruction network including a first sub-module, and obtaining a reconstructed image from which defects have been removed after reconstruction and a perceptual map of the reconstructed image, wherein the first sub-module is used to screen and process features by screening feature channels and extracting spatial relationships between each channel; and performing defect segmentation based on the detected image, the reconstructed image, and the perceptual map of the reconstructed image to obtain defect detection results.
[0013] Based on the above-mentioned neural network training method, apparatus, and computer-readable storage medium, and the method, apparatus, and computer-readable storage medium for performing defect detection using a neural network of the present invention, a first sub-module is introduced for feature screening and processing by screening feature channels and extracting the spatial relationship between each channel, or a second sub-module is further introduced for extracting the difference between the training image, the training reconstructed image, and the perceptual map of the training reconstructed image, thereby improving the neural network training method and improving the accuracy of defect detection, thereby realizing accurate generation and robust detection of disordered defects for different types of products or objects, avoiding the process of collecting large amounts of defect data and manual labeling, and improving the user experience. [Brief explanation of the drawings]
[0014] The above contents, objects, features, and advantages of the present application will become more apparent from the detailed description of the embodiments of the present application in conjunction with the drawings. [Figure 1] 1 is a flowchart of a neural network training method according to an embodiment of the present invention. [Figure 2] 1 is an illustration of a defect mask image according to an embodiment of the present invention. [Figure 3] 1 is an illustration of a defect control panel according to an embodiment of the present invention. [Figure 4] 1 is a diagram illustrating the structure of a reconfiguration network according to an embodiment of the present invention; [Figure 5] 1 is a diagram illustrating the structure of a split network according to an embodiment of the present invention; [Figure 6] 1 is a flowchart of a method for detecting defects using a neural network according to an embodiment of the present invention. [Figure 7] 1 is a block diagram of a neural network training device according to an embodiment of the present invention; [Figure 8] 1 is a block diagram of a neural network training device according to an embodiment of the present invention; [Figure 9] 1 is a block diagram of an apparatus for performing defect detection using a neural network according to an embodiment of the present invention. [Figure 10] 1 is a block diagram of an apparatus for performing defect detection using a neural network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, a neural network training method, apparatus, and computer-readable storage medium, as well as a defect detection method, apparatus, and computer-readable storage medium using a neural network according to embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, like numbers refer to like elements throughout. It is apparent that the embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the present invention.
[0016] When training a neural network for defect detection, a large number of training images containing various defects are generally introduced as samples to adjust the neural network parameters. However, in actual training processes, it is often necessary to train a corresponding model for each type of product or object. As the number of product types increases, the model training time and storage costs also increase significantly. Embodiments of the present invention provide a unified neural network training and defect detection method, which eliminates the need for manual labeling and uses only one model to improve the neural network training method and robustly detect defects in multiple types of products or objects. It should be noted that the method and apparatus provided by embodiments of the present invention are not limited to analyzing product scenes such as cables, tablets, and wooden boards, but can also include other computer vision tasks, such as detecting defects on road surfaces, chip surfaces, or any other object surfaces. To illustrate the application of the present invention through examples, only some features that readers can easily understand, such as color and shape, are listed. However, this does not mean that the present invention extracts and relies solely on these features. Advantageously, embodiments of the present invention can be implemented based on convolutional neural networks (CNNs) that can extract more complex features at multiple semantic levels, rather than extracting only limited low-level features.
[0017] 1 is a flowchart of a neural network training method 100 according to an embodiment of the present invention. The neural network training method according to an embodiment of the present invention will be described below with reference to FIG.
[0018] In step S101, a training image is input and a perceptual map of the training image is calculated.
[0019] In this step, the training images input to train the neural network may be images of defect-free products or objects, and the training images may be images of the same product or object taken at different angles, under different lighting conditions, and with different backgrounds, in order to enrich the training samples as much as possible, thereby further improving the defect detection accuracy of the trained neural network.
[0020] After the training images are input, a noticeable map (also referred to as a noticable map) of the training images may be calculated based on the input training images, and the calculated noticeable map of the training images may be a multi-scale perceptual map. Thus, the multi-scale perceptual map of the training images can represent the intensity changes of each image feature in the training images, i.e., the visual salience differences of each image feature in the training images. The multi-scale perceptual map of the training images can make the generated training defect images more natural and make the subsequent defect detection and localization processes more accurate.
[0021] In one example, a multi-scale just-noticeable image of the training image can be calculated using input training images, e.g., just-noticeable difference (JND) The minimum perceptual map of the training images can be calculated by a Just-in-Time Distortion (JND) model. Therefore, the JND value is an important visual saliency index that describes the degree of human perception of image quality. The JND value reveals the perceptual threshold of image intensity changes that the human visual system can notice. Any perceptible distortion level must be greater than the JND value. In an embodiment of the present invention, a training defect image sample set with diverse and multi-scale samples can be generated by the JND model and used in the subsequent defect detection and location process.
[0022] Preferably, the process of calculating the multi-scale minimum perceptual map of the training image using the just-noticeable difference model can include the following: First, calculate the minimum perceptual map of the training image based on the just-noticeable difference model, and then process the minimum perceptual map of the training image using multiple coefficients (e.g., multiply the minimum perceptual map of the training image by multiple coefficients) to obtain the multi-scale minimum perceptual map of the training image. For example, three coefficients (e.g., 0.33, 0.66, and 1) can be selected and multiplied by the minimum perceptual map of the training image, respectively, to obtain the three-scale minimum perceptual maps of the training image corresponding to three different scales. In another example, four coefficients (e.g., 0.25, 1, 0.75, and 1) can be selected and multiplied by the minimum perceptual map of the training image, respectively, to obtain the four-scale minimum perceptual maps of the training image corresponding to four different scales. The above-mentioned generation methods for the multi-scale minimum perceptual map of the training image are merely examples. In embodiments of the present invention, multi-scale minimum perceptual maps of training patterns at various scales can be generated in various ways according to the needs of different application scenarios. As described above, the multi-scale minimum perceptual map of the training image can represent the range and intensity variations of each image feature in the training image, and since it is perceptible to humans, the distribution and variations of grayscale values therein represent the differences in visual saliency of each image feature in the training image.
[0023] In step S102, a training defect image is generated from at least the training image and a perceptual map of the training image, where the training defect image includes one or more defects generated based on the perceptual map of the training image.
[0024] In this step, the process of generating training defect images can include first generating a defect mask image, and then generating one or more defect texture images based on the training image. Then, based on a defect control panel and the defect mask image, and using a multi-scale perceptual map of the training image, the training image can be fused with the one or more defect texture images to generate training defect images, where the defect control panel is used to control the proportion of the defect texture images in the training image.
[0025] Preferably, the defect mask image can be generated by selecting one or more defect regions of different sizes and shapes from the training image. Also, preferably, the defect texture image can be generated randomly from any image or by processing the training image. In one example, to enhance the realistic effect of defect simulation, the defect texture image can be generated by enhancing the training image. For example, various random processes, including inversion, scaling, rotation, area cropping, color conversion, brightness change, exposure change, etc., can be performed on the training image to obtain defect texture images with various textures. FIG. 2 is an example of a defect mask image according to one embodiment of the present invention. Specifically, FIG. 2 shows a multi-scale defect mask image in which defect regions of different sizes are marked and generated by selecting one or more defect regions of different sizes and shapes from the training image.
[0026] After obtaining the defect mask image and the defect texture image, the training image and one or more defect texture images are preferably fused using the multi-scale perceptual map of the training image as a weight image according to the defect area defined by combining the defect control panel and the defect mask image to generate a training defect image with various simulated defects. According to the principle of the perceptual map, the perceptual map of the training image can reveal areas with high color contrast and saturation in the training image, and areas with a larger grayscale in the perceptual map can occupy more defect components in the training defect image, i.e., can give more weight to the defect texture image. Conversely, areas with a smaller grayscale in the perceptual map of the training image can give more weight relative to the input training image. Based on this, applying the multi-scale perceptual map of the training image can obtain a fused training defect image at various scales, making the obtained training defect image more realistic and further used in the subsequent defect detection process.
[0027] FIG. 3 illustrates various examples of defect control panels according to an embodiment of the present invention. The defect control panel illustrated in FIG. 3 is used to limit the area occupied by a defect texture image when the defect texture image is fused with a training image. When used, a defect control panel to be used may be randomly selected from multiple defect control panels, or may be selected based on a preset order or criteria, and is not limited thereto. Preferably, the defect texture image occupies an area of n% width and m% height at a certain position (e.g., above, below, left, or right) in the training image, where m and n can be randomly selected from a preset interval, for example, the preset interval may be [50, 80]. Furthermore, the defect control panel may be, for example, an area covering the entire training image, and is not limited thereto. Of course, the above-described area setting and size selection of the defect control panel are merely examples. In actual applications, any control method for the defect control panel can be selected based on the application scenario, and the description thereof will be omitted here. By selecting and using the defect control panel, the proportion of the defect texture image that occupies the entire area of the training image can be controlled during the process of generating the training defect image, making the training defect image generation results more diverse and random, and obtaining more accurate neural network training and defect detection results.
[0028] Preferably, once the multi-scale perceptual map of the training image is generated, a multi-scale defect mask image can be generated accordingly, thereby fusing the training image and the defect texture image to finally obtain a training defect image. Specifically, the multi-scale defect mask image can be a defect mask image marked with defect areas of different sizes corresponding to each scale in the multi-scale perceptual map of the training image. For example, referring to the aforementioned process of generating a multi-scale minimally noticeable image perceptual map of the training image, the initially generated defect mask image can also be processed accordingly by multiple coefficients (e.g., multiplying the defect mask image by multiple coefficients respectively) to obtain corresponding multi-scale defect mask images, thereby defining the ranges of different defect areas at different scales. In the subsequent defect generation process, in the process of fusing the training image and the defect texture image using these multi-scale defect mask images based on the defect control panel, corresponding image processing can be performed on the defect areas marked in the multi-scale defect mask image, such as Gaussian blurring, in the fusion process to obtain a more realistic and natural-looking training defect image.
[0029] The specific methods for generating the training defect images described above are merely examples. In actual applications, various methods can be used to generate training defect images based on the corresponding scene, and are not limited here.
[0030] In step S103, a reconstruction network including a first submodule performs reconstruction on the training defect image to obtain a training reconstructed image with the defects removed after reconstruction and a perceptual map of the training reconstructed image, where the first submodule is used to screen and process features by screening feature channels and extracting the spatial relationship between each channel.
[0031] FIG. 4 illustrates the structure of a reconstruction network according to an embodiment of the present invention. In FIG. 4, training defect images are passed through a reconstruction network including an encoder and a decoder to generate training reconstructed images with defects removed and perceptual maps (or minimum perceptual maps) of the training reconstructed images. As illustrated, each layer of the encoder and decoder of the reconstruction network may include a base sub-module and a first sub-module, respectively, and the base sub-module and the first sub-module are connected to each other at each layer. Preferably, the base sub-module may include two consecutive 3×3 convolutional layers, a normalization layer, and a reLU active layer to extract features at each layer. The first sub-module may include a channel attention sub-block and a spatial attention sub-block. Preferably, the channel attention sub-block may extract relationships between different feature channels using an average pooling layer and two consecutive 1×1 convolutional layers, and may add the relationships between different feature channels back to the training defect images to screen for relatively important feature channels. The spatial attention sub-block can further extract the spatial relationships of each feature channel, and the extracted spatial relationships can be cascaded together after passing through an average pooling layer and a max pooling layer. Then, important feature regions are screened using a standard convolutional layer and a dilated convolutional network (e.g., a 3x3 convolutional layer with dilation), where the presence of the dilated convolutional network can expand the feature receptive field while maintaining spatial resolution.
[0032] When an input training defect image enters the reconstruction network, the input features can first be extracted using an encoder. Preferably, features of different levels can be extracted using convolution, normalization, and pooling processes with different parameters of the encoder in the reconstruction network. First, a convolution kernel can be used to perform convolution on the input training defect image to obtain a convolution mapping. Next, a linear correction unit and a batch normalization method can be used to normalize the convolution mapping to obtain a normalized convolution mapping. Then, a maximum or average pooling process can be applied to the normalized convolution mapping. To obtain rich multi-scale features, the encoder can adjust related parameters and repeat the above process multiple times to extract corresponding multi-scale feature maps through multiple downsampling processes. After the encoder processing of the reconstruction network, the decoder can restore the resolution of the multi-scale feature map through convolution, normalization, and upsampling processes. The reconstruction network combines features from the encoder and decoder with the same resolution and uses multiple sets of convolution processes. During the training phase of the neural network, the reconstruction network can convert the input training defect image into a corresponding training reconstruction image and a perceptual map of the corresponding training reconstruction image by learning and updating each parameter and weight of the neural network under the supervision of the input training image and the perceptual map of the training image.
[0033] In step S104, the neural network is trained and parameters of the neural network are adjusted by performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images.
[0034] Preferably, performing defect detection based on the training images, the training reconstructed images, and the perceptual map of the training reconstructed images may include: using a division network including a second sub-module to respectively extract differences between the training images and the training reconstructed images, between the training images and the perceptual map of the training reconstructed images, and between the training images and the training reconstructed images and the perceptual map of the training reconstructed images, to perform defect detection.
[0035] FIG. 5 shows a structural schematic diagram of a segmentation network according to an embodiment of the present invention. Preferably, the segmentation network of FIG. 5 can include a second sub-module and an encoder and a decoder connected to the second sub-module. In FIG. 5, the second sub-module is used to input training images, training reconstructed images, and perceptual maps of the training reconstructed images to extract difference features between the multi-layer training images, the training reconstructed images, and the perceptual maps of the training reconstructed images. The second sub-module first extracts the differences between the training images and the training reconstructed images, between the training images and the perceptual maps of the training reconstructed images, and between the training images and the training reconstructed images and the perceptual maps of the training reconstructed images, respectively. Then, the differences are serially combined to form a difference feature description whose number of feature channels is equal to the sum of the feature channels of these images. This is then fed into a network including an encoder and a decoder to perform defect segmentation, obtain defect detection results, and locate the defects. In FIG. 5, the specific network configuration of the encoder and decoder is the same as that of FIG. 4. A first sub-module may be installed at each layer of the encoder and decoder, and will not be described again here.
[0036] The defect segmentation and detection of the segmentation network of FIG. 5 can locate defects in training images and provide a decision or probability indication of whether a defect is present.
[0037] Preferably, when locating defects, the segmentation network can segment and classify pixel values in the resulting image for displaying defects, thereby more intuitively indicating the location of the defects. For example, when outputting an image for displaying defects, the pixels in the grayscale image can be divided into three parts according to their grayscale values, and these can be represented by normalized results of, for example, 0, 0.5, and 1, respectively, to more clearly indicate the location of the defects. Of course, the above process of segmenting, classifying, and displaying defects is merely an example. In actual applications, various pixel value classification and normalization operations can be performed on the image for displaying defects according to needs, and are not limited thereto. After locating and displaying the defects, the segmentation network can output a defect detection result, for example, a judgment on whether the training image is a defect image or whether the training image contains a defect.
[0038] The neural network training method according to an embodiment of the present invention improves the training method of the neural network and improves the accuracy of defect detection by introducing a first sub-module for feature screening and processing through screening for feature channels and extracting the spatial relationship between each channel, or further introducing a second sub-module for extracting the difference between the training image, the training reconstructed image, and the perceptual map of the training reconstructed image, thereby realizing accurate generation and robust detection of disordered defects for different types of products or objects, avoiding the process of collecting large amounts of defect data and manually labeling them, and improving the user experience.
[0039] Furthermore, according to another embodiment of the present invention, neural network training is not limited to the specific operation process of the neural network training method described above, and other needs in various application scenarios can be met through the neural network architecture and process. For example, the neural network architecture can be used to create diversified samples, such as human faces, animated images, etc. Alternatively, it can be applied to other scenes that require the introduction of other image defects or noise, thereby introducing various relatively realistic simulated defects or noise into the image. The method according to the embodiment of the present invention can combine training images and introduce various noises to obtain diversified samples for application in various application scenarios.
[0040] 6 is a flowchart of a method 600 for detecting defects using a neural network according to an embodiment of the present invention. In this embodiment, defect detection can be performed using a neural network that has been trained according to the process shown in FIG. 1. The defect detection method using a neural network according to an embodiment of the present invention will now be described with reference to FIG. 6.
[0041] In step S601, a detected image is input, and a reconstruction network including a first sub-module performs reconstruction on the detected image, and obtains a reconstructed image with defects removed after reconstruction and a perceptual map of the reconstructed image, where the first sub-module is used to screen and process features by screening feature channels and extracting spatial relationships between each channel.
[0042] In this step, the detected image can be input to a reconstruction network including a first submodule in the neural network trained by the process shown in FIG. 1 for reconstruction. Preferably, the detected image can be an image without defects or an image with one or more defects. In the process of detecting the detected image, a perceptual map of the reconstructed image can be obtained at the same time as obtaining a reconstructed image with the defects removed, or a multi-scale perceptual map of the reconstructed image can be further obtained and used in the subsequent defect detection and location process. The specific structure of the reconstruction network is shown in FIG. 4, and will not be described again here.
[0043] When an input detected image enters the reconstruction network, the input features can first be extracted using an encoder. Preferably, features of different layers can be extracted using convolution, normalization, and pooling processes of different parameters of the encoder in the reconstruction network. First, a convolution kernel can be used to perform convolution on the input detected image to obtain a convolution mapping. Next, a linear correction unit and a batch normalization method can be used to normalize the convolution mapping to obtain a normalized convolution mapping. Then, a maximum or average pooling process can be applied to the normalized convolution mapping. To obtain rich multi-scale features, the encoder can adjust related parameters and repeat the above process multiple times to extract corresponding multi-scale feature maps through multiple downsampling processes. After the encoder processing of the reconstruction network, the decoder can restore the resolution of the multi-scale feature maps through convolution, normalization, and upsampling processes. The reconstruction network combines features from the encoder and decoder with the same resolution and uses multiple sets of convolution processes to convert the input detected image into a corresponding reconstructed image with defects removed and a perceptual map of the corresponding reconstructed image.
[0044] In step S602, defect segmentation is performed based on the to-be-detected image, the reconstructed image, and a perceptual map of the reconstructed image, to obtain a defect detection result.
[0045] Preferably, performing defect segmentation based on at least the detected image, the reconstructed image, and the perceptual map of the reconstructed image may include: using a segmentation network including a second sub-module to respectively extract differences between the detected image and the reconstructed image, between the detected image and the perceptual map of the reconstructed image, and between the detected image and the reconstructed image and the perceptual map of the reconstructed image to perform defect detection. The specific structure of the segmentation network is shown in FIG. 5, and will not be described again here.
[0046] The defect segmentation and detection of the segmentation network of FIG. 5 can locate defects in the detected image and provide a decision or probability indication of whether a defect is present.
[0047] Preferably, when locating a defect, the segmentation network can segment and classify pixel values in the resulting image for displaying the defect, thereby more intuitively indicating the location of the defect. For example, when outputting an image for displaying a defect, the pixels in the grayscale image can be divided into three parts according to their grayscale values, and these can be represented by normalized results of, for example, 0, 0.5, and 1, respectively, to more clearly indicate the location of the defect. Of course, the above process of segmenting, classifying, and displaying defects is merely an example. In actual applications, various pixel value classification and normalization operations can be performed on the image for displaying defects according to needs, and are not limited thereto. After locating and displaying the defect, the segmentation network can output a defect detection result, for example, a judgment on whether the detected image is a defective image or whether the detected image contains a defect.
[0048] The above-mentioned neural network defect detection method according to an embodiment of the present invention improves the accuracy of defect detection by introducing a first sub-module for feature screening and processing through screening on feature channels and extracting the spatial relationship between each channel, or further introducing a second sub-module for extracting the difference between the training image, the training reconstructed image, and the perceptual map of the training reconstructed image, thereby realizing accurate generation and robust detection of disordered defects for different types of products or objects, avoiding the process of collecting large amounts of defect data and manually labeling, and improving user experience.
[0049] Next, a neural network training apparatus according to an embodiment of the present invention will be described with reference to FIG. 7. FIG. 7 is a block diagram of a neural network training apparatus 700 according to an embodiment of the present invention. As shown in FIG. 7, the neural network training apparatus 700 includes an input unit 710, a generation unit 720, a reconstruction unit 730, and a training unit 740. The neural network training apparatus 700 may include other components, but these components are not relevant to the present embodiment and will not be illustrated or described here. Furthermore, the specific details of the following operations performed by the neural network training apparatus 700 according to the present embodiment are the same as those described in FIG. 1 above, and therefore, to avoid redundancy, the same details will not be described again.
[0050] The input unit 710 of the neural network training apparatus 700 of FIG. 7 inputs training images and calculates perceptual maps of the training images.
[0051] The input unit 710 inputs training images and calculates the perceptible images of the training images.
[0052] The training images input by the input unit 710 for training the neural network may be images of products or objects without defects, and the training images may be images of the same product or object taken at different angles, under different lighting conditions, and with different backgrounds, in order to make the training samples as rich as possible, thereby further improving the defect detection accuracy of the trained neural network.
[0053] After the training images are input, a noticeable map of the training images may be calculated based on the input training images, and the calculated noticeable map of the training images may be a multi-scale perceptual map. The multi-scale perceptual map of the training images may represent the intensity variation of each image feature in the training images, i.e., the visual saliency difference of each image feature in the training images. The multi-scale perceptual map of the training images may make the generated training defect images more natural and make the subsequent defect detection and localization process more accurate.
[0054] In one example, given a training image, a multi-scale just-noticeable image of the training image can be calculated, e.g., just-noticeable difference (JND). The minimum perceptual map of the training images can be calculated by a Just-in-Time Distortion (JND) model. Therefore, the JND value is an important visual saliency index that describes the degree of human perception of image quality. The JND value reveals the perceptual threshold of image intensity changes that the human visual system can notice. Any perceptible distortion level must be greater than the JND value. In an embodiment of the present invention, a training defect image sample set with diverse and multi-scale samples can be generated by the JND model and used in the subsequent defect detection and location process.
[0055] Preferably, the process of calculating the multi-scale minimum perceptual map of the training image using the just-noticeable difference model can include the following: First, calculate the minimum perceptual map of the training image based on the just-noticeable difference model, and then process the minimum perceptual map of the training image using multiple coefficients (e.g., multiply the minimum perceptual map of the training image by multiple coefficients) to obtain the multi-scale minimum perceptual map of the training image. For example, three coefficients (e.g., 0.33, 0.66, and 1) can be selected and multiplied by the minimum perceptual map of the training image, respectively, to obtain the three-scale minimum perceptual maps of the training image corresponding to three different scales. In another example, four coefficients (e.g., 0.25, 1, 0.75, and 1) can be selected and multiplied by the minimum perceptual map of the training image, respectively, to obtain the four-scale minimum perceptual maps of the training image corresponding to four different scales. The above-mentioned generation methods for the multi-scale minimum perceptual map of the training image are merely examples. In embodiments of the present invention, multi-scale minimum perceptual maps of training patterns at various scales can be generated in various ways according to the needs of different application scenarios. As described above, the multi-scale minimum perceptual map of the training image can represent the range and intensity variations of each image feature in the training image, and since it is perceptible to humans, the distribution and variations of grayscale values therein represent the differences in visual saliency of each image feature in the training image.
[0056] A generation unit 720 generates training defect images from at least the training images and the perceptual map of the training images, where the training defect images include one or more defects generated based on the perceptual map of the training images.
[0057] The process of generating training defect images by the generation unit 720 may include first generating a defect mask image, and then generating one or more defect texture images based on the training image. Then, based on a defect control panel and the defect mask image, and using a multi-scale perceptual map of the training image, the training image can be fused with the one or more defect texture images to generate training defect images, where the defect control panel is used to control the proportion of the defect texture image in the training image.
[0058] Preferably, the defect mask image can be generated by selecting one or more defect regions of different sizes and shapes from the training image. Also, preferably, the defect texture image can be generated randomly from any image or by processing the training image. In one example, to enhance the realistic effect of defect simulation, the defect texture image can be generated by enhancing the training image. For example, various random processes, including inversion, scaling, rotation, area cropping, color conversion, brightness change, exposure change, etc., can be performed on the training image to obtain defect texture images with various textures. FIG. 2 is an example of a defect mask image according to one embodiment of the present invention. Specifically, FIG. 2 shows a multi-scale defect mask image in which defect regions of different sizes are marked and generated by selecting one or more defect regions of different sizes and shapes from the training image.
[0059] After obtaining the defect mask image and the defect texture image, the training image and one or more defect texture images are preferably fused using the multi-scale perceptual map of the training image as a weight image according to the defect area defined by combining the defect control panel and the defect mask image to generate a training defect image with various simulated defects. According to the principle of the perceptual map, the perceptual map of the training image can reveal areas with high color contrast and saturation in the training image, and areas with a larger grayscale in the perceptual map can occupy more defect components in the training defect image, i.e., can give more weight to the defect texture image. Conversely, areas with a smaller grayscale in the perceptual map of the training image can give more weight relative to the input training image. Based on this, applying the multi-scale perceptual map of the training image can obtain a fused training defect image at various scales, making the obtained training defect image more realistic and further used in the subsequent defect detection process.
[0060] FIG. 3 illustrates various examples of defect control panels according to an embodiment of the present invention. The defect control panel illustrated in FIG. 3 is used to limit the area occupied by a defect texture image when the defect texture image is fused with a training image. When used, a defect control panel to be used may be randomly selected from multiple defect control panels, or may be selected based on a preset order or criteria, and is not limited thereto. Preferably, the defect texture image occupies an area of n% width and m% height at a certain position (e.g., above, below, left, or right) in the training image, where m and n can be randomly selected from a preset interval, for example, the preset interval may be [50, 80]. Furthermore, the defect control panel may be, for example, an area covering the entire training image, and is not limited thereto. Of course, the above-described area setting and size selection of the defect control panel are merely examples. In actual applications, any control method for the defect control panel can be selected based on the application scenario, and the description thereof will be omitted here. By selecting and using the defect control panel, the proportion of the defect texture image that occupies the entire area of the training image can be controlled during the process of generating the training defect image, making the training defect image generation results more diverse and random, and obtaining more accurate neural network training and defect detection results.
[0061] Preferably, once the multi-scale perceptual map of the training image is generated, a multi-scale defect mask image can be generated accordingly, thereby fusing the training image and the defect texture image to finally obtain a training defect image. Specifically, the multi-scale defect mask image can be a defect mask image marked with defect areas of different sizes corresponding to each scale in the multi-scale perceptual map of the training image. For example, referring to the aforementioned process of generating a multi-scale minimum perceptual map of the training image, the initially generated defect mask image can also be processed accordingly by multiple coefficients (e.g., multiplying the defect mask image by multiple coefficients respectively) to obtain corresponding multi-scale defect mask images, thereby defining the ranges of different defect areas at different scales. In the subsequent defect generation process, in the process of fusing the training image and the defect texture image using these multi-scale defect mask images based on the defect control panel, corresponding image processing can be performed on the defect areas marked in the multi-scale defect mask image, such as Gaussian blurring, in the fusion process to obtain a more realistic and natural-looking training defect image.
[0062] The specific methods for generating the training defect images described above are merely examples. In actual applications, various methods can be used to generate training defect images based on the corresponding scene, and are not limited here.
[0063] The reconstruction unit 730 performs reconstruction on the training defect image through a reconstruction network including a first sub-module, and obtains a training reconstructed image with defects removed and a perceptual map of the training reconstructed image after reconstruction, where the first sub-module is used to screen and process features by screening feature channels and extracting spatial relationships between each channel.
[0064] FIG. 4 illustrates the structure of a reconstruction network according to an embodiment of the present invention. In FIG. 4, training defect images are passed through a reconstruction network including an encoder and a decoder to generate training reconstructed images with defects removed and perceptual maps (or minimum perceptual maps) of the training reconstructed images. As illustrated, each layer of the encoder and decoder of the reconstruction network may include a base sub-module and a first sub-module, respectively, and the base sub-module and the first sub-module are connected to each other at each layer. Preferably, the base sub-module may include two consecutive 3×3 convolutional layers, a normalization layer, and a reLU active layer to extract features at each layer. The first sub-module may include a channel attention sub-block and a spatial attention sub-block. Preferably, the channel attention sub-block may extract relationships between different feature channels using an average pooling layer and two consecutive 1×1 convolutional layers, and may add the relationships between different feature channels back to the training defect images to screen for relatively important feature channels. The spatial attention sub-block can further extract the spatial relationships of each feature channel, and the extracted spatial relationships can be cascaded together after passing through an average pooling layer and a max pooling layer. Then, a standard convolutional layer and a dilated convolutional network (e.g., a 3x3 convolutional layer with dilation) are used to screen important feature regions, where the presence of the dilated convolutional network can expand the feature receptive field while maintaining spatial resolution.
[0065] When an input training defect image enters the reconstruction network, the input features can first be extracted using an encoder. Preferably, features of different levels can be extracted using convolution, normalization, and pooling processes with different parameters of the encoder in the reconstruction network. First, a convolution kernel can be used to perform convolution on the input training defect image to obtain a convolution mapping. Next, a linear correction unit and a batch normalization method can be used to normalize the convolution mapping to obtain a normalized convolution mapping. Then, a maximum or average pooling process can be applied to the normalized convolution mapping. To obtain rich multi-scale features, the encoder can adjust related parameters and repeat the above process multiple times to extract corresponding multi-scale feature maps through multiple downsampling processes. After the encoder processing of the reconstruction network, the decoder can restore the resolution of the multi-scale feature map through convolution, normalization, and upsampling processes. The reconstruction network combines features from the encoder and decoder with the same resolution and uses multiple sets of convolution processes. During the training phase of the neural network, the reconstruction network can convert the input training defect image into a corresponding training reconstruction image and a perceptual map of the corresponding training reconstruction image by learning and updating each parameter and weight of the neural network under the supervision of the input training image and the perceptual map of the training image.
[0066] A training unit 740 trains the neural network and adjusts parameters of the neural network by performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images.
[0067] Preferably, the training unit 740 performs defect detection by extracting differences between the training image and the training reconstructed image, between the training image and the perceptual map of the training reconstructed image, and between the training image, the training reconstructed image, and the perceptual map of the training reconstructed image, respectively, using a division network including a second sub-module.
[0068] FIG. 5 shows a structural schematic diagram of a segmentation network according to an embodiment of the present invention. Preferably, the segmentation network of FIG. 5 can include a second sub-module and an encoder and a decoder connected to the second sub-module. In FIG. 5, the second sub-module is used to input training images, training reconstructed images, and perceptual maps of the training reconstructed images to extract difference features between the multi-layer training images, the training reconstructed images, and the perceptual maps of the training reconstructed images. The second sub-module first extracts the differences between the training images and the training reconstructed images, between the training images and the perceptual maps of the training reconstructed images, and between the training images and the training reconstructed images and the perceptual maps of the training reconstructed images, respectively. Then, the differences are serially combined to form a difference feature description whose number of feature channels is equal to the sum of the feature channels of these images. This is then fed into a network including an encoder and a decoder to perform defect segmentation, obtain defect detection results, and locate the defects. In FIG. 5, the specific network configuration of the encoder and decoder is the same as that of FIG. 4. A first sub-module may be installed at each layer of the encoder and decoder, and will not be described again here.
[0069] The defect segmentation and detection of the segmentation network of FIG. 5 can locate defects in training images and provide a decision or probability indication of whether a defect is present.
[0070] Preferably, when locating defects, the segmentation network can segment and classify pixel values in the resulting image for displaying defects, thereby more intuitively indicating the location of the defects. For example, when outputting an image for displaying defects, the pixels in the grayscale image can be divided into three parts according to their grayscale values, and these can be represented by normalized results of, for example, 0, 0.5, and 1, respectively, to more clearly indicate the location of the defects. Of course, the above process of segmenting, classifying, and displaying defects is merely an example. In actual applications, various pixel value classification and normalization operations can be performed on the image for displaying defects according to needs, and are not limited thereto. After locating and displaying the defects, the segmentation network can output a defect detection result, for example, a judgment on whether the training image is a defect image or whether the training image contains a defect.
[0071] The neural network training device of the embodiment of the present invention improves the training method of the neural network and improves the accuracy of defect detection by introducing a first sub-module for feature screening and processing through screening for feature channels and extracting the spatial relationship between each channel, or further introducing a second sub-module for extracting the difference between the training image, the training reconstructed image, and the perceptual map of the training reconstructed image, thereby realizing accurate generation and robust detection of disordered defects for different types of products or objects, avoiding the process of collecting large amounts of defect data and manually labeling, and improving the user experience.
[0072] Next, a neural network training apparatus according to an embodiment of the present invention will be described with reference to Fig. 8. Fig. 8 is a block diagram of a neural network training apparatus 800 according to an embodiment of the present invention. As shown in Fig. 8, the apparatus 800 can be a computer or a server.
[0073] 8, neural network training apparatus 800 includes one or more processors 810 and memory 820. In addition, neural network training apparatus 800 may include input devices, output devices (not shown), etc., which may be interconnected via a bus system and / or other types of connection mechanisms. Note that the components and configuration of neural network training apparatus 800 shown in FIG. 8 are exemplary only and are not limiting; neural network training apparatus 800 may have other components and configurations as needed.
[0074] The processor 810 may be a central processing unit (CPU) or other type of processing unit having data processing and / or command execution capabilities, and can perform desired functions using computer program commands stored in the memory 820, including: inputting a training image and calculating a perceptual map of the training image; generating a training defect image from at least the training image and the perceptual map of the training image, where the training defect image includes one or more defects generated based on the perceptual map of the training image; performing reconstruction on the training defect image through a reconstruction network including a first sub-module, and obtaining a defect-removed training reconstructed image and a perceptual map of the training reconstructed image after reconstruction, where the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; and training the neural network and adjusting parameters of the neural network by performing defect detection based on the training image, the training reconstructed image, and the perceptual map of the training reconstructed image.
[0075] The memory 820 may include one or more computer program products, which may include various types of computer-readable storage media, such as volatile and / or non-volatile memory. One or more computer program commands may be stored in the computer-readable storage media, and the processor 810 may execute the program commands to perform the functions of the neural network training apparatus according to the embodiment of the present invention and / or other desired functions and / or the neural network training method according to the embodiment of the present invention. Various application programs and various data may also be stored in the computer-readable storage media.
[0076] The following describes a computer-readable storage medium storing a computer program according to an embodiment of the present invention, which can be implemented by a processor in the following steps: input a training image and calculate a perceptual map for the training image; generate a training defect image from at least the training image and the perceptual map for the training image, where the training defect image includes one or more defects generated based on the perceptual map for the training image; perform reconstruction on the training defect image through a reconstruction network including a first sub-module, and obtain a training reconstructed image and a perceptual map for the training reconstructed image from which defects have been removed after reconstruction, where the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; and train the neural network and adjust parameters of the neural network by detecting defects based on the training image, the training reconstructed image, and the perceptual map for the training reconstructed image.
[0077] Hereinafter, a defect detection apparatus using a neural network according to an embodiment of the present invention will be described with reference to FIG. 9. FIG. 9 shows a block diagram of an apparatus 900 for performing defect detection using a neural network according to an embodiment of the present invention. As shown in FIG. 9, the defect detection apparatus 900 using a neural network includes an input unit 910 and a detection unit 920. The apparatus 900 may include other components, but these components are not relevant to the present embodiment and will not be illustrated or described here. In addition, the specific details of the following operations performed by the apparatus 900 according to the embodiment of the present invention are the same as those described in FIG. 6 above, and therefore, to avoid redundancy, the same details will not be described again.
[0078] The input unit 910 of the apparatus 900 for defect detection using a neural network in FIG. 9 inputs a detected image, and a reconstruction network including a first submodule performs reconstruction on the detected image to obtain a reconstructed image with defects removed after reconstruction and a perceptual map of the reconstructed image, where the first submodule is used to screen and process features by screening feature channels and extracting spatial relationships between each channel.
[0079] The input unit 910 can input the detected image to a reconstruction network including a first submodule in the neural network trained by the process shown in FIG. 1 for reconstruction. Preferably, the detected image may be an image without defects or an image with one or more defects. In the process of detecting the detected image, a perceptual map of the reconstructed image can be obtained at the same time as obtaining a reconstructed image with the defects removed, or a multi-scale perceptual map of the reconstructed image can be further obtained and used in the subsequent defect detection and location process. The specific structure of the reconstruction network is shown in FIG. 4 and will not be described again here.
[0080] When an input detected image enters the reconstruction network, the input features can first be extracted using an encoder. Preferably, features of different layers can be extracted using convolution, normalization, and pooling processes of different parameters of the encoder in the reconstruction network. First, a convolution kernel can be used to perform convolution on the input detected image to obtain a convolution mapping. Next, a linear correction unit and a batch normalization method can be used to normalize the convolution mapping to obtain a normalized convolution mapping. Then, a maximum or average pooling process can be applied to the normalized convolution mapping. To obtain rich multi-scale features, the encoder can adjust related parameters and repeat the above process multiple times to extract corresponding multi-scale feature maps through multiple downsampling processes. After the encoder processing of the reconstruction network, the decoder can restore the resolution of the multi-scale feature maps through convolution, normalization, and upsampling processes. The reconstruction network combines features from the encoder and decoder with the same resolution and uses multiple sets of convolution processes to convert the input detected image into a corresponding reconstructed image with defects removed and a perceptual map of the corresponding reconstructed image.
[0081] The detection unit 920 performs defect segmentation based on the detected image, the reconstructed image, and the perceptual map of the reconstructed image to obtain a defect detection result.
[0082] Preferably, the detection unit 920 uses a segmentation network including a second sub-module to respectively extract differences between the detected image and the reconstructed image, between the detected image and the perceptual map of the reconstructed image, and between the detected image, the reconstructed image, and the perceptual map of the reconstructed image to perform defect detection. The specific structure of the segmentation network is shown in Figure 5, and will not be described again here.
[0083] The defect segmentation and detection of the segmentation network of FIG. 5 can locate defects in the detected image and provide a decision or probability indication of whether a defect is present.
[0084] Preferably, when locating a defect, the segmentation network can segment and classify pixel values in the resulting image for displaying the defect, thereby more intuitively indicating the location of the defect. For example, when outputting an image for displaying a defect, the pixels in the grayscale image can be divided into three parts according to their grayscale values, and these can be represented by normalized results of, for example, 0, 0.5, and 1, respectively, to more clearly indicate the location of the defect. Of course, the above process of segmenting, classifying, and displaying defects is merely an example. In actual applications, various pixel value classification and normalization operations can be performed on the image for displaying defects according to needs, and are not limited thereto. After locating and displaying the defect, the segmentation network can output a defect detection result, for example, a judgment on whether the detected image is a defective image or whether the detected image contains a defect.
[0085] The neural network defect detection apparatus according to an embodiment of the present invention can improve the accuracy of defect detection by introducing a first sub-module for feature screening and processing through screening of feature channels and extracting the spatial relationship between each channel, or by further introducing a second sub-module for extracting the difference between the training image, the training reconstructed image, and the perceptual map of the training reconstructed image, thereby realizing accurate generation and robust detection of disordered defects for different types of products or objects, avoiding the process of collecting large amounts of defect data and manually labeling, and improving user experience.
[0086] A defect detection apparatus using a neural network according to an embodiment of the present invention will be described below with reference to Fig. 10. Fig. 10 shows a block diagram of an apparatus 900 for performing defect detection using a neural network according to an embodiment of the present invention. As shown in Fig. 10, the apparatus 1000 can be a computer or a server.
[0087] As shown in Figure 10, device 1000 includes one or more processors 1010 and memory 1020. In addition, device 1000 may include input devices, output devices (not shown), etc., which may be interconnected via a bus system and / or other type of connection mechanism. Note that the components and structure of device 1000 shown in Figure 9 are exemplary only and are not limiting. Device 1000 may have these components and structure as desired.
[0088] The processor 1010 may be a central processing unit (CPU) or other types of processing unit having data processing capabilities and / or command execution capabilities, and can perform desired functions using computer program commands stored in the memory 1020, including the following: inputting a detected image, performing reconstruction on the detected image through a reconstruction network including a first sub-module, obtaining a reconstructed image with defects removed after reconstruction and a perceptual map of the reconstructed image, where the first sub-module is used for screening and processing features by screening feature channels and extracting spatial relationships between each channel; performing defect segmentation based on the detected image, the reconstructed image, and the perceptual map of the reconstructed image, and obtaining a defect detection result;
[0089] The memory 1020 may include one or more computer program products, which may include various types of computer-readable storage media, such as, for example, volatile and / or non-volatile memory. One or more computer program commands may be stored in the computer-readable storage media, and the processor 1010 may execute the program commands to perform the functions of the apparatus for defect detection by neural network training according to the embodiment of the present invention and / or other desired functions and / or the method for defect detection by neural network training according to the embodiment of the present invention. Various application programs and various data may also be stored in the computer-readable storage media.
[0090] The following describes a computer-readable storage medium storing a computer program according to an embodiment of the present invention, which can be executed by a processor to perform the following steps: inputting a target image, performing reconstruction on the target image using a reconstruction network including a first sub-module, obtaining a reconstructed image from which defects have been removed after reconstruction and a perceptual map of the reconstructed image, the first sub-module being used to screen and process features by screening feature channels and extracting spatial relationships between each channel; and performing defect segmentation based on the target image, the reconstructed image, and the perceptual map of the reconstructed image to obtain a defect detection result.
[0091] Of course, the above specific embodiments are merely examples and are not limiting, and those skilled in the art may combine and integrate some steps and devices from the above separately described embodiments based on the concept of the present invention to achieve the effects of the present invention. Such combined and integrated embodiments are also included in the present invention and will not be described here.
[0092] The advantages, advantages, and effects mentioned in the present invention are merely illustrative and not limiting, and these advantages, advantages, and effects are not essential to each embodiment of the present invention. Furthermore, the specific details disclosed above are merely for illustrative purposes and easy understanding, and are not limiting. The details do not limit the following. In other words, it is essential to use the specific details to realize the present invention.
[0093] Block diagrams of components, devices, equipment, and systems according to the present invention are merely exemplary and do not require or suggest that they be connected, laid out, or arranged in the manner shown in the block diagrams. Those skilled in the art will recognize that these components, devices, equipment, and systems can be connected, laid out, or arranged in any manner. The terms "including," "including," "having," and "including" are open-ended terms and are used interchangeably. As used herein, the words "or" and "and" refer to the words "and / or" and are used interchangeably unless the context clearly dictates otherwise. As used herein, the word "for example" refers to "including but not limited to" and are used interchangeably.
[0094] The step flow charts and the above descriptions of the present invention are merely exemplary and are not intended to require or suggest that the steps of each embodiment be performed in the order presented. As will be recognized by those skilled in the art, the steps within the above embodiments can be performed in any order. Words such as "then," "thereafter," and "next" are not intended to limit the order of the steps. These words are merely to guide the reader in reading these method descriptions. Furthermore, any reference to a singular element using the articles "a," "one," or "the" does not limit that element to the singular.
[0095] Furthermore, the steps and devices in each embodiment of this specification are not limited to being performed in a particular embodiment, and in fact, new embodiments can be constructed by combining some of the relevant steps and devices of each embodiment of this specification based on the concept of the present invention, and these new embodiments are also included within the scope of the present invention.
[0096] Each operation in the methods described above can be implemented by any suitable means capable of performing the corresponding function, including various hardware and / or software components and / or modules, including, but not limited to, a circuit, an application specific integrated circuit (ASIC), or a processor.
[0097] Each illustrative logic block, module, circuit, etc. may be implemented or described using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any commercially available processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0098] The steps embodying the methods or algorithms described herein may be embodied directly in hardware, in a software module executed by a processor, or a combination of the two. The software module may be stored on any form of tangible storage medium. Possible storage media include, for example, random access memory (RAM), read-only memory (ROM), fast flash memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, etc. The storage medium may be coupled to the processor such that the processor reads information from, and writes information to, the storage medium. In the alternative, the storage medium may be integral to the processor. A software module may be a single command or multiple commands, and may be distributed over several different code segments, among different programs, and across multiple storage media.
[0099] The methods invented herein comprise one or more acts for achieving the described method. The methods and / or acts may be interchanged with one another without departing from the claims. In other words, except where a specific order of acts is specified, the order and / or execution of specific acts may be changed without departing from the claims.
[0100] The functions can be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the functions can be stored as one or more commands on a computer-readable medium. The storage medium can be any available medium that can be accessed by a computer. The following examples are not limiting. Such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic memory, or any other tangible medium that carries or stores desired program code in the form of commands or data structures and is accessible to a computer. As used herein, "disc" includes compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), soft magnetic discs, and Blu-ray discs.
[0101] Thus, a computer program product can perform the operations described herein. For example, such a computer program product can be a tangible computer-readable medium having instructions tangibly stored (and / or encoded) thereon, which instructions can be executed by one or more processors to perform the operations described herein. The computer program product can include packaging materials.
[0102] The software or commands may be transmitted over a transmission medium, such as coaxial cable, fiber optics, twisted pair cable, digital subscriber line (DSL), or wireless technologies such as infrared, radio, or microwave, to transmit the software from a website, server, or other remote source.
[0103] Alternatively, modules and / or other suitable means for performing the methods and techniques described herein may be downloaded and / or otherwise obtained by a user terminal and / or base station, as appropriate. For example, such devices may be coupled to a server to facilitate transmission of the means for performing the described methods. Alternatively, the methods described herein may be provided via a storage medium (e.g., RAM, ROM, physical storage medium such as a CD or soft magnetic disk) such that a user terminal and / or base station obtains the methods when coupled to or provides the storage medium to the device. Alternatively, any other suitable technology for providing the methods and techniques described herein to a device may be used.
[0104] Other examples and implementations are within the scope and spirit of the present invention and the claims. For example, depending on the nature of the software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Features implementing the functions may also be physically located, including distribution of functions so that portions of the functions are implemented in different physical locations. Moreover, as used herein, including in the claims, "or" in a list, such as "at least one," means a disjunctive list. That is, a list such as "at least one of A, B, or C" means A, B, or C, or AB, AC, or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not imply that the described example is optimal or better than other examples.
[0105] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings of the claims. Moreover, the claims are not limited to the specific processes, apparatus, manufacture, compositions of events, means, methods, and acts described above. Existing or developed processes, apparatus, manufacture, compositions of events, means, methods, or acts may be used that perform substantially the same function or achieve substantially the same result as those described herein. Accordingly, the claims include within their scope any such processes, apparatus, manufacture, compositions of events, means, methods, or acts.
[0106] The above teachings of the invention provided will enable one skilled in the art to make or use the invention. Various modifications of these teachings will be apparent to those skilled in the art, and the general principles defined herein may be applied in other ways without departing from the scope of the invention. Thus, the present invention is not intended to be limited to the teachings shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0107] The foregoing description has been presented for purposes of illustration and description, and is not intended to limit the embodiments of the present invention to the precise form disclosed herein. Having considered the foregoing examples and embodiments, certain variations, modifications, variations, additions, and subcombinations thereof will be apparent to those skilled in the art.
Claims
1. A neural network training method executed by a neural network training device, comprising: inputting training images and calculating a perceptual map of said training images; generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect image by a reconstruction network including a first sub-module to obtain a post-reconstruction defect-removed training reconstructed image and a perceptual map of the training reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; generating training defect images based on at least the training images and a perceptual map of the training images, generating a defect mask image and generating one or more defect texture images based on the training images; A method for training a neural network, comprising: generating a training defect image by fusing the training image with one or more defect texture images based on a defect control panel and the defect mask image and utilizing a multi-scale perceptual map of the training image, wherein the defect control panel is used to control the proportion of the defect texture image in the training image.
2. The step of calculating a perceptual map of the training images comprises:
2. The method of claim 1, further comprising computing a multi-scale perceptual map of the training images.
3. 2. The method of claim 1, wherein the first submodule includes a dilated convolutional network for expanding a feature receptive field.
4. 2. The neural network training method according to claim 1, wherein the reconstruction network comprises an encoder and a decoder, and a first sub-module is installed at each layer of the encoder and decoder, respectively.
5. A neural network training method executed by a neural network training device, comprising: inputting training images and calculating a perceptual map of said training images; generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect image by a reconstruction network including a first sub-module to obtain a post-reconstruction defect-removed training reconstructed image and a perceptual map of the training reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images, and performing defect detection by extracting, using a division network including a second submodule, differences between the training image and the training reconstructed image, differences between the training image and a perceptual map of the training reconstructed image, and differences between the training image and the training reconstructed image and the perceptual map of the training reconstructed image.
6. 6. The neural network training method according to claim 5, wherein the partitioned network comprises an encoder and a decoder, and a first sub-module is installed at each layer of the encoder and decoder, respectively.
7. A defect detection method executed by a defect detection device, comprising: a step of inputting a detected image, performing reconstruction on the detected image by a reconstruction network including a first sub-module, and obtaining a reconstructed image with defects removed after reconstruction and a perceptual map of the reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect segmentation based on the detected image, the reconstructed image, and a perceptual map of the reconstructed image to obtain a defect detection result; The step of performing defect segmentation based on the detected image, the reconstructed image, and a perceptual map of the reconstructed image to obtain a defect detection result includes: A defect detection method characterized by including extracting, by a division network including a second submodule, the differences between the detected image and the reconstructed image, the differences between the detected image and the perceptual map of the reconstructed image, and the differences between the detected image and the reconstructed image and the perceptual map of the reconstructed image, thereby detecting defects.
8. an input unit for inputting training images and computing perceptual maps for said training images; a generation unit for generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; a reconstruction unit that performs reconstruction on the training defect image using a reconstruction network including a first sub-module to obtain a reconstructed defect-removed training reconstructed image and a perceptual map of the training reconstructed image, the first sub-module being used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; a training unit that performs defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; The generating unit comprises: generating a defect mask image and generating one or more defect texture images based on the training images; A neural network training device, characterized in that a training defect image is generated by fusing the training image with one or more defect texture images based on a defect control panel and the defect mask image and utilizing a multi-scale perceptual map of the training image, and the defect control panel is used to control the proportion of the defect texture image in the training image. an input unit for inputting training images and calculating a perceptual map of said training images; a generation unit for generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; a reconstruction unit that performs reconstruction on the training defect image using a reconstruction network including a first sub-module to obtain a reconstructed defect-removed training reconstructed image and a perceptual map of the training reconstructed image, the first sub-module being used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; a training unit that performs defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; The training unit comprises: A neural network training device characterized in that defects are detected by extracting, using a division network including a second submodule, differences between the training image and the training reconstructed image, differences between the training image and the perceptual map of the training reconstructed image, and differences between the training image and the training reconstructed image and the perceptual map of the training reconstructed image.
10. a processor; a memory in which computer program commands are stored; The computer program instructions, when executed by the processor, cause the processor to: inputting training images and calculating a perceptual map of said training images; generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect image by a reconstruction network including a first sub-module to obtain a post-reconstruction defect-removed training reconstructed image and a perceptual map of the training reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; generating training defect images based on at least the training images and a perceptual map of the training images, generating a defect mask image and generating one or more defect texture images based on the training images; generating a training defect image by fusing the training image with one or more defect texture images based on a defect control panel and the defect mask image and utilizing a multi-scale perceptual map of the training image, wherein the defect control panel is used to control the proportion of the defect texture image in the training image.
11. A processor; a memory in which computer program commands are stored; The computer program instructions, when executed by the processor, cause the processor to: inputting training images and calculating a perceptual map of said training images; generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect image by a reconstruction network including a first sub-module to obtain a post-reconstruction defect-removed training reconstructed image and a perceptual map of the training reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images, and performing defect detection by extracting, using a division network including a second submodule, differences between the training image and the training reconstructed image, differences between the training image and a perceptual map of the training reconstructed image, and differences between the training image and the training reconstructed image and the perceptual map of the training reconstructed image.
12. A computer program that, when executed by a processor, inputting training images and calculating a perceptual map of said training images; generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect image by a reconstruction network including a first sub-module to obtain a post-reconstruction defect-removed training reconstructed image and a perceptual map of the training reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; generating training defect images based on at least the training images and a perceptual map of the training images, generating a defect mask image and generating one or more defect texture images based on the training images; generating a training defect image by fusing the training image with one or more defect texture images based on a defect control panel and the defect mask image and utilizing a multi-scale perceptual map of the training image, wherein the defect control panel is used to control the proportion of the defect texture image in the training image.
13. A computer program, the computer program, when executed by a processor, inputting training images and calculating a perceptual map of said training images; generating training defect images based on at least the training images and a perceptual map of the training images, the training defect images including one or more defects generated based on the perceptual map of the training images; performing reconstruction on the training defect image by a reconstruction network including a first sub-module to obtain a post-reconstruction defect-removed training reconstructed image and a perceptual map of the training reconstructed image, wherein the first sub-module is used for feature screening and processing by screening feature channels and extracting spatial relationships between each channel; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images to train a neural network and adjust parameters of the neural network; performing defect detection based on the training images, the training reconstructed images, and a perceptual map of the training reconstructed images, and performing defect detection by extracting, using a segmentation network including a second submodule, differences between the training image and the training reconstructed image, differences between the training image and a perceptual map of the training reconstructed image, and differences between the training image and the training reconstructed image and the perceptual map of the training reconstructed image.