A defect detection method and system based on image processing and deep learning
By combining image processing and deep learning in defect detection, shadow removal, color neutralization, brightness neutralization and contrast enhancement are solved, and the problem of degradation of defect detection effect in complex environments is achieved, achieving higher detection accuracy and reliability.
Patent Information
- Application Number
- CN202510512592.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Under complex lighting conditions, traditional defect detection methods are difficult to accurately identify and locate defects, and machine learning and deep learning methods also reduce detection effects in complex environments.
Defect detection methods based on image processing and deep learning are adopted to process images through shadow removal, color neutralization, brightness neutralization and contrast enhancement, reducing the influence of environmental factors, and using ResNet and Unet hybrid networks to locate defect locations.
It effectively avoids the impact of complex environments on detection effects and improves the accuracy and reliability of defect detection under complex lighting conditions.
Smart Images

Figure CN120031881B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of laser welding image processing, and specifically relates to a defect detection method and system based on image processing and deep learning. Background Art
[0002] In recent years, defect detection methods based on machine learning and deep learning have received extensive attention and applications, and have been proven to be effective in detecting defects in many industrial fields. However, complex lighting conditions can produce shadow occlusions, and contrast and color distortions make it more difficult to accurately identify and locate defects.
[0003] Traditional defect detection methods detect defects through computer vision techniques and statistical analysis of image processing. However, such solutions are difficult to detect random defects, especially in dynamic scenes with complex lighting conditions. Although the methods based on machine learning and deep learning have a certain generalization ability, can adapt to more scenarios, and the detection accuracy is also higher than traditional solutions, they rely on the quality of the training dataset and also encounter the problem of decreased detection effect in dynamic environments with complex lighting conditions. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a defect detection method and system based on image processing and deep learning to solve the technical problem of how to avoid the influence of complex environments on the detection effect.
[0005] This application discloses a defect detection method based on image processing and deep learning, including the following steps:
[0006] S1. Obtain a first image, and identify and remove the shadows in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image;
[0007] S2. Perform color neutralization on the second image based on the CIELAB color space to obtain a third image;
[0008] S3. Eliminate the brightness affecting feature extraction by separating and processing the illumination component and the reflection component of the third image to obtain a fourth image;
[0009] S4. Based on histogram equalization, convert the fourth image to the RGB channel to obtain a fifth image;
[0010] S5. Locate the defect positions of the fifth image based on a hybrid network of ResNet and Unet.
[0011] In some possible implementation manners, the identifying and removing the shadows in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image includes:
[0012] S11. Feed the first image into the basic architecture of VGG16, and extract the preliminary features of the first image based on the basic architecture;
[0013] S12. Based on the convolutional layer and pooling layer of VGG16, and apply dilated convolution at different levels to capture features of different scales for identifying the shadow area in the image;
[0014] S13. Aggregate the features of different scales based on the CAN encoder to form a feature representation for identifying the shadow area;
[0015] S14. Generate a mask for indicating the shadow position based on the feature representation;
[0016] S15. Through the inference of the DHAN network, output the second image with the shadow removed.
[0017] In some possible implementation manners, the color neutralization of the second image based on the CIELAB color space to obtain the third image includes:
[0018] S21. Convert the second image from the RGB color space to the CIELAB color space;
[0019] S22. Adjust the relative brightness of each pixel's color channel to adapt to the target lighting condition;
[0020] S23. Convert the second image with adjusted relative brightness back from the CIELAB color space to the RGB color space to obtain the third image.
[0021] In some possible implementation manners, the elimination of the brightness affecting feature extraction by separating and processing the illumination component and the reflection component of the third image to obtain the fourth image includes:
[0022] S31. Separate the illumination component and the reflection component, and the separation of the illumination component and the reflection component includes: Filter the third image at each resolution based on a Gaussian filter to estimate the illumination component of the third image; Obtain the reflection component by subtracting the illumination component from the third image;
[0023] S32. Reduce the brightness range of the illumination component through non - linear mapping to make the brightness distribution of the image more uniform;
[0024] S33. Combine the adjusted illumination component and the reflection component, and reconstruct to obtain the fourth image with neutralized brightness.
[0025] In some possible implementation manners, the positioning of the defect position of the fifth image based on the ResNet and Unet hybrid network includes:
[0026] S51. Input the fifth image into the ResNet and Unet hybrid network, and extract the feature representation of the fifth image based on the Resnet152 encoder. The feature representation captures the abstract patterns and structural information of the fifth image.
[0027] S52. The feature representation is fed into the Unet decoder. The Unet decoder gradually restores the spatial resolution of the image through upsampling and convolution operations, and at the same time fuses the feature representation to accurately locate the defects.
[0028] As a second aspect of the present application, there is also provided a defect detection system based on image processing and deep learning, including: a shadow removal module, a color neutralization module, a brightness neutralization module, a contrast enhancement module, and a defect detection module; wherein, the shadow removal module is used to obtain a first image, identify and remove the shadow in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image; the color neutralization module is used to perform color neutralization on the second image based on the CIELAB color space to obtain a third image; the brightness neutralization module is used to eliminate the brightness affecting feature extraction by separating and processing the illumination component and the reflection component of the third image to obtain a fourth image; the contrast enhancement module is used to convert the fourth image into RGB channels based on histogram equalization to obtain a fifth image; the defect detection module is used to locate the defect position of the fifth image based on the ResNet and Unet hybrid network.
[0029] In some possible implementation manners, the identifying and removing the shadow in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image includes: feeding the first image into the basic architecture of VGG16, and extracting the preliminary features of the first image based on the basic architecture; based on the convolutional layer and pooling layer of VGG16, and applying dilated convolutions at different levels to capture features of different scales for identifying the shadow area in the image; aggregating the features of different scales based on the CAN encoder to form a feature representation for identifying the shadow area; generating a mask for indicating the shadow position based on the feature representation; and through the inference of the DHAN network, outputting the second image with the shadow removed.
[0030] In some possible implementation manners, the performing color neutralization on the second image based on the CIELAB color space to obtain a third image includes: converting the second image from the RGB color space to the CIELAB color space; performing relative brightness adjustment on each pixel's color channel to adapt to the target illumination condition; and converting the second image with relative brightness adjustment back from the CIELAB color space to the RGB color space to obtain a third image.
[0031] In some possible embodiments, eliminating the brightness that affects feature extraction by separating and processing the illumination component and the reflection component of the third image to obtain a fourth image includes: separating the illumination component and the reflection component, and the separating of the illumination component and the reflection component includes: filtering the third image at each resolution based on a Gaussian filter to estimate the illumination component of the third image; obtaining the reflection component by subtracting the illumination component from the third image; reducing the brightness range of the illumination component through non-linear mapping to make the brightness distribution of the image more uniform; merging the adjusted illumination component and the reflection component to reconstruct the fourth image with neutralized brightness.
[0032] In some possible embodiments, locating the defect position of the fifth image based on a hybrid network of ResNet and Unet includes: inputting the fifth image into the hybrid network of ResNet and Unet, extracting the feature representation of the fifth image based on the Resnet152 encoder, and the feature representation captures the abstract patterns and structural information of the fifth image; the feature representation will be fed into the Unet decoder; the Unet decoder part gradually restores the spatial resolution of the image through upsampling and convolution operations while fusing the feature representation to facilitate accurate defect location.
[0033] Beneficial effects: By processing environmental factors that affect image defect detection through shadow removal, color neutralization, brightness neutralization, and contrast enhancement, and then using a deep learning solution for defect detection, the influence of complex environments on the detection effect is solved, and the decline in detection effect in a dynamic environment with complex lighting conditions is avoided.
[0034] Other advantages, objectives, and features of the present application will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on an examination of the following text, or can be taught from the practice of the present application. The objectives and other advantages of the present application can be realized and obtained through the following specification. Description of the Drawings
[0035] The embodiments described below with reference to the drawings are exemplary and are intended to explain and illustrate the present application, and should not be construed as limiting the protection scope of the present application.
[0036] Figure 1 is the system flow chart of the present application;
[0037] Figure 2 is the system structure diagram of the present application;
[0038] Among them: 1. Shadow removal module; 2. Color neutralization module; 3. Brightness neutralization module; 4. Contrast enhancement module; 5. Defect detection module. Specific implementation manner
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. The components of the embodiments of the present application described and illustrated herein generally may be arranged and designed in a variety of different configurations.
[0040] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0041] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0042] In the above description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0043] As Figure 1 shown, the present application discloses a defect detection method based on image processing and deep learning, including the following steps:
[0044] S1. Obtain a first image, and identify and remove the shadow in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image. The first image is an industrial part or product image captured by a camera on a production line, that is, the original image.
[0045] In some embodiments, the identifying and removing the shadow in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image includes:
[0046] S11. Send the first image into the basic architecture of VGG16, and extract preliminary features of the first image based on the basic architecture; through the convolutional layer and pooling layer of VGG16, the network starts to extract features of the image. As the network depth increases, the features gradually shift from local details to global context. VGG16 is essentially a convolutional neural network that uses convolutional kernels and pooling layers to represent the input image. The preliminary features include: edge and texture information of the image, shape information, and semantic information.
[0047] S12. Convolutional and pooling layers based on VGG16, and dilated convolutions are applied at different levels to capture features of different scales for identifying shadow regions in images. Features of different scales include large-scale structures: overall shape and contour; small-scale structures: local pattern details.
[0048] S13. Aggregate the features of different scales based on a CAN encoder to form a feature representation for identifying shadow regions.
[0049] S14. Generate a mask for indicating the shadow position based on the feature representation. The mask is generated based on the feature representation after learning the specific patterns and attributes of shadows.
[0050] S15. Through the inference of the DHAN network, output the second image with shadows removed. Since the DHAN network has original pre-trained weights, which means the network has been pre-trained on a large amount of data and can identify and process various shadow situations without training from scratch.
[0051] A pre-trained Dual Hierarchical Aggregation network (DHAN) is constructed using a convolutional neural network and a context aggregation network, and a series of convolutional operations and hierarchical aggregation of context features are used to identify and remove shadows in images.
[0052] S2. Perform color neutralization on the second image based on the CIELAB color space to obtain a third image. The result after color neutralization is more suitable for the human eye's visual color.
[0053] In some embodiments, performing color neutralization on the second image based on the CIELAB color space to obtain a third image includes:
[0054] S21. Convert the second image from the RGB color space to the CIELAB color space.
[0055] S22. Perform relative luminance adjustment on each pixel's color channel to adapt to the target lighting conditions.
[0056] In this embodiment, the CIELAB color space includes luminance (L*) and two color channels (a and b), which is more suitable for describing colors perceived by the human vision;
[0057] Based on the transformation formula, perform relative luminance adjustment on each of the two color channels a and b. The transformation formula is as follows:
[0058] ;
[0059] Among them, R, G, and B are the RGB values in the first image, namely the Red, Green, and Blue color values. R', G', and B' are the transformed RGB values, namely the new Red, Green, and Blue color values. , , is the white point correction factor based on the current lighting conditions. Here, W is a 3x1 vector that contains the correction factors for the three channels (X, Y, Z) respectively, depending on the color temperature and intensity of the light source. In the CIELAB color space, these white point correction factors can be determined by looking up the standard white point D65 corresponding to the color temperature of the light source. The transformation formula is used to adjust the color channels of each pixel to adapt to the target lighting conditions.
[0060] S23. Convert the second image with adjusted relative luminance back from the CIELAB color space to the RGB color space to obtain the third image.
[0061] The VonKries color adaptation transformation is adopted to improve the color clarity of the image. The idea is to use diagonal matrix transformation to describe the relationship between the colors of the same object surface under different illuminations. Based on VonKries color adaptation, the conversion is carried out from the source color to the target color in the (long), medium, and short) color spaces. The purpose of this conversion is to make the RPG light source color of the data sample adapt to different light sources, so as to maintain a constant white color. This will make the colors of the image more stable, improve the robustness of the system to color changes under different lighting conditions, enable it to extract features more conducive to defect recognition, and thus improve the accuracy and reliability of detection.
[0062] The above steps S21 - S23 are specifically as follows: First, normalize the RGB values to the range of [0, 1]:
[0063] ;
[0064] ;
[0065] ;
[0066] Subsequently, perform matrix transformation:
[0067] ;
[0068] Next, convert the result of the matrix transformation to the CIELAB color space and perform color neutralization:
[0069] ;
[0070] ;
[0071] ;
[0072] Among them, is a non - linear transformation function:
[0073] ;
[0074] , , are the X, Y, and Z values of the reference white point respectively. Here, the D65 white point is used:
[0075] ;
[0076] ;
[0077] ;
[0078] By adjusting the values of the a and b color channels, the white objects in the image can maintain a consistent visual effect under different lighting conditions.
[0079] Finally, the result after neutralization processing is converted back to the RGB color space through the following steps:
[0080] ;
[0081] ;
[0082] ;
[0083] Among them, is 's inverse function.
[0084] ;
[0085] Denormalization:
[0086] ;
[0087] ;
[0088] ;
[0089] That is, it is converted from the CIELAB color space back to the RGB color space.
[0090] S3. Eliminate the highlights that affect feature extraction by separating and processing the illumination component and the reflection component of the third image, and obtain the fourth image.
[0091] In some embodiments, obtaining a fourth image by separating and processing the illumination component and the reflection component of the third image to eliminate the brightness affecting feature extraction includes:
[0092] S31. Separating the illumination component and the reflection component, where separating the illumination component and the reflection component includes: filtering the third image at each resolution based on a Gaussian filter to estimate the illumination component of the third image; obtaining the reflection component by subtracting the illumination component from the third image;
[0093] S32. Reducing the brightness range of the illumination component through non - linear mapping to make the brightness distribution of the image more uniform; the brightness range here mainly removes the overly dark and overly bright outliers, and then narrows the gap in brightness at different positions. Generally, the brightness range is the brightness range of the pixel values at the 2% and 98% positions of the original. The non - linear change here is essentially taking the logarithm.
[0094] S33. Combining the adjusted illumination component and the reflection component and reconstructing to obtain the fourth image with neutralized brightness.
[0095] The above steps S31 - S33 are specifically as follows:
[0096] First, use a Gaussian filter to smooth the image at different resolution scales:
[0097] ,
[0098] where represents the resolution scale, represents the convolution operation.
[0099] Subsequently, for each scale , separate the illumination component and the reflection component. The image after Gaussian filtering is used as the illumination component, and the reflection component is calculated as the difference between the original image and the illumination component:
[0100] ;
[0101] ;
[0102] Immediately afterwards, perform non - linear mapping on the separated illumination component L to compress the brightness range:
[0103] ;
[0104] Here, 0 < r < 1 is an adjustable hyperparameter, usually 0.5 can be taken.
[0105] Next, recombine the adjusted illumination component and the reflection component to generate the output image at each scale:
[0106] ;
[0107] Finally, combine the output images Oσ at all scales to generate the final luminance-neutralized image O.
[0108] ;
[0109] Among them, n = 3 is the number of scales.
[0110] Process the image based on Multi-Scale Retinex (MSR), so that what the visual system perceives is not the absolute luminance, but the relative luminance, that is, the luminance change in the local image area. Neutralize the luminance on the premise of retaining the relative luminance and not changing the chromaticity and color components, which helps to eliminate abnormal light that may affect feature extraction.
[0111] S4. Based on histogram equalization, convert the fourth image to the RGB channel to obtain the fifth image.
[0112] Convert the fourth image to the RGB channel, which helps to make the differences between different objects in the fourth image more obvious, thus helping to detect defects and other abnormal situations. By equalizing the histogram, the overall image contrast increases, so as to obtain better features for the deep learning model.
[0113] S5. Locate the defect position of the fifth image based on the hybrid network of ResNet and Unet. The purpose is to improve the contrast.
[0114] The hybrid network of ResNet and Unet consists of two key components: Resnet and U-net networks. Using transfer learning, make Resnet act as the encoder of the U-net network. In this network structure, the last fully connected layer of the Resnet network is removed and directly integrated with the decoder.
[0115] The energy function of the Unet network is set to the softmax cross-entropy loss function:
[0116] ;
[0117] Among them, represents the activation of the feature channel k at the pixel position x, that is, the processing result after passing through the neural network. is the approximate maximum function, that is, for the k with the maximum activation of, is approximately equal to 1, and other k are approximately equal to 0. Based on this, the deviation of the cross-entropy from 1 at each position can be obtained. Among them is the true label of each pixel, is the weight map, and the initial weights during the training process are drawn from a Gaussian distribution with a standard deviation. The weight map is mainly used to compensate for the different frequencies of different types in the recognition types. The weight map is calculated as:
[0118] ;
[0119] where, represents the initial weight map assuming that the classes are balanced, represents the distance to the nearest boundary, represents the distance to the second nearest boundary.
[0120] In some embodiments, the defect location of the fifth image is located by the hybrid network based on ResNet and Unet, including:
[0121] S51. Input the fifth image into the hybrid network of ResNet and Unet, and extract the feature representation of the fifth image based on the Resnet152 encoder. The feature representation captures the abstract patterns and structural information of the fifth image;
[0122] S52. The feature representation will be fed into the Unet decoder; the Unet decoder part gradually restores the spatial resolution of the image through upsampling and convolution operations, and at the same time fuses the feature representation to facilitate the accurate location of defects. The input of the Unetr decoder is the feature representation of the encoder, and then through upsampling, the dimensions of the features are converted. Here, the transposed convolution method is used to supplement the missing values in the features. After passing through the upsampling layer, feature fusion will be performed in the skip connection layer. The upsampled features will be concatenated with the features of the corresponding levels of the encoder. The next step is the convolution layer. After fusing the feature representation, the decoder will apply a series of convolution layers to further refine the features and gradually reduce the number of features. After the convolution layer, a non-linear activation function (ReLU) is applied to increase the non-linear expression ability of the model. Finally, the decoder generates the final output map through a 1x1 convolution layer.
[0123] By processing environmental factors affecting image defect detection such as shadow removal, color neutralization, brightness neutralization, and contrast enhancement, and then using a deep learning solution for defect detection, the influence of complex environments on the detection effect is solved, and the decline in the detection effect in a dynamic environment with complex lighting conditions is avoided.
[0124] As a second aspect of the present application, there is also provided a defect detection system based on image processing and deep learning, including: a shadow removal module, a color neutralization module, a brightness neutralization module, a contrast enhancement module, and a defect detection module; wherein, the shadow removal module is configured to obtain a first image, and identify and remove the shadow in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image; the color neutralization module is configured to perform color neutralization on the second image based on the CIELAB color space to obtain a third image; the brightness neutralization module is configured to eliminate the brightness affecting feature extraction by separating and processing the illumination component and the reflection component of the third image to obtain a fourth image; the contrast enhancement module is configured to convert the fourth image into RGB channels based on histogram equalization to obtain a fifth image; the defect detection module is configured to locate the defect position of the fifth image based on a hybrid network of ResNet and Unet.
[0125] In some embodiments, the identifying and removing the shadow in the first image based on convolutional operations and hierarchical aggregation of context features to obtain a second image includes: sending the first image into the basic architecture of VGG16, and extracting preliminary features of the first image based on the basic architecture; based on the convolutional layer and pooling layer of VGG16, and applying dilated convolution at different levels to capture features of different scales for identifying the shadow area in the image; aggregating the features of different scales based on the CAN encoder to form a feature representation for identifying the shadow area; generating a mask for indicating the shadow position based on the feature representation; through the inference of the DHAN network, outputting the second image with the shadow removed.
[0126] In some embodiments, the performing color neutralization on the second image based on the CIELAB color space to obtain a third image includes: converting the second image from the RGB color space to the CIELAB color space; performing relative brightness adjustment on each pixel's color channel to adapt to the target illumination condition; converting the second image with relative brightness adjustment back from the CIELAB color space to the RGB color space to obtain a third image.
[0127] In some embodiments, the eliminating the brightness affecting feature extraction by separating and processing the illumination component and the reflection component of the third image to obtain a fourth image includes: separating the illumination component and the reflection component, and the separating the illumination component and the reflection component includes: filtering the third image at each resolution based on a Gaussian filter to estimate the illumination component of the third image; obtaining the reflection component by subtracting the illumination component from the third image; reducing the brightness range of the illumination component through non-linear mapping to make the brightness distribution of the image more uniform; merging the adjusted illumination component and the reflection component to reconstruct and obtain the fourth image with brightness neutralization.
[0128] In some embodiments, the defect location of the fifth image is located based on the ResNet and Unet hybrid network, including: inputting the fifth image into the ResNet and Unet hybrid network, extracting the feature representation of the fifth image based on the Resnet152 encoder, and the feature representation captures the abstract patterns and structural information of the fifth image; the feature representation will be fed into the Unet decoder; the Unet decoder part gradually restores the spatial resolution of the image through upsampling and convolution operations, while fusing the feature representation to facilitate the accurate location of defects.
[0129] Those skilled in the art can understand that although some embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of this application and forms different embodiments.
[0130] Those skilled in the art can understand that the descriptions of the various embodiments have their respective emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0131] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application. Such modifications and variations all fall within the scope defined by the appended claims. The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A defect detection method based on image processing and deep learning, characterized in that: The steps include: S1, obtaining a first image, identifying and removing shadows in the first image based on convolution operations and hierarchical aggregation of contextual features to obtain a second image, including: S11, sending the first image to the infrastructure of VGG16, and extracting preliminary features of the first image based on the infrastructure; S12, based on the convolution layer and pooling layer of VGG16, and applying dilated convolution at different levels to capture features of different scales for identifying shadow areas in the image; S13, aggregating the features of different scales based on the CAN encoder to form a feature representation for identifying shadow areas; S14, generating a mask for indicating the position of the shadow based on the feature representation; S15, after inference by the DHAN network, outputting the second image with the shadow removed; S2. Performing color neutralization on the second image based on the CIELAB color space to obtain a third image; S3, eliminating the light that affects feature extraction based on separating and processing the illumination component and the reflection component of the third image to obtain a fourth image; comprising: S31, separating an illumination component and a reflection component, wherein the separation of the illumination component and the reflection component comprises: filtering the third image at each resolution based on a Gaussian filter to estimate the illumination component of the third image; and obtaining the reflection component by subtracting the illumination component from the third image; S32, reducing the brightness range of the illumination component by nonlinear mapping, so that the brightness distribution of the image is more uniform; S33, combining the adjusted illumination component and reflection component to reconstruct the fourth image after brightness neutralization; S4. Based on histogram equalization, convert the fourth image into RGB channels to obtain a fifth image; S5. Locate the defect position of the fifth image based on a ResNet and Unet hybrid network.
2. The defect detection method based on image processing and deep learning according to claim 1, characterized in that: The step of performing color neutralization on the second image based on the CIELAB color space to obtain the third image includes: S21, converting the second image from the RGB color space to the CIELAB color space; S22, adjusting the relative brightness of the color channel of each pixel to adapt to the target lighting conditions; S23: Convert the second image after relative brightness adjustment from the CIELAB color space back to the RGB color space to obtain a third image.
3. The defect detection method based on image processing and deep learning according to claim 2, characterized in that: The locating the defect position of the fifth image based on the ResNet and Unet hybrid network includes: S51, inputting the fifth image into a ResNet and Unet hybrid network, and extracting a feature representation of the fifth image based on a Resnet152 encoder, wherein the feature representation captures abstract patterns and structural information of the fifth image; S52, the feature representation will be sent to the Unet decoder; the Unet decoder part gradually restores the spatial resolution of the image through upsampling and convolution operations, and at the same time fuses the feature representation.
4. A defect detection system based on image processing and deep learning, characterized in that: include: A shadow removal module, a color neutralization module, a brightness neutralization module, a contrast enhancement module and a defect detection module; wherein the shadow removal module is used to obtain a first image, and based on the convolution operation and the hierarchical aggregation of context features, identify and remove the shadow in the first image to obtain a second image, including: sending the first image to the infrastructure of VGG16, extracting preliminary features of the first image based on the infrastructure; based on the convolution layer and pooling layer of VGG16, and applying dilated convolution at different levels to capture features of different scales for identifying shadow areas in the image; based on the CAN encoder, aggregating the features of different scales to form a feature representation for identifying shadow areas; based on the feature representation, generating a mask for indicating the position of the shadow; after the inference of the DHAN network, outputting the second image with the shadow removed; the color neutralization module is used to perform a color sparse multiplication of the second image based on the CIELAB color space. The third image is obtained by color neutralization; the brightness neutralization module is used to eliminate the light that affects feature extraction based on separating and processing the illumination component and the reflection component of the third image to obtain a fourth image; including: separating the illumination component and the reflection component, wherein the separation of the illumination component and the reflection component includes: filtering the third image at each resolution based on a Gaussian filter to estimate the illumination component of the third image; obtaining the reflection component by subtracting the illumination component from the third image; reducing the brightness range of the illumination component by nonlinear mapping to make the brightness distribution of the image more uniform; merging the adjusted illumination component and reflection component to reconstruct the fourth image after brightness neutralization; the contrast enhancement module is used to convert the fourth image into RGB channels based on histogram equalization to obtain a fifth image; the defect detection module is used to locate the defect position of the fifth image based on a ResNet and Unet hybrid network.
5. The defect detection system based on image processing and deep learning according to claim 4, characterized in that: The method of color neutralizing the second image based on the CIELAB color space to obtain the third image includes: converting the second image from the RGB color space to the CIELAB color space; adjusting the relative brightness of the color channel of each pixel to adapt to the target lighting conditions; and converting the second image after the relative brightness adjustment from the CIELAB color space back to the RGB color space to obtain the third image.
6. The defect detection system based on image processing and deep learning according to claim 5, characterized in that: The defect position of the fifth image is located based on the ResNet and Unet hybrid network, including: inputting the fifth image into the ResNet and Unet hybrid network, extracting the feature representation of the fifth image based on the Resnet152 encoder, and the feature representation captures the abstract pattern and structural information of the fifth image; the feature representation is sent to the Unet decoder; the Unet decoder part gradually restores the spatial resolution of the image through upsampling and convolution operations, and fuses the feature representation.
Citation Information
Patent Citations
Chip quality detection system based on image processing and chip surface defect detection method
CN118464928A
Medical image processor and method of judging malignancy
JP2004254742A