Photoelectronic device surface defect visual detection method and system based on image recognition

By employing an image recognition-based visual inspection method for surface defects in optoelectronic devices, and utilizing a ring-shaped LED light source array and deep learning algorithms, automated defect detection has been achieved. This solves the problem of low efficiency in traditional manual inspection and improves both inspection speed and accuracy.

CN120876431APending Publication Date: 2025-10-31HENGYANG QIMING PLASTIC IND CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511032957.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional surface defect detection of optoelectronic devices relies on manual visual inspection, which leads to low efficiency, fatigue, false positives and false negatives, and inaccurate detection results, making it difficult to meet the needs of large-scale production.

Method used

A visual inspection method for surface defects in optoelectronic devices based on image recognition is adopted. It utilizes a ring-shaped LED light source array to output a three-band light source, combined with a DFD-Net dual-branch deep network model and an improved YOLOv5 target detection algorithm. Through multi-scale feature extraction and fusion, automated defect detection is achieved.

Benefits of technology

It achieves automated defect detection, improves detection speed and accuracy, adapts to the needs of large-scale production, reduces human judgment differences, and provides quantitative defect information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876431A_ABST
    Figure CN120876431A_ABST
Patent Text Reader

Abstract

The invention discloses a photoelectronic device surface defect visual detection method and system based on image recognition, and the method comprises the steps: carrying out the processing and fusion of three-waveband image data, and obtaining a three-waveband composite image; constructing a DFD-Net double-branch deep network model, inputting a three-band synthetic image into the model, capturing multi-scale defect geometric features through an ASPP module with increasing voidage, extracting reflective insensitive texture features through a phase-consistent convolutional layer, and performing feature fusion by using a gated cross attention mechanism to obtain a multi-scale defect feature fusion model; performing coarse positioning on the fused feature image based on an improved YOLOv5 target detection algorithm to obtain a defect area image, and calculating a pixel-level mask of the defect area image to obtain a segmented image; and positioning a defect boundary of the segmented image through an NMS non-maximum suppression algorithm, and outputting defect coordinates and size information. And the surface defect detection efficiency of the optoelectronic device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optoelectronic inspection technology, and in particular to a visual inspection method and system for surface defects in optoelectronic devices based on image recognition. Background Technology

[0002] Traditional surface defect detection of optoelectronic devices mainly relies on manual visual inspection. Inspectors observe the device surface with a magnifying glass or microscope and judge the presence and type of defects based on experience. This method leads to low efficiency in manual inspection, making it difficult to meet the needs of large-scale production. Furthermore, prolonged work can easily lead to fatigue, resulting in missed or false detections. In addition, different inspectors have different judgment standards, resulting in low accuracy and efficiency of inspection results. Summary of the Invention

[0003] The purpose of this invention is to solve the above-mentioned problems by designing a visual inspection method and system for surface defects of optoelectronic devices based on image recognition.

[0004] To achieve the above objectives, the technical solution of the present invention further includes the following steps in the above-mentioned visual detection method for surface defects of optoelectronic devices based on image recognition: Three characteristic band light sources are output from a ring-shaped LED light source array to illuminate the optoelectronic device, thereby obtaining three-band image data. The three-band image data are then processed and fused to obtain a three-band composite image. A DFD-Net dual-branch deep network model is constructed. The three-band synthetic image is input into the model. Multi-scale defect geometric features are captured by the ASPP module with increasing hole rate. Reflection-insensitive texture features are extracted by the phase-consistent convolutional layer. Feature fusion is performed by the gated cross-attention mechanism to obtain the fused feature image. The fused feature image is coarsely localized based on the improved YOLOv5 target detection algorithm to obtain the defect region image. The pixel-level mask of the defect region image is calculated using the Swing Transformer to obtain the segmented image. The defect boundaries of the segmented image are located using the Non-Maximum Suppression (NMS) algorithm, and the defect coordinates and size information are output.

[0005] Furthermore, in the above-mentioned visual inspection method for surface defects of optoelectronic devices based on image recognition, the step of illuminating the optoelectronic device with three characteristic band light sources output from a ring LED light source array to obtain three-band image data, and processing and fusing the three-band image data to obtain a three-band composite image, includes: The optoelectronic device is illuminated by a ring-shaped LED light source array that outputs three characteristic wavelengths: 450nm, 650nm, and 850nm. Image data of optoelectronic devices in various wavelengths are acquired using image sensors, including blue light images, red light images, and near-infrared light images; The image data of the optoelectronic device was denoised using a Gaussian filtering algorithm with a standard deviation of 1.5 to obtain denoised image data. After contrast enhancement of the denoised image data, a weighted fusion algorithm is used to assign weights based on the clarity of defect information in each band image to obtain a three-band composite image.

[0006] Furthermore, in the above-mentioned visual detection method for surface defects of optoelectronic devices based on image recognition, the construction of the DFD-Net dual-branch deep network model, inputting the three-band synthesized image into the model, capturing multi-scale defect geometric features through the ASPP module with increasing hole rate, and extracting reflectivity-insensitive texture features through phase-consistent convolutional layers, includes: The ASPP module uses a 4-way parallel dilated convolution structure to perform preliminary feature mapping on the three-band composite image through a 64-channel convolutional layer to obtain a basic feature map. The basic feature map is then convolved to obtain the defect geometric features. Phase-consistent convolutional layers are used to process the phase and amplitude information of the image. The three-band composite image is decomposed into a Gaussian pyramid to generate an image pyramid with five scales. Phase calculation and amplitude suppression are performed at each scale to obtain reflectivity-insensitive texture features.

[0007] Furthermore, in the above-mentioned visual detection method for surface defects of optoelectronic devices based on image recognition, the step of using a gated cross-attention mechanism to perform feature fusion to obtain a fused feature image includes: The multi-scale geometric feature map and the reflectivity-insensitive texture feature map are respectively converted into 128-channel feature vectors through convolutional layers, and the cosine similarity matrix between the two is calculated. Based on the similarity matrix, a spatial attention weight map is generated using the Softmax function, with weight values ​​ranging from 0 to 1. A learnable gating parameter vector is set, and the two types of features are dynamically weighted using the Sigmoid function to obtain a fused feature image.

[0008] Furthermore, in the above-mentioned visual detection method for surface defects of optoelectronic devices based on image recognition, the step of using Swing Transformer to calculate the pixel-level mask of the defect region image to obtain a segmented image includes: The fused feature image is input into the improved YOLOv5 algorithm, which divides the image into grids and combines the improved feature fusion structure to generate prediction results containing defect location and category probability. The confidence threshold is set to 0.65. When the confidence of the prediction result is higher than this threshold, it is determined to be a defective area.

[0009] Furthermore, in the above-mentioned visual detection method for surface defects of optoelectronic devices based on image recognition, the step of using Swing Transformer to calculate the pixel-level mask of the defect region image to obtain a segmented image includes: The image is divided into multiple non-overlapping image blocks, each of which is 16×16 pixels in size; The degree of correlation between each image patch is calculated through a self-attention mechanism, the features of the image patch are mapped back to the pixel level, and each pixel is assigned a probability value of a defect region. When the probability value is greater than 0.5, the pixel is determined to be a defective pixel; otherwise, it is a normal pixel. A pixel-level mask is formed to obtain a segmented image.

[0010] Furthermore, in the above-mentioned visual detection method for surface defects of optoelectronic devices based on image recognition, the step of locating the defect boundary of the segmented image using the NMS non-maximum suppression algorithm and outputting the defect coordinates and size information includes: Extract all edge pixels of the defect region in the binarized image, calculate the gradient value of each edge pixel, sort the edge pixels from high to low according to the gradient value, locate the defect boundary of the segmented image, and output the defect coordinates and size information.

[0011] Furthermore, in the image recognition-based visual inspection system for surface defects in optoelectronic devices, the system includes the following modules: The light source band adjustment module is used to illuminate the optoelectronic device by outputting three characteristic band light sources through a ring LED light source array to obtain three-band image data. The three-band image data is then processed and fused to obtain a three-band composite image. The feature image fusion module is used to construct the DFD-Net dual-branch deep network model. The three-band synthesized image is input into the model. The multi-scale defect geometric features are captured by the ASPP module with increasing hole rate. The reflective insensitive texture features are extracted by the phase-consistent convolutional layer. The feature fusion is performed by the gated cross attention mechanism to obtain the fused feature image. The feature image segmentation module is used to coarsely locate the fused feature image based on the improved YOLOv5 target detection algorithm to obtain the defect region image, and to calculate the pixel-level mask of the defect region image using the Swing Transformer to obtain the segmented image. The defect localization calculation module is used to locate the defect boundaries of the segmented image using the NMS non-maximum suppression algorithm, and output the defect coordinates and size information.

[0012] Furthermore, in the image recognition-based visual inspection system for surface defects in optoelectronic devices, the feature image segmentation module includes the following sub-modules: The partitioning submodule is used to input the fused feature image into the improved YOLOv5 algorithm, divide the image into grids, and combine it with the improved feature fusion structure to generate prediction results containing defect location and category probability. The confidence level submodule is used to set the confidence level threshold to 0.65. When the confidence level of the prediction result is higher than this threshold, it is determined to be a defect area.

[0013] Furthermore, in the image recognition-based visual inspection system for surface defects in optoelectronic devices, the defect localization calculation module includes the following sub-modules: The localization submodule is used to extract all edge pixels of the defect region in the binarized image, calculate the gradient value of each edge pixel, sort the edge pixels from high to low according to the gradient value, locate the defect boundary of the segmented image, and output the defect coordinates and size information.

[0014] Its beneficial effects are as follows: 1. The entire inspection process is automated. From image acquisition, feature extraction, defect localization to information output, no manual intervention is required, significantly improving inspection speed. Compared to manual inspection, it can adapt to the rapid inspection needs of large-scale production lines, shortening the product inspection cycle and improving inspection and production efficiency. 2. The inspection system can operate stably under different lighting conditions. Simultaneously, the multi-scale feature extraction and fusion mechanism allows the system to handle defects of different sizes and types, expanding the inspection range and improving its adaptability to complex inspection scenarios. 3. It avoids the judgment differences caused by human factors in manual inspection, resulting in more accurate inspection results. The output defect coordinates, size, category, and other information provide a quantitative basis for product quality assessment. Attached Figure Description

[0015] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0016] Figure 1 This is a schematic diagram of the first embodiment of the visual inspection method for surface defects of optoelectronic devices based on image recognition in this invention. Figure 2 This is a schematic diagram of a second embodiment of the visual inspection method for surface defects of optoelectronic devices based on image recognition in this invention. Figure 3 This is a schematic diagram of the first embodiment of the visual inspection system for surface defects of optoelectronic devices based on image recognition in this invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0018] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0019] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 As shown, a visual inspection method for surface defects in optoelectronic devices based on image recognition is described. This method includes the following steps: Step 101: Illuminate the optoelectronic device by outputting three characteristic band light sources through a ring LED light source array to obtain three-band image data. Process and fuse the three-band image data to obtain a three-band composite image. Specifically, in this embodiment, the optoelectronic device is illuminated by a ring-shaped LED light source array that outputs three characteristic wavelength light sources, including 450nm, 650nm, and 850nm. Image data of optoelectronic devices in various wavelengths are acquired using image sensors, including blue light images, red light images, and near-infrared light images; The Gaussian filtering algorithm was used to denoise the image data of optoelectronic devices, with the standard deviation set to 1.5, to obtain the denoised image data. After contrast enhancement of the denoised image data, a weighted fusion algorithm is used to assign weights based on the clarity of defect information in each band image to obtain a three-band composite image.

[0020] Specifically, (I) Design of Ring LED Light Source Array As a key illumination component for surface imaging of optoelectronic devices, the ring-shaped LED light source array adopts a multi-ring nested structure, with an inner ring diameter of 50mm and an outer ring diameter of 150mm, containing a total of 120 high-brightness LED beads. Its three characteristic wavelength light sources are precisely selected based on the reflection and absorption characteristics of different wavelengths of light by various defects on the surface of optoelectronic devices: 450nm (blue light), 660nm (red light), and 850nm (near-infrared light).

[0021] The 450nm blue light band can highlight surface scratches and other defects because the scratches significantly scatter blue light. The 660nm red light band can enhance the contrast of oxide spot defects, as the absorption of red light by the oxide area is significantly different from that of the normal area. The 850nm near-infrared light band can effectively penetrate the slight contamination layer on the surface of the device and clearly show internal crack defects.

[0022] (II) Image Data Acquisition Image acquisition was performed using a 12-megapixel industrial camera (4096×3072 resolution) and a telecentric lens (50mm focal length). The camera was coaxially positioned with the ring LED light source array, 300mm from the device surface, ensuring that the imaging field of view covered a 200mm diameter detection area. During the acquisition process, three characteristic wavelength light sources were sequentially illuminated individually, and three images were acquired for each wavelength to reduce noise interference. Clear image data for each wavelength, namely blue light image, red light image, and near-infrared light image, were obtained by averaging.

[0023] (III) Image Data Processing and Fusion The three-band image data were preprocessed, including denoising (using a Gaussian filtering algorithm with a standard deviation of 1.5) and contrast enhancement (using an adaptive histogram equalization method). Image fusion was then performed using a pixel-level weighted fusion algorithm, assigning weights based on the clarity of defect information in each band image (blue light image weight 0.3, red light image weight 0.3, near-infrared light image weight 0.4), ultimately yielding a composite three-band image.

[0024] Step 102: Construct a DFD-Net dual-branch deep network model. Input the three-band synthetic image into the model. Capture multi-scale defect geometric features through the ASPP module with increasing hole rate. Extract reflective insensitive texture features through phase-consistent convolutional layers. Use a gated cross-attention mechanism to fuse features and obtain a fused feature image. Specifically, in this embodiment, the ASPP module adopts a 4-way parallel dilated convolution structure. The three-band synthesized image is passed through a 64-channel convolutional layer for preliminary feature mapping to obtain a basic feature map. The basic feature map is then convolved to obtain the defect geometric features. Phase-consistent convolutional layers are used to process the phase and amplitude information of the image. The three-band composite image is decomposed into a Gaussian pyramid to generate an image pyramid with five scales. Phase calculation and amplitude suppression are performed at each scale to obtain reflectivity-insensitive texture features.

[0025] The multi-scale geometric feature map and the reflectivity-insensitive texture feature map are respectively converted into 128-channel feature vectors through convolutional layers, and the cosine similarity matrix between the two is calculated. Based on the similarity matrix, a spatial attention weight map is generated using the Softmax function, with weight values ​​ranging from 0 to 1. A learnable gating parameter vector is set, and the two types of features are dynamically weighted using the Sigmoid function to obtain a fused feature image.

[0026] Specifically, I. Model Input Preprocessing Before being input into the DFD-Net dual-branch deep network model, the three-band synthesized image needs to undergo standardization processing. Specifically, the image pixel values ​​are converted from a grayscale range of 0-255 to a normalized range of 0-1. This eliminates the influence of pixel value fluctuations under different lighting conditions, ensuring the model's adaptability to the input data. Simultaneously, data augmentation is performed using random horizontal flipping and random rotation within an angle range of -10° to 10° to improve the model's robustness to changes in device placement angles, providing more generalized input data for subsequent feature extraction. II. Dual-branch feature extraction process (a) Multi-scale defect geometric feature branch (ASPP module) The ASPP module, as the core component for capturing multi-scale defect geometric features, employs a 4-way parallel dilated convolutional structure. The input three-band synthesized image first undergoes preliminary feature mapping through a 64-channel 3×3 convolutional layer to obtain a basic feature map. Subsequently, this feature map is simultaneously fed into four convolutional paths with different dilation rates: The path with a hole rate of 1: using a 3×3 convolution kernel with 128 channels, it focuses on the edge contours of minute defects (pinholes with a diameter of 0.05mm-0.1mm) in the image without expanding the receptive field, and enhances local geometric features through two consecutive convolution operations. The path with a dilatation rate of 6: Employs a 3×3 convolutional kernel with 128 channels, expanding the receptive field to 51×51 pixels. This effectively captures the overall morphology of medium-scale defects (scratches 1mm-3mm in length). Padding is used during convolution to maintain the feature map size consistent with the input. The path with a dilatation rate of 12: Also uses a 3×3 convolutional kernel with 128 channels, expanding the receptive field to 99×99 pixels. It focuses on extracting the geometric distribution features of larger-scale defects (oxidation spots with an area of ​​5mm²-10mm²), and uses BatchNormalization layers to stabilize the training process. The path with a void ratio of 18: 128-channel 3×3 convolutional kernels form a receptive field of 147×147 pixels, which is specifically designed to capture features of ultra-large defects (cracks with a length of more than 5mm), and enhances nonlinear expression capabilities through the ReLU activation function. After the 4-channel output feature maps are spliced ​​(total number of channels is 128×4=512), they are compressed to 256 channels through a 1×1 convolutional layer to form a multi-scale geometric feature map. This feature map completely preserves the shape, size and location information of defects from the micrometer level to the millimeter level. (ii) Reflection-insensitive texture feature branch (phase-consistent convolutional layer) The phase-consistent convolutional layer employs a dual-channel parallel design, processing the phase and amplitude information of the image separately. The input three-band synthesized image is first decomposed into a Gaussian pyramid, generating image pyramids at five scales, each containing Gabor filter responses in eight directions. At each scale: Phase calculation channel: The phase angle of each pixel is extracted through Hilbert transform, and the phase consistency value is calculated by combining the phase difference between adjacent pixels. This value is between 0 and 1, and the higher the value, the greater the possibility of texture boundary at that location. For the highlight areas on the device surface caused by reflection, since their phase information is stable, the phase consistency value is not affected by the brightness change, and the defect texture under this area can be clearly preserved. Amplitude suppression channel: Adaptive threshold processing is applied to the amplitude response of Gabor filtering. When the amplitude exceeds 3 times the average value of the normal area (determined as a strong reflective area), the interference of the amplitude on the texture features is suppressed by an attenuation coefficient of 0.3, ensuring that the texture extraction focuses on the texture changes of the defect itself (stepped texture of crack edges, rough texture of oxide spot surface). After feature fusion at five scales, the phase-consistent convolutional layer outputs a 256-channel texture feature map. This feature map improves the signal-to-noise ratio in the reflective area by more than 40% compared to traditional convolutional layers, effectively distinguishing the texture differences between reflections and real defects. III. Feature Fusion of Gated Cross-Attention Mechanism The gated cross-attention mechanism achieves efficient fusion of two types of features through feature interaction and dynamic filtering, specifically in three stages: Feature interaction stage: The multi-scale geometric feature map and the reflectivity-insensitive texture feature map are converted into 128-channel feature vectors through 1×1 convolutional layers, and the cosine similarity matrix between them is calculated. Each element in the matrix represents the correlation strength between the geometric feature and the texture feature at the corresponding spatial location. For example, the similarity between the edge geometric feature and the texture feature of a scratch is usually higher than 0.7, while the similarity between the random noise region and the texture feature is lower than 0.2. Attention weight generation stage: Based on the similarity matrix, a spatial attention weight map is generated using the Softmax function, with weight values ​​ranging from 0 to 1. For high similarity regions (regions where the geometric contour of the crack coincides with the texture boundary), the weight value is close to 1, enhancing the contribution of the features in that region; for low similarity regions (geometric interference features of simple reflective areas), the weight value drops to below 0.3, weakening its influence. Gated fusion stage: A learnable gating parameter vector (128 dimensions) is set, and the two types of features are dynamically weighted using the Sigmoid function. When the gating parameter is greater than 0.5, geometric features are retained first; when it is less than 0.5, texture features are emphasized. Finally, a 256-channel fused feature image is obtained by element-wise addition. This image simultaneously contains the multi-scale geometric shape of the defect and anti-reflective texture details, providing comprehensive feature support for subsequent defect localization.

[0027] Step 103: Based on the improved YOLOv5 target detection algorithm, coarsely locate the fused feature image to obtain the defect region image. Use the Swing Transformer to calculate the pixel-level mask of the defect region image to obtain the segmented image. Specifically, in this embodiment, the fused feature image is input into the improved YOLOv5 algorithm, the image is divided into grids, and combined with the improved feature fusion structure, a prediction result containing the defect location and category probability is generated; The confidence threshold is set to 0.65. When the confidence of the prediction result is higher than this threshold, it is determined to be a defective area.

[0028] The image is divided into multiple non-overlapping image blocks, each of which is 16×16 pixels in size; The degree of correlation between each image patch is calculated through a self-attention mechanism, the features of the image patch are mapped back to the pixel level, and each pixel is assigned a probability value of a defect region. When the probability value is greater than 0.5, the pixel is determined to be a defective pixel; otherwise, it is a normal pixel. A pixel-level mask is formed to obtain a segmented image.

[0029] Specifically, (I) Details of YOLOv5 Algorithm Improvements To improve the coarse localization accuracy of defect regions in fused feature images, targeted improvements were made to the YOLOv5 algorithm. A lightweight attention module was introduced into the backbone network. This module evaluates the importance of feature channels and assigns higher weights to feature channels containing defect information, enhancing the network's sensitivity to defect features while avoiding localization errors caused by feature redundancy.

[0030] In the Neck section, the original feature fusion structure was adjusted to a cross-scale feature interaction structure, enabling features at different levels to exchange information more fully. Low-level features retain rich detail information, which helps to capture the location of small-sized defects; high-level features have stronger semantic information, which can help identify the overall range of large-sized defects. The effective interaction between the two improves the adaptability to defects of different scales.

[0031] In addition, the candidate box generation strategy was optimized. Based on the size distribution of common defects in optoelectronic devices (scratches are mostly between 0.5mm and 10mm in length and oxide spots are mostly between 0.3mm and 5mm in diameter), the size and proportion of the candidate boxes were redesigned to reduce the positioning error caused by the mismatch between the candidate boxes and the actual defect size.

[0032] (II) Coarse Positioning Implementation Process The fused feature image is input into the improved YOLOv5 algorithm. First, the image is divided into a grid, with each grid responsible for predicting potential defects within its coverage area. The network then progressively extracts deep features from the image through multiple convolutional and pooling operations, and combines this with the improved feature fusion structure to generate prediction results that include defect location and category probability.

[0033] A confidence threshold of 0.65 is set. When the confidence level of the prediction result is higher than this threshold, it is identified as a possible defect region. After screening, multiple candidate defect regions are obtained, which are presented in the form of rectangular boxes, representing the initial range of the defect region image. At this stage, the defect region image may contain some normal areas, but it can roughly define the location of the defect, laying the foundation for subsequent accurate segmentation.

[0034] II. Calculating pixel-level masks using Swin Transformer (a) Image preprocessing of defect areas Before performing pixel-level mask calculations using the Swing Transformer, the defect region image is preprocessed. The defect region image is uniformly adjusted to a fixed size (224×224 pixels). A scaling operation is used to ensure that the image size meets the input requirements of the Swing Transformer, while maintaining the shape and relative position of the defect.

[0035] The adjusted image is standardized to eliminate the impact of differences in pixel values ​​in different regions, making the pixel value distribution of the image more consistent with the training data distribution of the model, thereby improving the model's ability to identify defective region features.

[0036] (II) Pixel-level mask calculation process The Swin Transformer processes defect region images in a layered manner. First, the image is divided into multiple non-overlapping image patches, each 16×16 pixels in size, which serve as input tokens for the Transformer.

[0037] During processing, a self-attention mechanism is used to calculate the correlation between each image patch and other image patches. For defective regions, the features of pixels within them are highly consistent, and the correlation between image patches is high; while at the boundary between defective and normal regions, the features of image patches differ significantly, and the correlation is low.

[0038] Through iterative computation using a multi-layer Transformer encoder, the model gradually learns the pixel distribution patterns in defective regions. In the decoder, the features of image blocks are mapped back to the pixel level, assigning each pixel a probability value indicating it belongs to a defective region. When the probability value is greater than 0.5, the pixel is determined to be a defective pixel; otherwise, it is considered a normal pixel. This process ultimately forms a pixel-level mask, resulting in a segmented image. The segmented image clearly displays the specific shape and extent of the defective region, significantly improving accuracy compared to the rectangular bounding boxes obtained through coarse localization.

[0039] Step 104: Locate the defect boundaries of the segmented image using the Non-Maximum Suppression (NMS) algorithm, and output the defect coordinates and size information.

[0040] Specifically, in this embodiment, all edge pixels of the defect region in the binarized image are extracted, and the gradient value of each edge pixel is calculated. The edge pixels are sorted from high to low according to the gradient value, the defect boundary of the segmented image is located, and the defect coordinates and size information are output.

[0041] Specifically, (a) Image segmentation preprocessing Before inputting the segmented image into the Non-Maximum Suppression (NMS) algorithm, it needs to be binarized. Based on the probability values ​​of the pixel-level mask, the segmented image is converted into a black-and-white binary image, where the pixel value of the defective region is set to 255 (white) and the pixel value of the normal region is set to 0 (black). This step can eliminate grayscale interference, making the boundary between the defective region and the normal region clearer, and providing a clear image basis for subsequent boundary localization.

[0042] Simultaneously, morphological operations are performed on the binarized image, using a 3×3 structuring element for erosion and dilation. Erosion removes burrs and isolated noise points from the edges of defect areas, preventing these interfering factors from being misjudged as defect boundaries; dilation fills in small voids within the defect area, making the overall shape of the defect area more complete and ensuring accurate boundary positioning.

[0043] (II) Implementation process of NMS algorithm The Non-Maximum Suppression (NMS) algorithm is used to accurately locate defect boundaries from segmented images. The specific process is as follows: First, all edge pixels of the defect region in the binarized image are extracted, and the gradient value of each edge pixel is calculated. The gradient value reflects the rate of change of a pixel in the image. The pixel gradient value at the defect boundary is significantly higher than that of the inner and outer regions, which can serve as an important basis for judging boundary points.

[0044] Then, the edge pixels are sorted from highest to lowest gradient value, and the point with the highest gradient value is used as the initial boundary point. A neighborhood of a certain size (5×5 pixels) is set with this initial point as the center. If there are other edge pixels in the neighborhood, their gradient values ​​are compared, and the point with the highest gradient value is retained, while the other points are suppressed.

[0045] Repeat the above process, traversing all edge pixels, until all non-maximum edge points are suppressed. The remaining edge points then constitute the accurate boundary of the defect. This process effectively removes redundant edge points, ensuring that the obtained defect boundary is continuous, complete, and highly consistent with the actual defect contour.

[0046] II. Output of Defect Coordinates and Dimensions (a) Determination of defect coordinates Establish a two-dimensional Cartesian coordinate system with the top left corner of the segmented image as the origin, with the horizontal direction as the X-axis and the vertical direction as the Y-axis. Traverse all pixels on the defect boundary and record the coordinates (x, y) of each point.

[0047] For the defect boundary, the X-coordinate of the leftmost pixel is taken as the minimum X-coordinate of the defect, and the X-coordinate of the rightmost pixel is taken as the maximum X-coordinate; the Y-coordinate of the topmost pixel is taken as the minimum Y-coordinate, and the Y-coordinate of the bottommost pixel is taken as the maximum Y-coordinate. Using these four coordinate values, the location range of the defect in the image can be determined.

[0048] (ii) Defect size calculation The size information of the defect is calculated based on its minimum and maximum coordinate values. The length of the defect is the difference between the maximum and minimum X coordinates, and the width is the difference between the maximum and minimum Y coordinates. For irregularly shaped defects, in addition to length and width, the area of ​​the defect can also be calculated, which is the number of all pixels in the defect area multiplied by the actual physical size of a single pixel (pre-calibrated based on image resolution and camera focal length, e.g., 1 pixel corresponds to 0.01mm).

[0049] (III) Information Output Format The defect coordinates (minimum X, maximum X, minimum Y, maximum Y) and dimensional information (length, width, area) are organized according to a preset format and output as a text file or table. The output information may also include the defect category (combined with previous inspection results) and confidence level, providing comprehensive and accurate data support for the quality assessment and subsequent processing of optoelectronic devices. Simultaneously, the defect coordinates and boundaries are marked on the original image to generate a visual report, allowing staff to intuitively understand the specific location and shape of the defect.

[0050] Its beneficial effects are as follows: 1. The entire inspection process is automated. From image acquisition, feature extraction, defect localization to information output, no manual intervention is required, significantly improving inspection speed. Compared to manual inspection, it can adapt to the rapid inspection needs of large-scale production lines, shortening the product inspection cycle and improving inspection and production efficiency. 2. The inspection system can operate stably under different lighting conditions. Simultaneously, the multi-scale feature extraction and fusion mechanism allows the system to handle defects of different sizes and types, expanding the inspection range and improving its adaptability to complex inspection scenarios. 3. It avoids the judgment differences caused by human factors in manual inspection, resulting in more accurate inspection results. The output defect coordinates, size, category, and other information provide a quantitative basis for product quality assessment.

[0051] Please see Figure 2 In the visual inspection method for surface defects of optoelectronic devices based on image recognition, the optoelectronic device is illuminated by a ring-shaped LED light source array outputting three characteristic band light sources to obtain three-band image data. The three-band image data are then processed and fused to obtain a three-band composite image, including the following steps: Step 201: Irradiate the optoelectronic device by outputting three characteristic wavelength light sources through a ring LED light source array. The three characteristic wavelengths include 450nm, 650nm, and 850nm. Step 202: Acquire image data of optoelectronic devices in various wavelength bands using an image sensor, including blue light images, red light images, and near-infrared light images; Step 203: Use the Gaussian filtering algorithm to denoise the optoelectronic device image data, with the standard deviation set to 1.5, to obtain the denoised image data; Step 204: After enhancing the contrast of the denoised image data, a weighted fusion algorithm is used to assign weights based on the clarity of the defect information in each band image to obtain a three-band composite image.

[0052] The above describes embodiments of the visual inspection method for surface defects in optoelectronic devices based on image recognition according to the present invention. Please refer to [link / reference]. Figure 3 In the image recognition-based visual inspection system for surface defects in optoelectronic devices, the system includes the following modules: The light source band adjustment module is used to illuminate the optoelectronic device by outputting three characteristic band light sources through a ring LED light source array to obtain three-band image data. The three-band image data is then processed and fused to obtain a three-band composite image. The feature image fusion module is used to construct the DFD-Net dual-branch deep network model. The three-band synthetic image is input into the model. The multi-scale defect geometric features are captured by the ASPP module with increasing hole rate. The reflective insensitive texture features are extracted by the phase-consistent convolutional layer. The feature fusion is performed by the gated cross attention mechanism to obtain the fused feature image. The feature image segmentation module is used to coarsely locate the fused feature image based on the improved YOLOv5 target detection algorithm to obtain the defect region image. The Swing Transformer is used to calculate the pixel-level mask of the defect region image to obtain the segmented image. The defect localization calculation module is used to locate the defect boundaries of the segmented image using the NMS non-maximum suppression algorithm, and output the defect coordinates and size information.

[0053] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A visual inspection method for surface defects in optoelectronic devices based on image recognition, characterized in that, The visual inspection method for surface defects in optoelectronic devices includes the following steps: Three characteristic band light sources are output from a ring-shaped LED light source array to illuminate the optoelectronic device, thereby obtaining three-band image data. The three-band image data are then processed and fused to obtain a three-band composite image. A DFD-Net dual-branch deep network model is constructed. The three-band synthetic image is input into the model. Multi-scale defect geometric features are captured by the ASPP module with increasing hole rate. Reflection-insensitive texture features are extracted by the phase-consistent convolutional layer. Feature fusion is performed by the gated cross-attention mechanism to obtain the fused feature image. The fused feature image is coarsely localized based on the improved YOLOv5 target detection algorithm to obtain the defect region image. The pixel-level mask of the defect region image is calculated using the Swing Transformer to obtain the segmented image. The defect boundaries of the segmented image are located using the Non-Maximum Suppression (NMS) algorithm, and the defect coordinates and size information are output.

2. The visual inspection method for surface defects of optoelectronic devices based on image recognition as described in claim 1, characterized in that, The process involves illuminating the optoelectronic device with three characteristic wavelength light sources output from a ring-shaped LED light source array to obtain three-band image data. The three-band image data are then processed and fused to obtain a three-band composite image, including: The optoelectronic device is illuminated by a ring-shaped LED light source array that outputs three characteristic wavelengths: 450nm, 650nm, and 850nm. Image data of optoelectronic devices in various wavelengths are acquired using image sensors, including blue light images, red light images, and near-infrared light images; The image data of the optoelectronic device was denoised using a Gaussian filtering algorithm with a standard deviation of 1.5 to obtain denoised image data. After contrast enhancement of the denoised image data, a weighted fusion algorithm is used to assign weights based on the clarity of defect information in each band image to obtain a three-band composite image.

3. The visual inspection method for surface defects of optoelectronic devices based on image recognition as described in claim 1, characterized in that, The construction of the DFD-Net dual-branch deep network model involves inputting the three-band synthesized image into the model. Multi-scale defect geometric features are captured through the ASPP module with increasing hole rate, and reflectivity-insensitive texture features are extracted through phase-consistent convolutional layers. This includes: The ASPP module uses a 4-way parallel dilated convolution structure to perform preliminary feature mapping on the three-band composite image through a 64-channel convolutional layer to obtain a basic feature map. The basic feature map is then convolved to obtain the defect geometric features. Phase-consistent convolutional layers are used to process the phase and amplitude information of the image. The three-band composite image is decomposed into a Gaussian pyramid to generate an image pyramid with five scales. Phase calculation and amplitude suppression are performed at each scale to obtain reflectivity-insensitive texture features.

4. The visual inspection method for surface defects of optoelectronic devices based on image recognition as described in claim 1, characterized in that, The feature fusion using a gated cross-attention mechanism to obtain a fused feature image includes: The multi-scale geometric feature map and the reflectivity-insensitive texture feature map are respectively converted into 128-channel feature vectors through convolutional layers, and the cosine similarity matrix between the two is calculated. Based on the similarity matrix, a spatial attention weight map is generated using the Softmax function, with weight values ​​ranging from 0 to 1. A learnable gating parameter vector is set, and the two types of features are dynamically weighted using the Sigmoid function to obtain a fused feature image.

5. The visual inspection method for surface defects of optoelectronic devices based on image recognition as described in claim 1, characterized in that, The step of calculating the pixel-level mask of the defect region image using the Swing Transformer to obtain the segmented image includes: The fused feature image is input into the improved YOLOv5 algorithm, which divides the image into grids and combines the improved feature fusion structure to generate prediction results containing defect location and category probability. The confidence threshold is set to 0.

65. When the confidence of the prediction result is higher than this threshold, it is determined to be a defective area.

6. The visual inspection method for surface defects of optoelectronic devices based on image recognition as described in claim 1, characterized in that, The step of calculating the pixel-level mask of the defect region image using the Swing Transformer to obtain the segmented image includes: The image is divided into multiple non-overlapping image blocks, each of which is 16×16 pixels in size; The degree of correlation between each image patch is calculated through a self-attention mechanism, the features of the image patch are mapped back to the pixel level, and each pixel is assigned a probability value of a defect region. When the probability value is greater than 0.5, the pixel is determined to be a defective pixel; otherwise, it is a normal pixel. A pixel-level mask is formed to obtain a segmented image.

7. The visual inspection method for surface defects of optoelectronic devices based on image recognition as described in claim 1, characterized in that, The step of locating the defect boundary of the segmented image using the Non-Maximum Suppression (NMS) algorithm and outputting the defect coordinates and size information includes: Extract all edge pixels of the defect region in the binarized image, calculate the gradient value of each edge pixel, sort the edge pixels from high to low according to the gradient value, locate the defect boundary of the segmented image, and output the defect coordinates and size information.

8. A visual inspection system for surface defects in optoelectronic devices based on image recognition, characterized in that, The visual inspection system for surface defects in optoelectronic devices includes the following modules: The light source band adjustment module is used to illuminate the optoelectronic device by outputting three characteristic band light sources through a ring LED light source array to obtain three-band image data. The three-band image data is then processed and fused to obtain a three-band composite image. The feature image fusion module is used to construct the DFD-Net dual-branch deep network model. The three-band synthesized image is input into the model. The multi-scale defect geometric features are captured by the ASPP module with increasing hole rate. The reflective insensitive texture features are extracted by the phase-consistent convolutional layer. The feature fusion is performed by the gated cross attention mechanism to obtain the fused feature image. The feature image segmentation module is used to coarsely locate the fused feature image based on the improved YOLOv5 target detection algorithm to obtain the defect region image, and to calculate the pixel-level mask of the defect region image using the Swing Transformer to obtain the segmented image. The defect localization calculation module is used to locate the defect boundaries of the segmented image using the NMS non-maximum suppression algorithm, and output the defect coordinates and size information.

9. The visual inspection system for surface defects of optoelectronic devices based on image recognition as described in claim 8, characterized in that, The feature image segmentation module includes the following sub-modules: The partitioning submodule is used to input the fused feature image into the improved YOLOv5 algorithm, divide the image into grids, and combine it with the improved feature fusion structure to generate prediction results containing defect location and category probability. The confidence level submodule is used to set the confidence level threshold to 0.

65. When the confidence level of the prediction result is higher than this threshold, it is determined to be a defect area.

10. The visual inspection system for surface defects of optoelectronic devices based on image recognition as described in claim 8, characterized in that, The defect location calculation module includes the following sub-modules: The localization submodule is used to extract all edge pixels of the defect region in the binarized image and calculate the gradient value of each edge pixel. The edge pixels are sorted from high to low according to their gradient values ​​to locate the defect boundaries of the segmented image and output the defect coordinates and size information.

Citation Information

Cited By

  • Paper roll raised strip detection system and method

    CN121259305A

  • Automobile part detection method and system based on machine vision

    CN121298745A

  • A machine vision-based automobile accessory detection method and system

    CN121298745B

  • Cable defect monitoring and risk grading method based on multi-modal deep fusion

    CN121350851A

  • Image recognition system for defect detection of industrial parts

    CN121582224A