VCSEL defect detection method and system based on multi-modal semantic segmentation model
Through the VCSEL defect detection method based on the multimodal semantic segmentation model, the problems of low detection efficiency and poor accuracy of VCSEL chips in the prior art are solved, and high-precision defect detection and positioning of VCSEL chips are realized, and detection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510660150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing VCSEL chip defect detection methods are inefficient and have poor accuracy, making it difficult to fully reflect the chip's defect damage information, resulting in limited detection accuracy.
The VCSEL defect detection method based on the multimodal semantic segmentation model is adopted, and the VCSEL multimodal semantic segmentation model is designed by obtaining the multimodal defect data set. The pre-trained convolution neural network ResNet and SE attention module are used to combine the hollow convolution and depth separation convolution to realize the identification and positioning of the multi-class defects of the VCSEL chip.
It realizes high-precision defect detection of VCSEL chips, and can detect surface and internal defects of the chip at the same time, improves detection efficiency and accuracy, and reduces production costs and detection time.
Smart Images

Figure CN120182274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of reliability analysis of vertical cavity surface emitting laser (VCSEL) chips, and specifically to a VCSEL defect detection method and system based on a multi-modal semantic segmentation model. Background Art
[0002] As a high-performance semiconductor laser, the vertical cavity surface emitting laser (VCSEL) has been widely used in the fields of optical communication, 3D sensing, lidar, etc. due to its advantages such as low power consumption, high modulation rate, and easy two-dimensional integration. However, during the manufacturing and use of VCSEL chips, due to factors such as material defects, process fluctuations, and thermal stress, damages such as electrode detachment, oxide layer damage, and active region defects are likely to occur, seriously affecting its performance and service life.
[0003] Traditional VCSEL chip defect damage detection methods mainly rely on electrical tests and optical microscope observations. The electrical test method judges its performance by measuring parameters such as the I-V characteristics and L-I characteristics of the VCSEL chip, but it is difficult to accurately locate the defect position and type. Although the optical microscope observation method can directly observe the surface morphology of the VCSEL, it is difficult to detect internal defects, and it depends on manual experience, with low efficiency and is easily affected by subjective factors.
[0004] In recent years, with the rapid development of artificial intelligence technology, image recognition technology based on deep learning has shown great potential in the field of industrial defect detection. However, most of the existing VCSEL chip defect detection methods based on deep learning use single-modal data (such as optical images or electrical signals), which are difficult to comprehensively reflect the defect damage information of VCSEL chips, resulting in limited detection accuracy. Summary of the Invention
[0005] In order to solve the technical problems of low detection efficiency and poor detection accuracy of VCSEL chips caused by the current cumbersome defect detection process and backward detection technology, the present invention provides a VCSEL defect detection method and system based on a multi-modal semantic segmentation model.
[0006] The present invention is implemented by adopting the following technical solutions: A VCSEL defect detection method based on a multi-modal semantic segmentation model includes the following steps: S1. Obtain a multi-modal defect data set; S11. Collect visible light images and infrared images of multiple VCSEL sample chips; The collected VCSEL sample chips include VCSEL chips without defects and VCSEL chips with defects; S12. Respectively fuse the visible light image and the infrared image of each VCSEL sample chip through a multimodal image fusion algorithm to obtain the multimodal defect fusion image of each VCSEL sample chip; S13. Mark the multimodal defect fusion images of the VCSEL sample chips obtained in step S12, eliminate the multimodal defect fusion images without defects, and mark the defect categories, defect positions, and defect morphologies of the defective multimodal defect fusion images. The marked multimodal defect fusion images are used as chip defect samples to form a multimodal defect dataset; S2. Design of the VCSEL multimodal semantic segmentation model; S21. Use the pre-trained convolutional neural network ResNet as the encoder, which includes multiple layers of ordinary convolutions, and an SE attention module is added after each layer of ordinary convolution; S22. Use dilated convolutions with the same number and one-to-one correspondence as the number of encoder layers to form the decoder; Except for the last layer, each layer of ordinary convolution in the encoder is skip-connected to the dilated convolution of the corresponding layer in the decoder; The last layer of ordinary convolution in the encoder is connected to the last layer of dilated convolution in the decoder through depthwise separable convolution to form the VCSEL multimodal semantic segmentation model; S3. Training of the VCSEL multimodal semantic segmentation model Input the chip defect samples of the multimodal defect dataset obtained in S1 into the VCSEL multimodal semantic segmentation model to train the VCSEL multimodal semantic segmentation model. The VCSEL multimodal semantic segmentation model can output the multimodal defect segmentation result map of each chip defect sample, and the multimodal defect segmentation result map contains the defect category, defect position, and defect morphology of the chip defect sample; When the defect category, defect position, and defect morphology in the multimodal defect segmentation result map output by the VCSEL multimodal semantic segmentation model are compared with the chip defect sample and the accuracy reaches more than 95%, it is regarded as the completion of the training of the VCSEL multimodal semantic segmentation model; S4. Detection of the VCSEL multimodal semantic segmentation model Collect the visible light image and the infrared image of the VCSEL chip to be tested, obtain the multimodal defect fusion image of the VCSEL chip to be tested through the multimodal image fusion algorithm, and then input the multimodal defect fusion image of the VCSEL chip to be tested into the VCSEL multimodal semantic segmentation model completed in training in S3, and the VCSEL multimodal semantic segmentation model outputs a multimodal defect segmentation result map containing the defect category, defect position, and defect morphology of the VCSEL chip to be tested.
[0007] Further, the defect categories include surface contamination, mechanical scratches, epitaxial defects, and dark defects in the light-emitting holes.
[0008] Further, in step S3, before training, the following training parameters and configurations are set for the VCSEL multi-modal semantic segmentation model: the initial learning rate is set to 1e -4 , and the cosine annealing learning rate scheduler is used to dynamically adjust the learning rate; the AdamW optimizer is used, combined with weight decay to prevent overfitting; the batch size is set to 16, and the number of training epochs is 150; after training is completed, the model weights are saved.
[0009] A VCSEL defect detection system based on a multi-modal semantic segmentation model is used to implement a VCSEL defect detection method based on a multi-modal semantic segmentation model, including a probe station for placing VCSEL sample chips, a conveyor belt for placing VCSEL chips to be tested, an image acquisition system for collecting images of VCSEL sample chips and VCSEL chips to be tested, and a computer; the image acquisition system includes an infrared CCD for collecting infrared images of VCSEL sample chips and VCSEL chips to be tested, a visible light CCD for collecting visible light images of VCSEL sample chips and VCSEL chips to be tested, a visible light source and an infrared light source for providing light sources for VCSEL sample chips and VCSEL chips to be tested; the signal output ends of the infrared CCD and the visible light CCD are connected to the computer, and the probe station, the conveyor belt, the visible light source, and the infrared light source are all electrically connected to the computer and controlled by the computer to act; the computer is built-in with a multi-modal image fusion algorithm and a VCSEL multi-modal semantic segmentation model.
[0010] Further, before the defect detection system performs detection, first set the running speed of the conveyor belt, the intensities of the visible light source and the infrared light source, the shooting angles of the infrared CCD and the visible light CCD, and assign an identification number to each VCSEL chip to be tested; then start automatic detection. The multi-modal image fusion algorithm built into the computer fuses the visible light image and the infrared image of the VCSEL chip to be tested collected, and after obtaining the fused image of the VCSEL chip to be tested, the trained VCSEL multi-modal semantic segmentation model built into the computer detects the fused image of the VCSEL chip to be tested. If there are no defects in the fused image of the VCSEL chip to be tested, the computer outputs the identification number of the VCSEL chip to be tested and marks it as qualified; if there are defects in the fused image of the VCSEL chip to be tested, the computer outputs the identification number, defect category, defect location, and defect form of the VCSEL chip to be tested and marks it as unqualified.
[0011] In view of the multi-modal defect image features of VCSEL, a multi-modal semantic segmentation model for VCSEL is proposed. An innovative feature extraction convolutional module is designed, combined with the SE attention mechanism. At the same time, multi-modal data (visible light defect images and infrared light defect images) are used to train the semantic segmentation model to achieve accurate recognition and positioning of various types of VCSEL defects, and for the first time, simultaneous detection of internal and external defects of VCSEL is realized, ultimately improving the detection efficiency and accuracy. Aiming at the problems of low efficiency and high cost of repeated screening in the traditional detection process, the defect detection system described in the present invention effectively improves the automation degree of the detection process, significantly reduces the production cost and detection time, and provides an efficient and low-cost solution for the defect detection of laser chips. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 Schematic diagram of the structure of the VCSEL defect detection system based on the multi-modal semantic segmentation model of the present invention Figure 1 。
[0013] Figure 2 Schematic diagram of the structure of the VCSEL defect detection system based on the multi-modal semantic segmentation model of the present invention Figure 2 。
[0014] Figure 3 Schematic diagram of the process for obtaining the multi-modal defect data set of the present invention.
[0015] Figure 4 Schematic diagrams of the defects of various VCSEL chips in the multi-modal defect data set.
[0016] Figure 5 Schematic diagram of the structure of the VCSEL multi-modal semantic segmentation model.
[0017] Figure 6 Flow chart of the operation of the defect detection system.
[0018] Figure 7 Experimental comparison diagram for the production of the multi-modal defect data set in step S1 of the present invention.
[0019] Figure 8 Schematic diagram of the multi-modal defect detection results in step S4 of the present invention.
[0020] 1 - VCSEL sample chip, 2 - probe station, 3 - VCSEL chip to be tested, 4 - conveyor belt, 5 - computer, 6 - infrared CCD, 7 - visible light CCD, 8 - surface defect, 9 - internal defect, 10 - light output hole. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] A VCSEL defect detection method based on a multi-modal semantic segmentation model includes the following steps: S1. Obtain a multi-modal defect dataset; S11. Collect visible light images and infrared images of multiple VCSEL sample chips 1; The collected VCSEL sample chips 1 include VCSEL chips without defects and VCSEL chips with defects; S12. Respectively fuse the visible light image and the infrared image of each VCSEL sample chip 1 through a multi-modal image fusion algorithm to obtain a multi-modal defect fusion image of each VCSEL sample chip 1; S13. Mark the multi-modal defect fusion images of the VCSEL sample chips 1 obtained in step S12, eliminate the multi-modal defect fusion images without defects, and mark the defect categories, defect positions, and defect morphologies of the multi-modal defect fusion images with defects. The marked multi-modal defect fusion images are used as chip defect samples to form a multi-modal defect dataset; The defect categories include surface contamination, mechanical scratches, epitaxial defects, and dark defects in the light-emitting holes; S2. Design of the VCSEL multi-modal semantic segmentation model; S21. Use the pre-trained convolutional neural network ResNet as the encoder, and this encoder includes multiple layers of ordinary convolutions, and an SE attention module is added after each layer of ordinary convolution; S22. Use dilated convolutions that are the same in number and one-to-one correspondence with the number of encoder layers to form the decoder; Except for the last layer, each layer of ordinary convolution in the encoder is skip-connected to the dilated convolution of the corresponding layer in the decoder; The last layer of ordinary convolution in the encoder is connected to the last layer of dilated convolution in the decoder through depthwise separable convolution to form the VCSEL multi-modal semantic segmentation model; S3. Training of the VCSEL multi-modal semantic segmentation model Input the chip defect samples of the multi-modal defect dataset obtained in S1 into the VCSEL multi-modal semantic segmentation model to train the VCSEL multi-modal semantic segmentation model. The VCSEL multi-modal semantic segmentation model can output a multi-modal defect segmentation result map of each chip defect sample. The multi-modal defect segmentation result map contains the defect category, defect position, and defect morphology of the chip defect sample; When the defect category, defect position, and defect morphology in the multi-modal defect segmentation result map output by the VCSEL multi-modal semantic segmentation model are compared with the chip defect sample, and the precision reaches more than 95%, it is regarded as the VCSEL multi-modal semantic segmentation model completing the training; The following training parameters and configurations are set for the VCSEL multi-modal semantic segmentation model before training: The initial learning rate is set to 1e -4 , and use a cosine annealing learning rate scheduler to dynamically adjust the learning rate; Use the AdamW optimizer and combine weight decay to prevent overfitting; The batch size is set to 16, and the number of training epochs is 150; Save the model weights after training; S4. Detection of VCSEL Multimodal Semantic Segmentation Model Collect the visible light image and infrared image of the VCSEL chip 3 to be tested. Obtain the multimodal defect fusion image of the VCSEL chip 3 to be tested through the multimodal image fusion algorithm. Then input the multimodal defect fusion image of the VCSEL chip 3 to be tested into the VCSEL multimodal semantic segmentation model completed in S3, and the VCSEL multimodal semantic segmentation model outputs a multimodal defect segmentation result map including the defect category, defect location and defect morphology of the VCSEL chip 3 to be tested. The multimodal image fusion algorithm adopts but is not limited to the method disclosed in Chinese Patent Publication No. CN117830235A.
[0022] Currently, the quality screening links on the chip production line are all independently separated. First, surface defects are detected, and then internal defects are detected, which is very time-consuming. At the same time, the detection method is also manual detection, so the misdetection rate will be high. The innovation of the present invention lies in being able to detect the surface defects and internal defects of the chip simultaneously, greatly improving the detection efficiency and accuracy; the significance of the VCSEL multimodal semantic segmentation model designed by the present invention is: one is that it can greatly simplify the detection process on the chip production line, realize "one-step" detection, and reduce the detection efficiency; the other is that the detection of the algorithm can avoid the problem of high misdetection rate caused by manual detection.
[0023] A VCSEL defect detection system based on a multimodal semantic segmentation model includes a probe station 2 for placing the VCSEL sample chip 1, a conveyor belt 4 for placing the VCSEL chip 3 to be tested, an image acquisition system for collecting images of the VCSEL sample chip 1 and the VCSEL chip 3 to be tested, and a computer 5; the image acquisition system includes an infrared CCD 6 for collecting infrared images of the VCSEL sample chip 1 and the VCSEL chip 3 to be tested, a visible light CCD 7 for collecting visible light images of the VCSEL sample chip 1 and the VCSEL chip 3 to be tested, a visible light source and an infrared light source for providing light sources for the VCSEL sample chip 1 and the VCSEL chip 3 to be tested; the signal output ends of the infrared CCD 6 and the visible light CCD 7 are connected to the computer 5, and the probe station 2, the conveyor belt 4, the visible light source, and the infrared light source are all electrically connected to the computer 5 and controlled by the computer 5 to act (connected through an HDMI data cable); the computer 5 internally stores a multimodal image fusion algorithm and a VCSEL multimodal semantic segmentation model.
[0024] When obtaining a multi-modal defect data set, the VCSEL sample chip 1 is placed on the probe station 2, the light source is adjusted, and then the infrared image and visible light image of the VCSEL sample chip 1 are collected through the infrared CCD 6 and the visible light CCD 7. The probe station 2, the conveyor belt 4, the visible light source, and the infrared light source are all electrically connected to the computer 5 and the actions are controlled by the computer 5, which means that the actions of the probe station 2 and the conveyor belt 4 (such as the lifting of the probe station 2, the running direction and speed of the conveyor belt 4) and the intensities of the visible light source and the infrared light source can be controlled through the computer; the visible light source and the infrared light source are supported by corresponding brackets and the angles and heights are adjusted. Of course, according to the needs of actual applications, the angles of the visible light source and the infrared light source can also be adjusted through the computer 5, and all of these are easily achieved by those skilled in the art.
[0025] Before the defect detection system detects, the infrared CCD 6 and the visible light CCD 7 are placed above the conveyor belt, and then the running speed of the conveyor belt 4, the intensities of the visible light source and the infrared light source, and the shooting angles of the infrared CCD 6 and the visible light CCD 7 are set, and an identification number is assigned to each VCSEL chip 3 to be tested; then the automatic detection starts. The multi-modal image fusion algorithm built in the computer 5 fuses the visible light image and the infrared image of the VCSEL chip 3 to be tested collected, and after obtaining the multi-modal defect fusion image of the VCSEL chip 3 to be tested, the trained VCSEL multi-modal semantic segmentation model built in the computer 5 detects the multi-modal defect fusion image of the VCSEL chip 3 to be tested. If there is no defect in the multi-modal defect fusion image of the VCSEL chip 3 to be tested, the computer 5 outputs the identification number of the VCSEL chip 3 to be tested and marks it as qualified; if there is a defect in the multi-modal defect fusion image of the VCSEL chip 3 to be tested, the computer outputs the identification number, defect category, defect location, and defect form of the VCSEL chip 3 to be tested and marks it as unqualified.
[0026] Due to the complexity of the VCSEL chip production process, various defects often occur, but the defect recognition rate is low and the scale is small. Traditional detection methods are difficult to meet the requirements for accuracy and detection speed. The present invention improves the U-Net network, and then deploys the algorithm to the computer, which is connected to related devices such as the light source, the infrared CCD 6, the conveyor belt 4, and the probe station 2 to form a complete VCSEL defect detection system based on a multi-modal semantic segmentation model.
[0027] The following further describes the present invention with reference to the accompanying drawings.
[0028] 1. Preparation of multi-modal defect data set As Figure 1As shown in the figure, in the defect detection system of the present invention, the probe station 2, the infrared CCD 6, the visible light CCD 7, the visible light source and the infrared light source, and the computer 5 constitute a VCSEL multi-modal defect acquisition system for collecting visible light images and infrared images of a plurality of VCSEL sample chips 1. During collection, the VCSEL sample chip 1 is placed on the probe station 2, and the visible light source and the infrared light source irradiate the VCSEL sample chip 1 respectively. At the same time, the visible light CCD 7 and the infrared CCD 6 collect the visible light image and the infrared image of the VCSEL sample chip 1 respectively.
[0029] As Figure 3 shown, after obtaining the visible light image and infrared image datasets of a large number of VCSEL sample chips 1, a multi-modal defect dataset required by the present invention is obtained by using a multi-modal image fusion algorithm. The defect types mainly include 4 categories: surface contamination, mechanical scratches, epitaxial defects, and dark defects in the light-emitting holes, as Figure 4 shown.
[0030] 2. Design and Training of VCSEL Multi-modal Semantic Segmentation Model As Figure 5 shown, an improved U-Net model is adopted, and the following optimizations are introduced on the basis of the original U-Net: (1) Encoder part (Encoder) The pre-trained convolutional neural network ResNet is used as the encoder to replace the simple convolutional layer of the original U-Net. It is convenient to extract multi-level feature information, from low-level features (defect edges, defect textures) to high-level semantic features (defect shapes). An SE (Squeeze-and-Excitation) attention module is added after each layer of the encoder to enhance the weight of tiny defect features.
[0031] (2) Decoder part (Decoder) At each layer, dilated convolution is used to replace ordinary convolution to expand the receptive field and capture more context information. The features of each layer of the encoder are directly connected to the corresponding layer of the decoder to enhance the reuse of defect features. At the same time, the multi-scale features extracted by the encoder are fused to achieve accurate multi-modal defect pixel-level segmentation.
[0032] At the last layer, depthwise separable convolution is used to reduce the number of parameters, optimize the calculation efficiency, improve the inference speed, and finally output a high-resolution defect segmentation result.
[0033] (3) Model Hyperparameter Setting and Training Relevant parameters for training: The initial learning rate is set to 1e -4, and use the CosineAnnealing LR Scheduler to dynamically adjust the learning rate to avoid getting stuck in local optima; use the AdamW optimizer and combine it with Weight Decay to prevent overfitting; set the batch size to 16 and the number of training epochs to 150. After training is completed, save the model weights and conduct tests on the VCSEL multi-modal semantic segmentation model.
[0034] 3. VCSEL Defect Detection System Based on Multi-modal Semantic Segmentation Model Deploy the VCSEL multi-modal semantic segmentation model to computer 5 and connect it to related devices such as visible light sources, infrared light sources, infrared CCD 6, visible light CCD 7, and conveyor belt 4 to form a complete automated VCSEL defect detection system. The specific implementation is as Figure 2 shown. The VCSEL chip 3 to be tested is placed on the conveyor belt 4, and the positions and angles of the infrared light source and visible light source are adjusted so that the infrared CCD 6 and visible light CCD 7 can capture clear infrared images and visible light images.
[0035] The flow block diagram of the operation of the defect detection system is as Figure 6 shown: 1) The system starts running; 2) Parameter initialization, set network parameters, the running speed of the conveyor belt 4, the number of VCSEL chips 3 to be tested, light source intensity, voltage and current parameters, shooting angles, so that the light source and CCD can obtain the best images and cooperate well with the operation of the conveyor belt 4; 3) Start automatic detection. Computer 3 determines whether the VCSEL chip 3 to be tested is qualified through the built-in VCSEL multi-modal semantic segmentation model; the VCSEL multi-modal semantic segmentation model is a model that has been trained and completed; 4) If the system determines that the VCSEL chip 3 to be tested is qualified, record the chip identification number; if it is unqualified, save the detection log and output the chip identification number, defect category, defect location, defect form, detection speed and accuracy; 5) The system operation ends.
[0036] The present invention also provides relevant experimental results and data to prove that the method described in the present invention has high detection accuracy.
[0037] Experimental Results 1) Production of multi-modal defect data sets (corresponding to step S1): The method proposed in step S1 of the present invention is qualitatively and quantitatively compared with eight image fusion methods. These eight methods represent typical models in the field of deep learning image fusion. First of all, these methods are all based on deep learning technology and can effectively handle multi-target image fusion tasks, such as adaptively extracting multi-modal defect features, etc. At the same time, they achieve a good balance between retaining image details and enhancing image contrast, and are suitable for complex image scenarios and multi-modal target segmentation. Secondly, these eight methods are different in terms of network architecture, loss function design, feature fusion strategy, etc., which enhances the comparability of the fusion technology and makes the fusion results more persuasive. These eight algorithms are all trained on the multi-modal defect dataset built by the present invention (as described in step S1) and fine-tuned according to their respective official training strategies to obtain the best fusion effect. The present invention selects 6 groups of defective chips for multi-modal defect fusion operation, and the fusion results are as shown in Figure 7 and Table 1.
[0038] Table 1 Comparison table of quality evaluation results of multi-modal defect dataset in step S1
[0039] Table 1 and Figure 7 "Proposed" in represents the method for obtaining the multi-modal defect dataset designed in step S1 of the present invention. Among them, "EN" in Table 1 represents the information entropy of the image, "SD" represents the standard deviation of the image, "MI" represents the mutual information between the source image and the fused image, and "Q abf " represents the fusion quality; it can be seen from Table 1 that the multi-modal defect dataset obtained by using step S1 of the present invention is superior to the current state-of-the-art fusion methods in terms of information entropy, standard deviation, mutual information and fusion quality indicators, which indicates that the multi-modal defect dataset obtained in step S1 of the present invention has high image quality and also lays a good foundation for improving the detection accuracy of subsequent VCSEL multi-modal defects.
[0040] 2) VCSEL multi-modal defect model training and detection (corresponding to steps S3 and S4): According to the model training settings in step S3, a trained VCSEL multi-modal semantic segmentation model is obtained. The present invention compares the results of this model with 7 high-performance segmentation models (U-Net, UNet++, Attention U-Net, Swin-Unet, SegNet, Fast-SegNet, DeepLabv3). The comparison of the detection results is shown in Table 2.
[0041] Table 2 Comparison table of detection results of VCSEL multi-modal semantic segmentation model in step S4
[0042] In Table 2, "Proposed" represents the method of the VCSEL multi-modal semantic segmentation model designed in step S4 of the present invention. Among them, "DSC" is used to evaluate the similarity between the predicted value and the true value, "mIoU" is used to evaluate the overall overlap between the predicted result and the true result, "Recall" is used to measure the recognition ability of the model for positive defect examples, and "Precision" represents the prediction accuracy. The higher the value, the higher the reliability of the model in predicting positive defect examples.
[0043] Figure 8 It is a schematic diagram of the multi-modal defect detection result in step S4 of the present invention; during the process of labeling the data set, the present invention classifies four types of defects, namely surface contamination, mechanical scratches, epitaxial defects, and dark defects in the light-emitting aperture, into two categories. Among them, surface contamination, mechanical scratches, and epitaxial defects are all classified as surface defects 8, and dark defects in the light-emitting aperture are classified as internal defects 9. In addition, in order to enable the segmentation algorithm to accurately locate the dark defects at the light-emitting aperture, the present invention additionally labels a category of light-emitting aperture 10. Figure 8 It can be seen from the detection result that the VCSEL multi-modal semantic segmentation model can accurately detect surface defects 8 and internal defects 9, and label 10 is the light-emitting aperture of the chip.
[0044] The above results show that the VCSEL defect detection method based on the multi-modal semantic segmentation model proposed by the present invention far exceeds other models in terms of four indicators, and the detection accuracy reaches 95.74%. This detection method can well segment the dark area with complex texture and fuzzy edges. The surface defect information and the active area defect information in the dark area help engineers to conduct the next failure analysis on the failed VCSEL chips, contributing to improving the stability and reliability of semiconductor laser chips.
Claims
1. A method for detecting VCSEL defects based on a multimodal semantic segmentation model, characterized in that: It includes the following steps: S1. Obtain a multi-modal defect dataset; S11. Collect visible light images and infrared images of multiple VCSEL sample chips (1); the collected VCSEL sample chips (1) include defect-free VCSEL chips and defective VCSEL chips; S12. Respectively fuse the visible light image and the infrared image of each VCSEL sample chip (1) through a multi-modal image fusion algorithm to obtain a multi-modal defect fusion image of each VCSEL sample chip (1); S13. Mark the multi-modal defect fusion images of the VCSEL sample chips (1) obtained in step S12, eliminate the defect-free multi-modal defect fusion images, and mark the defective multi-modal defect fusion images with defect categories, defect positions, and defect morphologies. The marked multi-modal defect fusion images are used as chip defect samples to form a multi-modal defect dataset; S2. Design of a VCSEL multi-modal semantic segmentation model; S21. Use the pre-trained convolutional neural network ResNet as the encoder, and the encoder includes multiple layers of ordinary convolutions, and an SE attention module is added after each layer of ordinary convolution; S22. Use dilated convolutions that are the same in number and in one-to-one correspondence with the number of encoder layers to form a decoder; except for the last layer, each layer of ordinary convolution in the encoder jumps to the corresponding layer of dilated convolution in the decoder; the last layer of ordinary convolution in the encoder is connected to the last layer of dilated convolution in the decoder through depthwise separable convolution to form a VCSEL multi-modal semantic segmentation model; S3. Training of the VCSEL multi-modal semantic segmentation model Input the chip defect samples of the multi-modal defect dataset obtained in S1 into the VCSEL multi-modal semantic segmentation model to train the VCSEL multi-modal semantic segmentation model. The VCSEL multi-modal semantic segmentation model can output a multi-modal defect segmentation result map of each chip defect sample, and the multi-modal defect segmentation result map contains the defect category, defect position, and defect morphology of the chip defect sample; when the defect category, defect position, and defect morphology in the multi-modal defect segmentation result map output by the VCSEL multi-modal semantic segmentation model are compared with the chip defect sample and the accuracy reaches more than 95%, it is regarded as the completion of the training of the VCSEL multi-modal semantic segmentation model; S4. Detection of the VCSEL multi-modal semantic segmentation model Collect the visible light image and the infrared image of the VCSEL chip to be tested (3), obtain the multi-modal defect fusion image of the VCSEL chip to be tested (3) through a multi-modal image fusion algorithm, and then input the multi-modal defect fusion image of the VCSEL chip to be tested (3) into the VCSEL multi-modal semantic segmentation model completed in training in S3, and the VCSEL multi-modal semantic segmentation model outputs a multi-modal defect segmentation result map including the defect category, defect position, and defect morphology of the VCSEL chip to be tested (3).
2. The method for detecting VCSEL defects based on a multimodal semantic segmentation model according to claim 1, characterized in that, The defect categories include surface contamination, mechanical scratches, epitaxial defects, and dark defects in the light-emitting holes.
3. The method for detecting VCSEL defects based on a multimodal semantic segmentation model according to claim 1, characterized in that, In step S3, the following training parameters and configurations are set for the VCSEL multi-modal semantic segmentation model before training: the initial learning rate is set to 1e -4 , and the cosine annealing learning rate scheduler is used to dynamically adjust the learning rate; the AdamW optimizer is used, combined with weight decay to prevent overfitting; the batch size is set to 16, and the number of training epochs is 150; the model weights are saved after training is completed.
4. A system for detecting VCSEL defects based on a multimodal semantic segmentation model, used to implement the method for detecting VCSEL defects based on a multimodal semantic segmentation model according to any one of claims 1-3, characterized in that, It includes a probe station (2) for placing the VCSEL sample chip (1), a conveyor belt (4) for placing the VCSEL chip (3) to be tested, an image acquisition system for acquiring images of the VCSEL sample chip (1) and the VCSEL chip (3) to be tested, and a computer (5); the image acquisition system includes an infrared CCD (6) for acquiring infrared images of the VCSEL sample chip (1) and the VCSEL chip (3) to be tested, a visible light CCD (7) for acquiring visible light images of the VCSEL sample chip (1) and the VCSEL chip (3) to be tested, a visible light source and an infrared light source for providing light sources for the VCSEL sample chip (1) and the VCSEL chip (3) to be tested; the signal output ends of the infrared CCD (6) and the visible light CCD (7) are connected to the computer (5), and the probe station (2), the conveyor belt (4), the visible light source, and the infrared light source are all electrically connected to the computer (5) and controlled by the computer (5) to act; the computer (5) is built-in with a multi-modal image fusion algorithm and a VCSEL multi-modal semantic segmentation model.
5. The system for detecting VCSEL defects based on a multimodal semantic segmentation model according to claim 4, characterized in that, Before the defect detection system conducts detection, first set the running speed of the conveyor belt (4), the intensities of the visible light source and the infrared light source, the shooting angles of the infrared CCD (6) and the visible light CCD (7), and assign an identification number to each VCSEL chip (3) to be tested; then start automatic detection. The multi-modal image fusion algorithm built in the computer (5) fuses the visible light image and the infrared image of the VCSEL chip (3) to be tested collected, and after obtaining the multi-modal defect fusion image of the VCSEL chip (3) to be tested, the completed trained VCSEL multi-modal semantic segmentation model built in the computer (5) detects the multi-modal defect fusion image of the VCSEL chip (3) to be tested. If there are no defects in the multi-modal defect fusion image of the VCSEL chip (3) to be tested, the computer (5) outputs the identification number of the VCSEL chip (3) to be tested and marks it as qualified; If there are defects in the multi-modal defect fusion image of the VCSEL chip (3) to be tested, the computer outputs the identification number, defect category, defect location, and defect form of the VCSEL chip (3) to be tested and marks it as unqualified.
Citation Information
Patent Citations
VCSEL quality screening method based on image fusion algorithm
CN117830235A
Automatic segmentation method for residual UNet rectal cancer tumor magnetic resonance image
CN112785617A
Laser chip defect detection and classification method and system based on SSD algorithm
CN113077430A
High-voltage transmission line identification method
CN115424016A
Building exterior wall defect detection method based on deep learning multi-modal image fusion
CN116091477A