Automatic Detection Method and System for Semiconductor Chip Defects Based on Hyperspectral Imaging
Through the combination of hyperspectral imaging technology and neural network, the problem of feature extraction and distinction in semiconductor chip defect detection is solved, and high-precision and reliable defect detection is achieved, which can identify multiple defects and optimize the detection results.
Patent Information
- Application Number
- CN202510029830.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing semiconductor chip defect detection methods based on deep learning are difficult to effectively extract and distinguish different types of defect characteristics, resulting in low detection accuracy and ignoring the influence of process parameters.
Using a hyperspectral imaging-based method, image enhancement is performed through improved neuromorphic adaptive processing networks and dual-path diffusion models, combined with dynamic hybrid feature enhancement networks of neural architecture search and multi-branch expert decision-making modules, the reliability evaluation and optimization of defect detection results is used using spatiotemporal knowledge inference maps.
The accuracy and reliability of defect detection are improved, and multiple types of defects can be identified, false detection rates and missed detection rates can be reduced. The reliability of detection results is enhanced by analyzing the morphological similarity, spatial position relationship and timing correlation between defects.
Smart Images

Figure CN119417838B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chip detection, and particularly to an automatic semiconductor chip defect detection method and system based on hyperspectral imaging. Background Art
[0002] The manufacturing process of semiconductor chips is complex, and various defects are likely to occur during production, such as scratches, cracks, and contamination. These defects will seriously affect the performance and reliability of the chips. Traditional defect detection methods mainly rely on manual visual inspection, which is inefficient, costly, and easily affected by subjective factors, resulting in inconsistent detection results.
[0003] With the development of deep learning technology, image recognition and object detection technologies based on deep learning have achieved remarkable results in various fields. Applying deep learning technology to semiconductor chip defect detection can achieve automated, high-precision, and high-efficiency defect detection. However, existing deep learning-based semiconductor chip defect detection methods still have problems such as difficulty in effectively extracting and distinguishing different types of defect features, resulting in low detection accuracy, and ignoring the influence of process parameters.
[0004] Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0005] Embodiments of the present invention provide an automatic semiconductor chip defect detection method and system based on hyperspectral imaging, which can at least solve some of the problems existing in the prior art.
[0006] In a first aspect of the embodiments of the present invention, an automatic semiconductor chip defect detection method based on hyperspectral imaging is provided, including:
[0007] Collect hyperspectral image data of a semiconductor chip to be detected, add the hyperspectral image data to an improved neuromorphic adaptive processing network and calculate a noise distribution density map and a texture complexity map of a local area. Based on the noise distribution density map and the texture complexity map, construct an adaptive enhancement matrix. Based on the adaptive enhancement matrix, calculate wavelet decomposition coefficients of different scales, perform a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, and reconstruct through inverse wavelet transform to obtain preliminarily enhanced image data; input the preliminarily enhanced image data into a dual-path diffusion model, extract high-frequency detail features through a first path, extract low-frequency structure features through a second path, and adaptively fuse the high-frequency detail features and the low-frequency structure features to obtain a compensation feature map. Based on the compensation feature map, correct the uneven illumination of the preliminarily enhanced image data to obtain corrected image data.
[0008] Construct a dynamic hybrid feature enhancement network based on the neural architecture search method, add the corrected image data to the dynamic hybrid feature enhancement network and generate a multi-level feature pyramid, perform cross-layer attention operations on the feature maps in adjacent layers of the multi-level feature pyramid, generate inter-layer correlation weights, reconstruct and enhance the feature maps of each level, generate a fused feature map and add it to a pre-set multi-branch expert decision-making module, where the multi-branch expert decision-making module includes multiple sub-networks for different types of defects, each sub-network independently generates an independent detection result corresponding to the fused feature map, and combines the independent detection results by a dynamic integration strategy based on uncertainty to obtain an initial defect detection result;
[0009] Build a spatio-temporal knowledge inference graph based on the initial defect detection result, set the initial defect detection result as a graph node, calculate the morphological similarity and spatial position relationship between each initial defect detection result and construct the edge connection strength, perform hierarchical clustering on the initial defect detection result based on the edge connection strength to generate similar defect groups, perform graph attention network operations on the spatio-temporal knowledge inference graph, generate defect correlation feature vectors, perform temporal correlation analysis on the defect correlation feature vectors and the process parameter sequence, construct a process parameter influence model, and perform reliability evaluation and optimization on the initial defect detection result based on the process parameter influence model and the similar defect groups to obtain the final defect detection result.
[0010] In an alternative embodiment,
[0011] Collect the hyperspectral image data of the semiconductor chip to be detected, add the hyperspectral image data to the improved neuromorphic adaptive processing network and calculate the noise distribution density map and texture complexity map of the local area, construct an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculate the wavelet decomposition coefficients at different scales based on the adaptive enhancement matrix, and perform non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, including:
[0012] Collect the hyperspectral image data of the semiconductor chip to be detected, where the hyperspectral image data includes image data in the visible light band, near-infrared band, and short-wave infrared band;
[0013] Input the hyperspectral image data into the improved neuromorphic adaptive processing network, calculate the local statistical features of the hyperspectral image data at different scales, extract the noise distribution feature vector and texture feature vector, construct a multi-layer neural network based on the noise distribution feature vector and calculate the noise distribution density map, construct a convolutional neural network based on the texture feature vector and calculate the texture complexity map, and input the noise distribution density map and the texture complexity map into the attention mechanism module to generate an attention weight map;
[0014] Perform feature reconstruction on the noise distribution density map and the texture complexity map according to the attention weight map, fuse the reconstructed feature maps to construct an adaptive enhancement matrix, perform adaptive enhancement processing on different bands of the hyperspectral image data based on the adaptive enhancement matrix, construct a multi-scale directional filter bank and a wavelet decomposition unit, perform directional decomposition, scale decomposition and wavelet decomposition on the enhanced hyperspectral image data simultaneously, obtain multiple directional sub-bands and scale sub-bands through the multi-scale directional filter bank, and obtain an approximation coefficient sub-band and a detail coefficient sub-band through the wavelet decomposition unit;
[0015] Perform adaptive threshold segmentation on the directional sub-band, the scale sub-band, the approximation coefficient sub-band and the detail coefficient sub-band, extract significant feature coefficients, perform adaptive weight modulation on the significant feature coefficients and the adaptive enhancement matrix, and perform a non-linear activation function operation on the modulated coefficients to generate enhancement coefficients.
[0016] In an alternative embodiment,
[0017] Reconstruct the initially enhanced image data through inverse wavelet transform; input the initially enhanced image data into a dual-path diffusion model, extract high-frequency detail features through the first path, extract low-frequency structural features through the second path, and adaptively fuse the high-frequency detail features and the low-frequency structural features to obtain a compensation feature map, and correct the illumination non-uniformity of the initially enhanced image data based on the compensation feature map to obtain the corrected image data including:
[0018] Construct a deep residual reconstruction network, input the enhancement coefficients into the deep residual reconstruction network, and reconstruct the initially enhanced image data through residual learning and skip connection structures;
[0019] Construct a dual-branch attention diffusion network, the dual-branch attention diffusion network includes a spatial attention branch and a channel attention branch, input the initially enhanced image data into the dual-branch attention diffusion network, extract a high-frequency detail feature map through the spatial attention branch, and extract a low-frequency structural feature map through the channel attention branch;
[0020] Construct a hierarchical feature pyramid network, where the hierarchical feature pyramid network includes a feature mapping module, a feature reconstruction module, and a feature optimization module. Input the high-frequency detail feature map and the low-frequency structure feature map into the feature mapping module to generate a set of feature maps at multiple scales. Input the set of feature maps into the feature reconstruction module, calculate the similarity matrix between features at different scales by combining the self-attention mechanism, and perform feature reconstruction based on the similarity matrix. Input the reconstructed feature map into the feature optimization module, and iteratively optimize the features through a recurrent neural network and a feedback mechanism to obtain an optimized compensation feature map;
[0021] Construct an adaptive illumination compensation network, where the adaptive illumination compensation network includes an illuminance analysis module, a reflectance analysis module, and an illumination equalization module. Input the compensation feature map into the illuminance analysis module, and estimate the image illuminance component and the illuminance distribution map through a multi-scale convolutional neural network. Input the compensation feature map into the reflectance analysis module, and restore the image reflectance component and the texture detail map through a depthwise separable convolutional network. Input the illuminance distribution map and the texture detail map into the illumination equalization module, and correct the non-uniformity of illumination for the preliminarily enhanced image data through an adaptive exposure compensation algorithm and a local contrast enhancement algorithm, and output the corrected image data.
[0022] In an optional implementation manner,
[0023] Construct a dynamic hybrid feature enhancement network based on the neural architecture search method. Add the corrected image data to the dynamic hybrid feature enhancement network and generate a multi-level feature pyramid. Perform cross-layer attention operations on the feature maps in adjacent layers in the multi-level feature pyramid to generate inter-layer correlation weights, reconstruct and enhance the feature maps at each level, generate a fused feature map, and add it to a pre-set multi-branch expert decision-making module. Among them, the multi-branch expert decision-making module includes multiple sub-networks for different types of defects. Each sub-network independently generates an independent detection result corresponding to the fused feature map, and combines the independent detection results through a dynamic integration strategy based on uncertainty to obtain an initial defect detection result, including:
[0024] Construct a dynamic hybrid feature enhancement network based on the neural architecture search method. The neural architecture search method selects and combines multiple convolutional layers, pooling layers, and activation function components from a predefined network structure space through a search algorithm, generates a candidate network, evaluates the performance of the candidate network, and selects the network architecture with the optimal performance as the structure of the dynamic hybrid feature enhancement network;
[0025] Input the corrected image data into the dynamic hybrid feature enhancement network. Perform multiple downsampling operations through the encoder module in the dynamic hybrid feature enhancement network to convert the corrected image data into low-resolution features. Perform multiple upsampling operations through the decoder module in the dynamic hybrid feature enhancement network to restore the low-resolution features to the original resolution features. Extract feature maps of multiple different scales at different levels of the encoder module and the decoder module to construct a multi-level feature pyramid;
[0026] Perform cross-layer attention operations on the feature maps of adjacent layers in the multi-level feature pyramid. Map the corrected feature maps of adjacent layers to the same number of channels through convolution operations, calculate the query matrix and the key matrix of the corrected feature maps in adjacent layers, generate an attention weight matrix based on the query matrix and the key matrix, reconstruct the corrected feature maps based on the attention weight matrix, and perform the cross-layer attention operation and feature reconstruction operation on the corrected feature maps in each layer of the multi-level feature pyramid to generate a fused pyramid feature map;
[0027] Construct a multi-branch expert decision-making module. The multi-branch expert decision-making module includes multiple independent sub-networks. Each sub-network adopts an encoder-decoder structure and is equipped with a classifier. Input the fused pyramid feature map into the multi-branch expert decision-making module, and detect different types of defects through multiple sub-networks respectively. Each sub-network outputs a corresponding independent detection result;
[0028] Obtain the confidence levels corresponding to the independent detection results output by multiple sub-networks. Based on the confidence levels, perform weighted average fusion on the independent detection results through a dynamic integration strategy to obtain an initial defect detection result.
[0029] In an alternative embodiment,
[0030] Perform cross-layer attention operations on the feature maps of adjacent layers in the multi-level feature pyramid as shown in the following formula:
[0031] ;
[0032] where A represents the element of the calculated attention weight matrix, represents the k-th element of the query vector at the m-th position in the feature map of the i-th layer, represents the weight matrix of the query matrix of the i-th layer, converting the query vector from dimension d h to dimension d l dimension, represents the k-th element of the key vector at the n-th position in the feature map of the j-th layer, represents the weight matrix of the key matrix of the i-th layer, converting the key vector from dimension dh Convert the dimension to d l dimension, d h represents the original dimension of the query vector and the key vector, d l represents the dimension after the query vector and the key vector are linearly transformed, h represents the number of heads in the multi-head attention mechanism represents the attention weight between the m-th position in the feature map of the i-th layer and the n-th position in the feature map of the j-th layer in the g-th attention head represents the weight matrix of the output matrix, O is used to indicate the output
[0033] In an alternative embodiment
[0034] Based on the initial defect detection result, establish a spatio-temporal knowledge inference graph, set the initial defect detection result as a graph node, calculate the morphological similarity and spatial position relationship between each initial defect detection result and construct the edge connection strength, perform hierarchical clustering on the initial defect detection result based on the edge connection strength to generate similar defect groups, perform graph attention network operations on the spatio-temporal knowledge inference graph, generate defect correlation feature vectors and perform temporal correlation analysis on the defect correlation feature vectors and the process parameter sequence, construct a process parameter influence model, and perform reliability evaluation and optimization on the initial defect detection result based on the process parameter influence model and the similar defect groups to obtain the final defect detection result, including:
[0035] Obtain the initial defect detection result of the chip image, the initial defect detection result includes the morphological features and spatial position coordinates of each defect, the morphological features include length, width and area, and the spatial position coordinates include the row number and column number of the center point of the defect;
[0036] Set each defect in the initial defect detection result as a graph node, establish a spatio-temporal knowledge inference graph, each graph node includes the morphological features and spatial position coordinates of the corresponding defect, for each pair of graph nodes in the spatio-temporal knowledge inference graph, obtain the morphological similarity by calculating the cosine similarity between the morphological features of different graph nodes, obtain the spatial position distance by calculating the Euclidean distance between the spatial position coordinates of different graph nodes, construct the edge connection strength based on the morphological similarity and the spatial position distance, if the morphological similarity between two graph nodes is greater than a pre-set similarity threshold and the spatial position distance is less than a pre-set distance threshold, calculate the edge connection strength according to the spatial position distance;
[0037] Based on the edge connection strength, cluster the graph nodes by a bottom-up hierarchical clustering method. Initially, set each graph node as an independent cluster, iteratively merge the two clusters with the maximum edge connection strength, and repeat the merging until the edge connection strength between all clusters is less than a preset connection strength threshold, generating multiple similar defect groups;
[0038] Construct a graph attention network, which includes an attention layer and a feature propagation layer. For the spatio-temporal knowledge inference graph, calculate the attention weights between each graph node and its corresponding adjacent nodes through the attention layer. The attention weights are determined based on the similarity degree of node features. Based on the attention weights, update the features of the graph nodes through the feature propagation layer, so that the features of each graph node aggregate the weighted information of the corresponding adjacent nodes, generating a node feature matrix;
[0039] Extract the defect correlation feature vectors corresponding to each similar defect group in the node feature matrix, obtain the process parameter sequence during the corresponding time period of the similar defect group. The process parameter sequence includes temperature parameters and pressure parameters. Calculate the cross-correlation coefficient between the defect correlation feature vector and the process parameter sequence through the Granger causality test method. At the same time, construct a process parameter influence model according to the cross-correlation coefficient. Input each defect in the initial defect detection result into the process parameter influence model, determine the corresponding similar defect group based on the morphological features in the initial defect detection result, and judge whether the current defect is caused by process abnormality according to the prediction result output by the process parameter influence model. Mark the defects determined to be caused by process abnormality as reliable defects, filter out the defects not caused by process abnormality, and obtain the final defect detection result.
[0040] In an optional implementation manner,
[0041] Calculating the cross-correlation coefficient between the defect correlation feature vector and the process parameter sequence through the Granger causality test method includes:
[0042] Obtain the process parameter sequence and the defect correlation feature vector in the chip manufacturing process. Divide the process parameter sequence into a training data set and a test data set in chronological order. The training data set is used to construct a prediction model, and the test data set is used to evaluate the performance of the prediction model;
[0043] Construct a first prediction model based on the training data set, and predict the process parameter value at the current moment using the historical data of the process parameter sequence within a preset time window. Construct a second prediction model based on the training data set, and simultaneously predict the process parameter value at the current moment using the historical data of the process parameter sequence and the defect correlation feature vector within a preset time window;
[0044] Input the test data set into the first prediction model and the second prediction model, calculate the sum of squared residuals of the two prediction models by summing the squares of the differences between the predicted values and the actual values, and construct a Granger causality test statistic, which is calculated based on the sum of squared residuals of the first prediction model, the sum of squared residuals of the second prediction model, the number of parameters of the first prediction model, the number of parameters of the second prediction model, and the number of samples in the test data set;
[0045] Determine the critical value based on a preset significance level. When the Granger causality test statistic is greater than the critical value, determine that the defect-associated feature vector is the Granger cause of the process parameter sequence, and calculate the cross-correlation coefficient between each feature in the defect-associated feature vector and the process parameter sequence at different time lags based on the covariance and variance of the defect-associated feature vector and the process parameter sequence.
[0046] In a second aspect of the embodiments of the present invention, a semiconductor chip defect automatic detection system based on hyperspectral imaging is provided, including:
[0047] A first unit for collecting hyperspectral image data of a semiconductor chip to be detected, adding the hyperspectral image data to an improved neuromorphic adaptive processing network and calculating a noise distribution density map and a texture complexity map of a local area, constructing an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculating wavelet decomposition coefficients at different scales based on the adaptive enhancement matrix, performing a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, and reconstructing through inverse wavelet transform to obtain preliminarily enhanced image data; inputting the preliminarily enhanced image data into a dual-path diffusion model, extracting high-frequency detail features through a first path, extracting low-frequency structure features through a second path, adaptively fusing the high-frequency detail features and the low-frequency structure features to obtain a compensation feature map, and correcting the illumination non-uniformity of the preliminarily enhanced image data based on the compensation feature map to obtain corrected image data;
[0048] A second unit is configured to construct a dynamic hybrid feature enhancement network based on a neural architecture search method, add the corrected image data to the dynamic hybrid feature enhancement network and generate a multi-level feature pyramid, perform cross-layer attention operations on the feature maps in adjacent layers of the multi-level feature pyramid, generate inter-layer correlation weights, reconstruct and enhance the feature maps of each level, generate fused feature maps, and add the fused feature maps to a pre-set multi-branch expert decision module. The multi-branch expert decision module includes multiple sub-networks for different types of defects. Each sub-network independently generates an independent detection result corresponding to the fused feature map, and combines the independent detection results by using a dynamic integration strategy based on uncertainty to obtain an initial defect detection result;
[0049] A third unit is configured to establish a spatio-temporal knowledge inference graph based on the initial defect detection result, set the initial defect detection result as a graph node, calculate the morphological similarity and spatial position relationship between each initial defect detection result and construct edge connection strength, perform hierarchical clustering on the initial defect detection results based on the edge connection strength to generate similar defect groups, perform graph attention network operations on the spatio-temporal knowledge inference graph, generate defect association feature vectors, perform temporal correlation analysis on the defect association feature vectors and a process parameter sequence, construct a process parameter influence model, and perform reliability evaluation and optimization on the initial defect detection result based on the process parameter influence model and the similar defect groups to obtain a final defect detection result.
[0050] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: a processor;
[0051] a memory for storing instructions executable by the processor;
[0052] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0053] In a fourth aspect of the embodiments of the present invention,
[0054] there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0055] In the present invention, the wavelet decomposition coefficients are optimized by an adaptive enhancement matrix, and the uneven illumination is corrected by combining a dual-path diffusion model, effectively enhancing the image quality and highlighting the defect features, thereby improving the accuracy of defect detection. The dynamic hybrid feature enhancement network constructed based on neural architecture search, combined with a cross-layer attention mechanism and a multi-branch expert decision-making module, can effectively extract and fuse multi-level features to achieve accurate recognition of different types of defects. The initial detection results are evaluated and optimized using a spatio-temporal knowledge inference graph. By analyzing the morphological similarity, spatial position relationship, and temporal correlation with process parameters among defects, the false detection rate and missed detection rate are effectively reduced, and the reliability of the detection results is improved. Hyperspectral imaging can simultaneously obtain spatial and spectral information. Each pixel contains a continuous spectral curve, providing richer spectral feature information than traditional RGB images. Different types of defects have unique spectral features in specific bands, and accurate identification and classification of defects can be achieved through spectral features. By using the images obtained by hyperspectral imaging, various surface defects such as scratches, cracks, contamination, and oxidation can be simultaneously identified, defects of different depths and sizes can be recognized, and quantitative information about the defects can be provided. Hyperspectral data has high-dimensional features and is insensitive to interference factors such as environmental illumination changes and surface reflections. Irrelevant information interference can be removed through band selection to improve the detection reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a schematic flow chart of the method for automatically detecting semiconductor chip defects based on hyperspectral imaging according to an embodiment of the present invention;
[0057] Figure 2 is a schematic structural diagram of the system for automatically detecting semiconductor chip defects based on hyperspectral imaging according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0060] Figure 1 is a schematic flow chart of the method for automatically detecting semiconductor chip defects based on hyperspectral imaging according to an embodiment of the present invention, as Figure 1As shown, the method includes:
[0061] S1. Collect hyperspectral image data of the semiconductor chip to be detected, add the hyperspectral image data to an improved neuromorphic adaptive processing network and calculate the noise distribution density map and texture complexity map of the local area. Construct an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculate wavelet decomposition coefficients of different scales based on the adaptive enhancement matrix, perform a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, and reconstruct through inverse wavelet transform to obtain preliminarily enhanced image data; Input the preliminarily enhanced image data into a dual-path diffusion model, extract high-frequency detail features through the first path, extract low-frequency structure features through the second path, and adaptively fuse the high-frequency detail features and the low-frequency structure features to obtain a compensation feature map. Correct the illumination non-uniformity of the preliminarily enhanced image data based on the compensation feature map to obtain corrected image data;
[0062] The neuromorphic adaptive processing network is a network structure based on a neuromorphic computing model, aiming to imitate the structure and processing mechanism of the biological nervous system. By adaptively adjusting network parameters, it adapts to the characteristics of different input data, thereby realizing effective processing of noise and signals. The noise distribution density map is an image describing the noise distribution in a signal or image, usually used to represent the noise intensity distribution of different regions or pixels. The texture complexity map is an image generated by analyzing the texture features of different regions in an image, reflecting the texture complexity. The inverse wavelet transform reconstruction is a signal or image reconstruction method that restores the processed data in the wavelet domain to the original time domain or spatial domain by reversely performing the wavelet transform process. The dual-path diffusion model is a mathematical model used for image processing or signal transmission, especially widely applied in image denoising, diffusion, smoothing, etc. The non-uniformity correction is a technique used for signal processing and image processing, aiming to correct the non-uniformity of signals or images caused by factors such as the environment and equipment.
[0063] In an alternative embodiment,
[0064] Collecting hyperspectral image data of the semiconductor chip to be detected, adding the hyperspectral image data to an improved neuromorphic adaptive processing network and calculating the noise distribution density map and texture complexity map of the local area, constructing an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculating wavelet decomposition coefficients of different scales based on the adaptive enhancement matrix, and performing a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients includes:
[0065] Collect hyperspectral image data of the semiconductor chip to be detected, where the hyperspectral image data includes image data in the visible light band, near-infrared band, and short-wave infrared band;
[0066] Input the hyperspectral image data into an improved neuromorphic adaptive processing network, calculate the local statistical features of the hyperspectral image data at different scales, extract the noise distribution feature vector and texture feature vector, construct a multi-layer neural network based on the noise distribution feature vector and calculate the noise distribution density map, construct a convolutional neural network based on the texture feature vector and calculate the texture complexity map, and input the noise distribution density map and the texture complexity map into the attention mechanism module to generate an attention weight map;
[0067] According to the attention weight map, perform feature reconstruction on the noise distribution density map and the texture complexity map, fuse the reconstructed feature maps to construct an adaptive enhancement matrix, perform adaptive enhancement processing on different bands of the hyperspectral image data based on the adaptive enhancement matrix, construct a multi-scale direction filter bank and a wavelet decomposition unit, perform direction decomposition, scale decomposition, and wavelet decomposition on the enhanced hyperspectral image data simultaneously, obtain multiple direction sub-bands and scale sub-bands through the multi-scale direction filter bank, and obtain an approximation coefficient sub-band and detail coefficient sub-bands through the wavelet decomposition unit;
[0068] Perform adaptive threshold segmentation on the direction sub-bands, the scale sub-bands, the approximation coefficient sub-bands, and the detail coefficient sub-bands, extract significant feature coefficients, perform adaptive weight modulation on the significant feature coefficients and the adaptive enhancement matrix, and perform a non-linear activation function operation on the modulated coefficients to generate enhancement coefficients.
[0069] The multi-scale direction filter bank is a set of filters that processes input data at multiple scales and directions. The wavelet decomposition unit is a module for performing wavelet transform on signals or images. The significant feature coefficients refer to the numerical values or coefficients with significant features extracted from the input data during image processing or signal processing.
[0070] Collect hyperspectral image data of the semiconductor chip to be detected. Use a hyperspectral camera equipped with a filter wheel to obtain image data of the chip at different bands. These bands cover the visible light band (e.g., 400 - 700nm, e.g., collect red, green, and blue color images), the near-infrared band (e.g., 700 - 1000nm, e.g., collect images at three bands of 750nm, 850nm, and 950nm), and the short-wave infrared band (e.g., 1000 - 2500nm, e.g., collect images at three bands of 1200nm, 1600nm, and 2200nm). The image data of each band is saved as a grayscale image, and finally a hyperspectral image dataset is formed. For example, for a chip, we collect image data of nine bands including the three visible light bands of red, green, and blue, the three near-infrared bands, and the three short-wave infrared bands.
[0071] Input the hyperspectral image data into an improved neuromorphic adaptive processing network. The network structure contains multiple modules for extracting local statistical features at different scales, performing multi-scale Gaussian filtering on the input hyperspectral image data. For example, use Gaussian filters with standard deviations of 1, 2, and 4 respectively for filtering to obtain smoothed images at different scales. At each scale, calculate the statistical features within the neighborhood of each pixel (e.g., neighborhoods of 3x3, 5x5, 7x7), such as mean, variance, gradient, etc. These statistical features constitute a noise distribution feature vector and a texture feature vector. Build a multi-layer neural network based on the noise distribution feature vector, such as a multi-layer perceptron with two hidden layers, learn the pattern of the noise distribution through training, and output a noise distribution density map. Build a convolutional neural network based on the texture feature vector, such as a convolutional neural network with three convolutional layers and two pooling layers, learn the pattern of the texture through training, and output a texture complexity map. Input the noise distribution density map and the texture complexity map into an attention mechanism module, such as using a spatial attention mechanism, calculate the weight at each position, and generate an attention weight map. For example, if a pixel point has a high noise density and complex texture, the value of this pixel point on the attention weight map is high.
[0072] Perform feature reconstruction on the noise distribution density map and the texture complexity map according to the attention weight map. Multiply the attention weight map with the noise distribution density map and the texture complexity map to obtain weighted feature maps. Fuse the weighted noise distribution density map and the texture complexity map, such as adding them in a certain proportion to construct an adaptive enhancement matrix. For example, multiply the weighted noise distribution density map and the texture complexity map by 0.6 and 0.4 respectively and then add them to obtain the final adaptive enhancement matrix. Perform adaptive enhancement processing on different bands of the hyperspectral image data based on the adaptive enhancement matrix. Multiply the adaptive enhancement matrix with the image data of each band to obtain enhanced image data.
[0073] Construct a multi-scale directional filter bank and a wavelet decomposition unit. The multi-scale directional filter bank contains filters with multiple directions and scales, such as Gabor filters, which are used to extract texture information in different directions and scales. The wavelet decomposition unit decomposes the image using wavelet transform to obtain an approximation coefficient sub-band and detail coefficient sub-bands. For example, perform a three-level decomposition using the Haar wavelet. Perform adaptive threshold segmentation on the directional sub-bands, scale sub-bands, approximation coefficient sub-band, and detail coefficient sub-bands. According to the statistical characteristics of each sub-band, such as the mean and standard deviation, adaptively determine the threshold, set the coefficients less than the threshold to zero, and retain the coefficients greater than the threshold, thereby extracting significant feature coefficients. For example, for a certain sub-band, calculate its mean and standard deviation, and set the threshold to the mean plus twice the standard deviation.
[0074] Perform adaptive weight modulation on the significant feature coefficients and an adaptive enhancement matrix. Multiply the significant feature coefficients by the elements at the corresponding positions of the adaptive enhancement matrix to obtain the modulated coefficients. Perform a non-linear activation function operation on the modulated coefficients, such as the ReLU function or the Sigmoid function, to generate enhanced coefficients. For example, use the ReLU function to perform non-linear activation on the modulated coefficients, set the values less than zero to zero, and keep the values greater than zero unchanged.
[0075] In this embodiment, through hyperspectral data fusion and adaptive enhancement, the characteristics of semiconductor chip defects are effectively highlighted, improving the accuracy and reliability of defect detection. By constructing the noise distribution density map and the texture complexity map, the interference of background noise and irrelevant textures can be effectively suppressed, improving the robustness of defect detection. Through multi-scale directional filtering and wavelet decomposition, rich multi-scale multi-directional features are extracted, and the expression ability of the features is enhanced through non-linear mapping, which is beneficial to more accurately identify and locate defects.
[0076] In an alternative embodiment,
[0077] Reconstruct the initially enhanced image data through inverse wavelet transform; input the initially enhanced image data into a dual-path diffusion model, extract high-frequency detail features through the first path, extract low-frequency structure features through the second path, and adaptively fuse the high-frequency detail features and the low-frequency structure features to obtain a compensation feature map. Based on the compensation feature map, correct the illumination non-uniformity of the initially enhanced image data to obtain the corrected image data, including:
[0078] Construct a deep residual reconstruction network, input the enhanced coefficients into the deep residual reconstruction network, and reconstruct the initially enhanced image data through residual learning and skip connection structures;
[0079] Construct a dual-branch attention diffusion network, where the dual-branch attention diffusion network includes a spatial attention branch and a channel attention branch. Input the preliminarily enhanced image data into the dual-branch attention diffusion network, extract a high-frequency detail feature map through the spatial attention branch, and extract a low-frequency structural feature map through the channel attention branch;
[0080] Construct a hierarchical feature pyramid network, where the hierarchical feature pyramid network includes a feature mapping module, a feature reconstruction module, and a feature optimization module. Input the high-frequency detail feature map and the low-frequency structural feature map into the feature mapping module to generate a set of feature maps at multiple scales. Input the set of feature maps into the feature reconstruction module, calculate the similarity matrix between features at different scales by combining the self-attention mechanism and perform feature reconstruction based on the similarity matrix. Input the reconstructed feature map into the feature optimization module, and iteratively optimize the features through a recurrent neural network and a feedback mechanism to obtain an optimized compensated feature map;
[0081] Construct an adaptive illumination compensation network, where the adaptive illumination compensation network includes an illuminance analysis module, a reflectance analysis module, and an illumination equalization module. Input the compensated feature map into the illuminance analysis module, estimate the image illuminance component and the illuminance distribution map through a multi-scale convolutional neural network. Input the compensated feature map into the reflectance analysis module, restore the image reflectance component and the texture detail map through a depthwise separable convolutional network. Input the illuminance distribution map and the texture detail map into the illumination equalization module, and correct the non-uniformity of the illumination of the preliminarily enhanced image data through an adaptive exposure compensation algorithm and a local contrast enhancement algorithm, and output the corrected image data.
[0082] The hierarchical feature pyramid network is a deep learning network structure that constructs a feature pyramid through multi-level and different-scale feature maps, gradually extracts and fuses multi-level features in the image. The adaptive exposure compensation algorithm is an image processing algorithm used to automatically adjust the exposure parameters of an image to compensate for possible underexposure or overexposure problems during shooting. The local contrast enhancement algorithm is an image enhancement technique aimed at enhancing local details and contrast in an image.
[0083] Obtain low-quality image data to be processed, which may have problems such as noise, blur, and uneven illumination. Use inverse wavelet transform to preliminarily enhance the image data. Specifically, decompose the low-quality image data by wavelet transform to obtain sub-band coefficients in multiple different frequency bands. According to the requirements of image quality, enhance each sub-band coefficient. For example, enhance the high-frequency sub-band coefficients to improve image details, and enhance the low-frequency sub-band coefficients to improve the overall brightness of the image. Reconstruct the processed sub-band coefficients by inverse wavelet transform to obtain preliminarily enhanced image data. For example, perform three-level wavelet decomposition on a low-quality image with a resolution of 512x512, use the db4 wavelet basis function to obtain the low-frequency sub-band LL3 and high-frequency sub-bands LH3, HL3, HH3. Multiply the high-frequency sub-band coefficients by 1.2 for enhancement, multiply the low-frequency sub-band coefficients by 1.1 for enhancement, and then perform inverse wavelet transform reconstruction to obtain preliminarily enhanced image data.
[0084] Input the preliminarily enhanced image data into a dual-path diffusion model, which includes two paths: the first path is used to extract high-frequency detail features of the image, such as edges, textures, etc.; the second path is used to extract low-frequency structural features of the image, such as shapes, contours, etc. The first path can be implemented using a deep residual reconstruction network. Through residual learning and skip connection structures, it can effectively extract the high-frequency detail features of the image and suppress noise interference. The second path can be implemented using a dual-branch attention diffusion network. This network includes a spatial attention branch and a channel attention branch, which are used to focus on the spatial details and channel information of the image respectively, so as to extract low-frequency structural features. For example, assume that the size of the preliminarily enhanced image data is 256x256. The deep residual reconstruction network of the first path includes 3 residual blocks, and each residual block includes 2 convolutional layers with a kernel size of 3x3. The dual-branch attention diffusion network of the second path includes a spatial attention module and a channel attention module, which process the input features respectively.
[0085] The extracted high-frequency detailed feature map and low-frequency structural feature map are input into a hierarchical feature pyramid network, which includes a feature mapping module, a feature reconstruction module, and a feature optimization module. The feature mapping module maps the input feature map to multiple scales to form a group of feature maps. The feature reconstruction module combines the self-attention mechanism, calculates the similarity matrix between features at different scales, and performs feature reconstruction based on this matrix to fuse feature information at different scales. The feature optimization module iteratively optimizes the reconstructed features through a recurrent neural network and a feedback mechanism to obtain the final compensated feature map. For example, assume that the sizes of the high-frequency detailed feature map and the low-frequency structural feature map are both 128x128. They are respectively mapped to three scales of 64x64, 32x32, and 16x16 through the feature mapping module. The feature reconstruction module calculates the similarity matrix between features at different scales and performs feature reconstruction. The feature optimization module uses a recurrent neural network with the number of iterations set to 5 to optimize the features and obtain the final compensated feature map.
[0086] The compensated feature map is input into an adaptive illumination compensation network, which includes an illuminance analysis module, a reflectance analysis module, and an illumination equalization module. The illuminance analysis module estimates the illuminance component and illuminance distribution map of the image through a multi-scale convolutional neural network. The reflectance analysis module restores the reflectance component and texture detail map of the image through a depthwise separable convolutional network. The illumination equalization module inputs the illuminance distribution map and the texture detail map into an adaptive exposure compensation algorithm and a local contrast enhancement algorithm to correct the illumination non-uniformity of the preliminarily enhanced image data and outputs the final corrected image data. For example, assume that the size of the compensated feature map is 256x256. The illuminance analysis module uses a three-layer convolutional neural network with convolutional kernel sizes of 7x7, 5x5, and 3x3 respectively to estimate the image illuminance component and illuminance distribution map. The reflectance analysis module uses two layers of depthwise separable convolution with a convolutional kernel size of 3x3 to restore the image reflectance component and texture detail map. The illumination equalization module corrects the illumination non-uniformity of the preliminarily enhanced image data according to the illuminance distribution map and the texture detail map.
[0087] In this embodiment, the depth residual reconstruction network and the dual-branch attention diffusion network are used to effectively extract the details and structural features of the image, and the hierarchical feature pyramid network is used for feature fusion and optimization, thereby improving the clarity and detail expressiveness of the image. Through the adaptive illumination compensation network, the illuminance component and reflectance component of the image can be accurately estimated and the illumination equalization process can be performed, thereby effectively correcting the illumination non-uniformity problem of the image and making the image more natural and beautiful. By performing preliminary enhancement on the image through wavelet transform, noise and blur can be effectively suppressed, and the robustness of the image can be improved. The introduction of the adaptive illumination compensation network also makes this method more adaptable to images under different illumination conditions.
[0088] S2. Construct a dynamic hybrid feature enhancement network based on the neural architecture search method, add the corrected image data to the dynamic hybrid feature enhancement network and generate a multi-level feature pyramid, perform cross-layer attention operations on the feature maps in adjacent layers of the multi-level feature pyramid, generate inter-layer correlation weights, reconstruct and enhance the feature maps of each level, generate fused feature maps and add them to a pre-set multi-branch expert decision-making module, where the multi-branch expert decision-making module includes multiple sub-networks for different types of defects, each sub-network independently generates an independent detection result corresponding to the fused feature map, and combines the independent detection results by a dynamic integration strategy based on uncertainty to obtain an initial defect detection result;
[0089] The neural architecture search method is a technique for automatically designing neural network structures. By exploring different combinations of network structures, it automatically selects the optimal architecture parameters to improve the performance and efficiency of the model. The inter-layer correlation weight refers to the connection weight between different layers in a neural network. The cross-layer attention operation is a mechanism for establishing an attention mechanism between different layers of a neural network to enhance the mutual attention of features at different levels. The multi-branch expert decision-making module is a neural network module containing multiple independent branches, where each branch acts as an expert responsible for processing different types of information or tasks.
[0090] In an optional implementation manner,
[0091] Construct a dynamic hybrid feature enhancement network based on the neural architecture search method, add the corrected image data to the dynamic hybrid feature enhancement network and generate a multi-level feature pyramid, perform cross-layer attention operations on the feature maps in adjacent layers of the multi-level feature pyramid, generate inter-layer correlation weights, reconstruct and enhance the feature maps of each level, generate fused feature maps and add them to a pre-set multi-branch expert decision-making module, where the multi-branch expert decision-making module includes multiple sub-networks for different types of defects, each sub-network independently generates an independent detection result corresponding to the fused feature map, and combines the independent detection results by a dynamic integration strategy based on uncertainty to obtain an initial defect detection result, including:
[0092] Construct a dynamic hybrid feature enhancement network based on the neural architecture search method. The neural architecture search method selects and combines multiple convolutional layers, pooling layers, and activation function components from a predefined network structure space through a search algorithm, generates candidate networks, evaluates the performance of the candidate networks, and selects the network architecture with the optimal performance as the structure of the dynamic hybrid feature enhancement network;
[0093] Input the corrected image data into the dynamic hybrid feature enhancement network. Perform multiple downsampling operations through the encoder module in the dynamic hybrid feature enhancement network to convert the corrected image data into low-resolution features. Perform multiple upsampling operations through the decoder module in the dynamic hybrid feature enhancement network to restore the low-resolution features to the original resolution features. Extract feature maps of multiple different scales at different levels of the encoder module and the decoder module to construct a multi-level feature pyramid;
[0094] Perform cross-layer attention operations on the feature maps of adjacent layers in the multi-level feature pyramid. Map the corrected feature maps of adjacent layers to the same number of channels through convolutional operations. Calculate the query matrix and key matrix of the corrected feature maps in adjacent layers. Generate an attention weight matrix based on the query matrix and the key matrix. Reconstruct the corrected feature maps based on the attention weight matrix. Perform the cross-layer attention operation and feature reconstruction operation on the corrected feature maps in each layer of the multi-level feature pyramid to generate the fused pyramid feature maps;
[0095] Construct a multi-branch expert decision-making module. The multi-branch expert decision-making module includes multiple independent sub-networks. Each sub-network adopts an encoder-decoder structure and is equipped with a classifier. Input the fused pyramid feature maps into the multi-branch expert decision-making module. Detect different types of defects through multiple sub-networks respectively. Each sub-network outputs a corresponding independent detection result;
[0096] Obtain the confidence levels corresponding to the independent detection results output by multiple sub-networks. Based on the confidence levels, perform weighted average fusion on the independent detection results through a dynamic integration strategy to obtain the initial defect detection result.
[0097] The activation function component is a function module in a neural network used to activate neurons. The upsampling operation is a commonly used operation in signal processing or image processing, aiming to increase the resolution or size of data.
[0098] Construct a dynamic hybrid feature enhancement network using the neural architecture search method. Pre-define a network structure search space that includes various convolutional layers (e.g., different convolutional kernel sizes, different convolutional strides), pooling layers (e.g., max pooling, average pooling), and activation functions (e.g., ReLU, Sigmoid). Employ a search algorithm, such as an evolutionary algorithm or a reinforcement learning algorithm, to select and combine these components from the search space to generate multiple candidate network structures. For each candidate network structure, train it using an image dataset with defect annotations and evaluate its performance on the validation set, such as the mean average precision (mAP). Select the network architecture with the best performance as the final dynamic hybrid feature enhancement network. For example, the network structure finally selected by the search algorithm consists of three convolutional layers with convolutional kernel sizes of 3x3, 5x5, and 7x7 respectively, all with a stride of 1. After each convolutional layer, there is a ReLU activation function and a 2x2 max pooling layer.
[0099] Preprocess the collected image data that may contain defects, such as grayscaling, normalization, etc. The corrected image data is input into the constructed dynamic hybrid feature enhancement network. This network includes an encoder module and a decoder module. The encoder module converts the input image into a low-resolution feature map through multiple downsampling operations. For example, if the input image size is 256x256, it becomes 16x16 after passing through the encoder. The decoder module then restores the low-resolution feature map to the original resolution through multiple upsampling operations. For example, it restores the 16x16 feature map to 256x256. At different levels of the encoder and decoder, extract multiple feature maps of different scales to construct a multi-level feature pyramid. For example, extract the feature maps from the third layer of the encoder and the second layer of the decoder to form a two-layer feature pyramid.
[0100] Perform cross-layer attention operations on the feature maps of adjacent layers in the multi-level feature pyramid. Taking the two-layer feature pyramid as an example, map the feature maps from the third layer of the encoder and the second layer of the decoder to the same number of channels through convolutional operations, such as mapping them to 64 channels. Calculate the query matrix and the key matrix of these two feature maps respectively. Based on the query matrix and the key matrix, calculate the attention weight matrix. Use the attention weight matrix to perform weighted reconstruction on the feature map to obtain the enhanced feature map. Perform cross-layer attention operations and feature reconstruction operations on each layer in the pyramid, and finally generate the fused pyramid feature map.
[0101] Construct a multi-branch expert decision-making module, which contains multiple independent sub-networks. Each sub-network adopts an encoder-decoder structure and is equipped with a classifier. Input the fused pyramid feature map into the multi-branch expert decision-making module. Each sub-network focuses on detecting different types of defects, such as scratches, cracks, stains, etc. Each sub-network independently processes the input feature map and outputs corresponding independent detection results, such as the location and category of the defect. Suppose there are three sub-networks, which are used to detect scratches, cracks, and stains respectively.
[0102] Obtain the confidence corresponding to the independent detection results output by each sub-network. Based on the confidence, adopt a dynamic integration strategy, such as weighted average based on uncertainty, to fuse the independent detection results. Detection results with high confidence will be given greater weights. Fuse the independent detection results of multiple sub-networks by weighted average to obtain the final defect detection result. For example, if the confidence of the sub-network 1 detecting a scratch is 0.9, the confidence of the sub-network 2 detecting a crack is 0.6, and the confidence of the sub-network 3 detecting a stain is 0.3, then the final result is more inclined to the scratch.
[0103] In this embodiment, by performing neural architecture search to find the optimal network structure, and combining a dynamic hybrid feature enhancement network and a cross-layer attention mechanism, it is possible to effectively extract and fuse feature information of different scales, improve the accuracy of defect detection. Especially in the case of complex backgrounds and tiny defects, the multi-branch expert decision-making module and the dynamic integration strategy can effectively handle different types of defects and reduce the impact of misjudgments of a single sub-network on the final result, enhancing the robustness and generalization ability of the model. Through neural architecture search, the network structure can be optimized, reducing the number of model parameters and computational complexity, improving the efficiency of defect detection, and meeting the real-time requirements.
[0104] In an alternative embodiment,
[0105] Perform cross-layer attention operation on the feature maps of adjacent layers in the multi-level feature pyramid as shown in the following formula:
[0106] ;
[0107] where, A represents the element of the calculated attention weight matrix, represents the k-th element of the query vector at the m-th position in the feature map of the i-th layer, represents the weight matrix of the query matrix of the i-th layer, which converts the query vector from dimension d h to dimension d l dimension, represents the k-th element of the key vector at the n-th position in the feature map of the j-th layer, represents the weight matrix of the key matrix of the i-th layer, which converts the key vector from dimension d h dimension to dl Dimension, d h Represents the original dimension of the query vector and the key vector, d l Represents the dimension after the query vector and the key vector are linearly transformed. h represents the number of heads in the multi-head attention mechanism. Represents the attention weight between the m-th position in the feature map of the i-th layer and the n-th position in the feature map of the j-th layer in the g-th attention head. Represents the weight matrix of the output matrix. O is used to indicate the output.
[0108] In this embodiment, by fusing the features of adjacent layers, the context information in the image can be better captured, thereby improving the expression ability of the features. The attention mechanism can make the model pay more attention to important features and ignore irrelevant features, thereby improving the robustness of the model. The cross-layer attention mechanism can promote the information flow between features of different levels, so that the detailed information of low-level features can be effectively transmitted to high-level features, and the semantic information of high-level features can also be effectively transmitted to low-level features.
[0109] S3. Establish a spatio-temporal knowledge inference graph based on the initial defect detection results. Set the initial defect detection results as graph nodes, calculate the morphological similarity and spatial position relationship between each initial defect detection result and construct the edge connection strength. Based on the edge connection strength, perform hierarchical clustering on the initial defect detection results to generate similar defect groups. Perform graph attention network operations on the spatio-temporal knowledge inference graph to generate defect correlation feature vectors and perform temporal correlation analysis on the defect correlation feature vectors and the process parameter sequence. Construct a process parameter influence model. Based on the process parameter influence model and the similar defect groups, perform reliability evaluation and optimization on the initial defect detection results to obtain the final defect detection results.
[0110] The morphological similarity is a measurement method for measuring the similarity between different morphological features, usually used in fields such as image analysis and pattern recognition. By calculating the shape similarity of different objects or graphics, it is used for classification, matching, or clustering. The hierarchical clustering is an unsupervised learning algorithm used to divide data into different hierarchical structures. By calculating the similarity between data points, hierarchical clustering can generate a tree-like structure (also called a dendrogram) representing the hierarchical relationship of the data, and is often used in data analysis and pattern recognition tasks. The similar defect group refers to grouping different defects into the same group based on the similarity of defects in defect detection or classification tasks. The temporal correlation analysis is a technique for analyzing the mutual relationship between different variables in time series data, usually by calculating the correlation coefficient of time series data or using other statistical methods to evaluate the dependence and temporal relationship between different data points.
[0111] In an alternative embodiment,
[0112] Based on the initial defect detection results, establish a spatio-temporal knowledge inference graph. Set the initial defect detection results as graph nodes, calculate the morphological similarity and spatial position relationship between each initial defect detection result, and construct the edge connection strength. Perform hierarchical clustering on the initial defect detection results based on the edge connection strength to generate similar defect groups. Perform graph attention network operations on the spatio-temporal knowledge inference graph to generate defect correlation feature vectors, and perform temporal correlation analysis on the defect correlation feature vectors and the process parameter sequence to construct a process parameter influence model. Based on the process parameter influence model and the similar defect groups, perform reliability evaluation and optimization on the initial defect detection results to obtain the final defect detection results, including:
[0113] Obtain the initial defect detection results of the chip image. The initial defect detection results include the morphological features and spatial position coordinates of each defect. The morphological features include length, width, and area, and the spatial position coordinates include the row number and column number of the center point of the defect.
[0114] Set each defect in the initial defect detection results as a graph node, and establish a spatio-temporal knowledge inference graph. Each graph node includes the morphological features and spatial position coordinates of the corresponding defect. For each pair of graph nodes in the spatio-temporal knowledge inference graph, calculate the morphological similarity by calculating the cosine similarity between the morphological features of different graph nodes, calculate the spatial position distance by calculating the Euclidean distance between the spatial position coordinates of different graph nodes, and construct the edge connection strength based on the morphological similarity and the spatial position distance. If the morphological similarity between two graph nodes is greater than a pre-set similarity threshold and the spatial position distance is less than a pre-set distance threshold, calculate the edge connection strength based on the spatial position distance.
[0115] Based on the edge connection strength, cluster the graph nodes by a bottom-up hierarchical clustering method. Initially, set each graph node as an independent cluster, iteratively merge the two clusters with the maximum edge connection strength, and repeat the merging until the edge connection strength between all clusters is less than a pre-set connection strength threshold, generating multiple similar defect groups.
[0116] Construct a graph attention network, which includes an attention layer and a feature propagation layer. For the spatio-temporal knowledge inference graph, calculate the attention weights between each graph node and its corresponding adjacent nodes through the attention layer. The attention weights are determined based on the similarity degree of node features. Based on the attention weights, update the features of the graph nodes through the feature propagation layer, so that the features of each graph node aggregate the weighted information of the corresponding adjacent nodes, generating a node feature matrix.
[0117] Extract the defect correlation feature vectors corresponding to each of the similar defect groups in the node feature matrix, obtain the process parameter sequences within the time periods corresponding to the similar defect groups, where the process parameter sequences include temperature parameters and pressure parameters, calculate the cross-correlation coefficients between the defect correlation feature vectors and the process parameter sequences through the Granger causality test method. At the same time, construct a process parameter influence model based on the cross-correlation coefficients, input each defect in the initial defect detection result into the process parameter influence model, determine the corresponding similar defect group based on the morphological features in the initial defect detection result, and judge whether the current defect is caused by process anomalies according to the prediction results output by the process parameter influence model. Mark the defects determined to be caused by process anomalies as reliable defects, filter out the defects not caused by process anomalies, and obtain the final defect detection result.
[0118] The graph node is one of the basic elements in graph theory, representing a data point or entity in a graph. In a graph model, nodes usually represent objects, events, or other entities, and are connected by edges. The edge connection strength is an attribute of the edges in the graph, used to measure the tightness or strength of the connection between two nodes. The edge connection strength can reflect the relationship strength between nodes and is usually used in tasks such as network analysis and social network analysis. The independent cluster refers to the situation where there are no direct connections or relationships between nodes or data points in a graph or dataset. The Granger causality test method is a statistical method used to determine whether one time series can predict another time series.
[0119] Obtain the initial defect detection result of the chip image. Use a defect detection device, such as an optical microscope or a scanning electron microscope, to scan the chip image and identify and extract the defect information in the image. This information includes the morphological features and spatial position coordinates of each defect. The morphological features include the length, width, and area of the defect, which can be calculated using image processing techniques, such as edge detection and region segmentation algorithms. The spatial position coordinates record the row number and column number of the defect center point in the image. For example, in a chip image, three defects are detected. The length of defect 1 is 10 microns, the width is 5 microns, the area is 50 square microns, and the center point coordinates are (100, 200); the length of defect 2 is 8 microns, the width is 6 microns, the area is 48 square microns, and the center point coordinates are (150, 250); the length of defect 3 is 12 microns, the width is 4 microns, the area is 48 square microns, and the center point coordinates are (300, 400).
[0120] Build a spatio-temporal knowledge inference graph. Each detected defect is regarded as a graph node to form a spatio-temporal knowledge inference graph. Each graph node contains the morphological features (length, width, area) and spatial position coordinates (row number, column number) of the corresponding defect. For example, according to three detected defects, a spatio-temporal knowledge inference graph containing three nodes is built.
[0121] Calculate the edge connection strength. For each pair of nodes in the graph, calculate the morphological similarity and spatial position distance between them. The morphological similarity is measured by calculating the cosine similarity between the morphological feature vectors (length, width, area) of the two nodes. The spatial position distance is measured by calculating the Euclidean distance between the center point coordinates of the two nodes. If the morphological similarity between two nodes is greater than a preset similarity threshold (e.g., 0.9), and the spatial position distance is less than a preset distance threshold (e.g., 100 microns), then calculate the edge connection strength according to the spatial position distance. The smaller the spatial position distance, the greater the edge connection strength. For example, the morphological similarity between node 1 and node 2 is 0.95, and the spatial position distance is 50 microns, which meets the conditions, and the calculated edge connection strength is 0.9; the morphological similarity between node 1 and node 3 is 0.8, and the spatial position distance is 200 microns, which does not meet the conditions, so there is no edge connection.
[0122] Perform hierarchical clustering. According to the calculated edge connection strength, use the bottom-up hierarchical clustering method to cluster the graph nodes. In the initial state, each node is regarded as an independent cluster. Then iteratively merge the two clusters with the maximum edge connection strength. Repeat the merging process until the edge connection strength between all clusters is less than a preset connection strength threshold (e.g., 0.5). Finally, the nodes in the graph are divided into multiple groups of similar defects. For example, the edge connection strength between node 1 and node 2 is 0.9, which is greater than the threshold 0.5, so node 1 and node 2 are merged into a cluster.
[0123] Perform graph attention network operations. Construct a graph attention network, which includes an attention layer and a feature propagation layer. Input the spatio-temporal knowledge inference graph into the graph attention network. The attention layer calculates the attention weights between each node and its adjacent nodes. The attention weights are determined based on the similarity degree of node features. The feature propagation layer updates the features of each node according to the attention weights, so that the features of each node aggregate the weighted information of its adjacent nodes to generate a node feature matrix.
[0124] Build a process parameter influence model. Extract the defect correlation feature vectors corresponding to each similar defect group in the node feature matrix. Obtain the process parameter sequences during the corresponding time period of the similar defect groups, including temperature parameters and pressure parameters. Use the Granger causality test method to calculate the cross-correlation coefficients between the defect correlation feature vectors and the process parameter sequences. Build a process parameter influence model based on the cross-correlation coefficients. For example, assume that the defect correlation feature vector of a certain similar defect group is [0.1, 0.2, 0.3], the temperature parameter sequence during the corresponding time period is [200, 210, 220], and the pressure parameter sequence is [1.0, 1.1, 1.2]. The cross-correlation coefficient between the temperature parameter and the defect correlation feature vector calculated by the Granger causality test method is 0.8, and the cross-correlation coefficient between the pressure parameter and the defect correlation feature vector is 0.6. Then build a process parameter influence model, which reflects the influence degrees of temperature and pressure on defects.
[0125] Perform defect reliability evaluation and optimization. Input each defect in the initial defect detection results into the process parameter influence model. Determine the similar defect group to which the defect belongs according to its morphological features. Judge whether the defect is caused by process abnormalities according to the prediction results of the process parameter influence model. Mark the defects determined to be caused by process abnormalities as reliable defects, filter out the defects not caused by process abnormalities, and obtain the final defect detection results. For example, input defect 1 into the process parameter influence model, and the model predicts that this defect is caused by temperature abnormality, then mark defect 1 as a reliable defect.
[0126] In this embodiment, by combining the morphological features, spatial position relationships, and process parameters of defects, defects caused by process abnormalities can be identified more accurately, misdetections not caused by process abnormalities can be effectively filtered out, thereby improving the accuracy of defect detection. By grouping similar defects through hierarchical clustering, the efficiency of defect classification can be effectively improved, which is convenient for targeted analysis and processing of different types of defects. By building a process parameter influence model, the correlation relationship between process parameters and defects can be analyzed, providing guidance for optimizing the process flow and reducing the occurrence of defects.
[0127] In an alternative embodiment,
[0128] Calculating the cross-correlation coefficients between the defect correlation feature vectors and the process parameter sequences by the Granger causality test method includes:
[0129] Obtain the process parameter sequences and defect correlation feature vectors in the chip manufacturing process, divide the process parameter sequences into a training data set and a test data set in chronological order, where the training data set is used to build a prediction model, and the test data set is used to evaluate the performance of the prediction model;
[0130] Construct a first prediction model based on the training data set, use the historical data of the process parameter sequence within a preset time window to predict the process parameter value at the current moment, construct a second prediction model based on the training data set, and simultaneously use the historical data of the process parameter sequence and the defect correlation feature vector within a preset time window to predict the process parameter value at the current moment;
[0131] Input the test data set into the first prediction model and the second prediction model, calculate the sum of squared residuals of the two prediction models by summing the squares of the differences between the predicted values and the actual values, and construct a Granger causality test statistic, which is calculated based on the sum of squared residuals of the first prediction model, the sum of squared residuals of the second prediction model, the number of parameters of the first prediction model, the number of parameters of the second prediction model, and the number of samples in the test data set;
[0132] Determine the critical value based on a preset significance level. When the Granger causality test statistic is greater than the critical value, it is determined that the defect correlation feature vector is the Granger cause of the process parameter sequence, and calculate the cross-correlation coefficient between each feature in the defect correlation feature vector and the process parameter sequence at different time lags based on the covariance and variance between the defect correlation feature vector and the process parameter sequence.
[0133] The Granger causality test statistic is a statistic that measures the strength of the causal relationship between time series, used to decide whether to reject the null hypothesis (i.e., no causal relationship) and further quantify the strength of the causal relationship. The Granger cause refers to a variable or time series that can provide causal prediction ability in the Granger causality test. The time lag refers to the degree of influence of the value of a certain variable at a certain moment in a time series on the value of another variable at a subsequent moment. The cross-correlation coefficient is a statistic that measures the similarity between two time series. By calculating the cross-correlation coefficient, the correlation between two time series at different time lags can be analyzed, and it is commonly used in fields such as signal processing and system modeling.
[0134] Collect process parameter data and defect-related feature vector data during the chip manufacturing process. The process parameter data includes, for example, temperature, pressure, deposition rate, etc., and the defect correlation feature vector can include information such as the type, size, and location of the defect, and needs to be cleaned and preprocessed, such as removing outliers and missing values.
[0135] Divide the collected process parameter sequence into a training data set and a test data set in chronological order. For example, use the first 80% of the data as the training data set, and the remaining 20% as the test data set. The training data set is used to construct the prediction model, and the test data set is used to evaluate the performance of the model.
[0136] Two prediction models are constructed. The first prediction model uses only the historical data of the process parameter sequence to predict the process parameter value at the current moment. For example, using the temperature data of the past 10 time points to predict the temperature at the current moment, methods such as autoregressive models or recurrent neural networks can be adopted. The second prediction model uses both the historical data of the process parameter sequence and the defect-related feature vector within a preset time window to predict the process parameter value at the current moment. For example, using the temperature data of the past 10 time points and the defect size data at the corresponding time points to predict the temperature at the current moment. This model can also adopt a prediction method similar to that of the first model.
[0137] The test data set is input into these two prediction models, and the sum of the squares of the differences between the predicted values and the actual values is calculated respectively to obtain the residual sum of squares of the two prediction models. Suppose the residual sum of squares of the first model on the test data set is 100, and that of the second model is 80.
[0138] Using the residual sum of squares of the two models, the number of parameters of the two models, and the number of samples in the test data set, a Granger causality test statistic is constructed. This statistic can reflect the degree of improvement in the prediction accuracy of the defect-related feature vector for the process parameter sequence. For example, if the number of parameters of the first model is 5, the number of parameters of the second model is 7, and the number of samples in the test data set is 1000, the Granger causality test statistic can be calculated according to a specific formula.
[0139] According to the preset significance level (e.g., 0.05), the critical value is determined, which can be obtained by looking up tables or using statistical software. If the calculated Granger causality test statistic is greater than the critical value, it is determined that the defect-related feature vector is the Granger cause of the process parameter sequence. For example, if the calculated Granger causality test statistic is 10 and the critical value is 5, it can be determined that there is a Granger causal relationship between the defect-related feature vector and the process parameter sequence.
[0140] If the defect-related feature vector is determined to be the Granger cause of the process parameter sequence, the cross-correlation coefficients between each feature in the defect-related feature vector and the process parameter sequence at different time lags are further calculated. For example, calculating the cross-correlation coefficients between the defect size and the temperature at lags of 1 time point, 2 time points, etc., can be obtained by calculating the covariance and variances of the defect size and the temperature respectively, and then calculating the correlation coefficients, which helps to understand how the defect-related feature vector affects the process parameter sequence and the time lag at which the influence occurs.
[0141] In this embodiment, by identifying the defect-related feature vectors that have Granger causality with the process parameter sequence, the root cause of the defect can be quickly located, thereby improving the efficiency of defect diagnosis. It can help engineers understand which process parameters have the greatest impact on product quality, so as to more targeted optimize the process parameter control strategy, reduce the occurrence of defects. By effectively controlling the key process parameters, the yield of chip products can ultimately be improved and the production cost can be reduced.
[0142] Figure 2 FIG. is a schematic structural diagram of an automatic semiconductor chip defect detection system based on hyperspectral imaging according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0143] A first unit for collecting hyperspectral image data of a semiconductor chip to be detected, adding the hyperspectral image data to an improved neuromorphic adaptive processing network and calculating a noise distribution density map and a texture complexity map of a local area, constructing an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculating wavelet decomposition coefficients of different scales based on the adaptive enhancement matrix, performing a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, and reconstructing through inverse wavelet transform to obtain preliminarily enhanced image data; inputting the preliminarily enhanced image data into a dual-path diffusion model, extracting high-frequency detail features through a first path, extracting low-frequency structure features through a second path, and adaptively fusing the high-frequency detail features and the low-frequency structure features to obtain a compensated feature map, and performing illumination non-uniformity correction on the preliminarily enhanced image data based on the compensated feature map to obtain corrected image data;
[0144] A second unit for constructing a dynamic hybrid feature enhancement network based on a neural architecture search method, adding the corrected image data to the dynamic hybrid feature enhancement network and generating a multi-level feature pyramid, performing cross-layer attention operations on the feature maps in adjacent layers in the multi-level feature pyramid, generating inter-layer correlation weights and reconstructing and enhancing the feature maps of each level, generating a fused feature map and adding it to a preset multi-branch expert decision module, wherein the multi-branch expert decision module includes multiple sub-networks for different types of defects, each sub-network independently generates an independent detection result corresponding to the fused feature map, and fuses the independent detection results by combining a dynamic integration strategy based on uncertainty to obtain an initial defect detection result;
[0145] A third unit is used to establish a spatio-temporal knowledge inference graph based on the initial defect detection results. Set the initial defect detection results as graph nodes, calculate the morphological similarity and spatial position relationship between each initial defect detection result, construct the edge connection strength, perform hierarchical clustering on the initial defect detection results based on the edge connection strength to generate similar defect groups, perform graph attention network operations on the spatio-temporal knowledge inference graph, generate defect association feature vectors, perform temporal correlation analysis on the defect association feature vectors and the process parameter sequence, construct a process parameter influence model, and perform reliability evaluation and optimization on the initial defect detection results based on the process parameter influence model and the similar defect groups to obtain the final defect detection results.
[0146] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: a processor;
[0147] a memory for storing instructions executable by the processor;
[0148] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0149] In a fourth aspect of the embodiments of the present invention,
[0150] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0151] The present invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present invention.
[0152] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An automatic defect detection method for semiconductor chips based on hyperspectral imaging, characterized in that, Including: Collecting hyperspectral image data of a semiconductor chip to be detected, adding the hyperspectral image data to an improved neuromorphic adaptive processing network and calculating a noise distribution density map and a texture complexity map of a local area, constructing an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculating wavelet decomposition coefficients of different scales based on the adaptive enhancement matrix, performing a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, and reconstructing through inverse wavelet transform to obtain preliminarily enhanced image data; Inputting the preliminarily enhanced image data into a dual-path diffusion model, extracting high-frequency detail features through a first path, extracting low-frequency structure features through a second path, adaptively fusing the high-frequency detail features and the low-frequency structure features to obtain a compensation feature map, and correcting the uneven illumination of the preliminarily enhanced image data based on the compensation feature map to obtain corrected image data; Constructing a dynamic hybrid feature enhancement network based on a neural architecture search method, adding the corrected image data to the dynamic hybrid feature enhancement network and generating a multi-level feature pyramid, performing cross-layer attention operations on feature maps in adjacent layers in the multi-level feature pyramid, generating inter-layer association weights and reconstructing and enhancing feature maps of each level, generating a fused feature map and adding it to a pre-set multi-branch expert decision module, where the multi-branch expert decision module includes multiple sub-networks for different types of defects, each sub-network independently generates an independent detection result corresponding to the fused feature map, and fusing the independent detection results by combining a dynamic integration strategy based on uncertainty to obtain an initial defect detection result; Establishing a spatio-temporal knowledge inference graph based on the initial defect detection result, setting the initial defect detection result as a graph node, calculating the morphological similarity and spatial position relationship between each initial defect detection result and constructing an edge connection strength, hierarchically clustering the initial defect detection results based on the edge connection strength to generate similar defect groups, performing graph attention network operations on the spatio-temporal knowledge inference graph, generating a defect association feature vector and performing temporal correlation analysis on the defect association feature vector and a process parameter sequence, constructing a process parameter influence model, and performing reliability evaluation and optimization on the initial defect detection result based on the process parameter influence model and the similar defect groups to obtain a final defect detection result.
2. The method according to claim 1, wherein Collecting hyperspectral image data of a semiconductor chip to be detected, adding the hyperspectral image data to an improved neuromorphic adaptive processing network and calculating a noise distribution density map and a texture complexity map of a local area, constructing an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculating wavelet decomposition coefficients of different scales based on the adaptive enhancement matrix, and performing a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, including: Collecting hyperspectral image data of a semiconductor chip to be detected, where the hyperspectral image data includes image data in the visible light band, near-infrared band, and short-wave infrared band; Input the hyperspectral image data into an improved neuromorphic adaptive processing network, calculate the local statistical features of the hyperspectral image data at different scales, extract the noise distribution feature vector and the texture feature vector, construct a multi-layer neural network based on the noise distribution feature vector and calculate the noise distribution density map, construct a convolutional neural network based on the texture feature vector and calculate the texture complexity map, and input the noise distribution density map and the texture complexity map into the attention mechanism module to generate an attention weight map; According to the attention weight map, perform feature reconstruction on the noise distribution density map and the texture complexity map, fuse the reconstructed feature maps to construct an adaptive enhancement matrix, perform adaptive enhancement processing on different bands of the hyperspectral image data based on the adaptive enhancement matrix, construct a multi-scale direction filter bank and a wavelet decomposition unit, and perform direction decomposition, scale decomposition, and wavelet decomposition on the enhanced hyperspectral image data simultaneously. Obtain multiple direction sub-bands and scale sub-bands through the multi-scale direction filter bank, and obtain an approximation coefficient sub-band and a detail coefficient sub-band through the wavelet decomposition unit; Perform adaptive threshold segmentation on the direction sub-band, the scale sub-band, the approximation coefficient sub-band, and the detail coefficient sub-band, extract significant feature coefficients, perform adaptive weight modulation on the significant feature coefficients and the adaptive enhancement matrix, and perform a non-linear activation function operation on the modulated coefficients to generate enhancement coefficients.
3. The method according to claim 1, wherein Reconstruct the initially enhanced image data through inverse wavelet transform; input the initially enhanced image data into a dual-path diffusion model, extract high-frequency detail features through the first path, extract low-frequency structural features through the second path, and adaptively fuse the high-frequency detail features and the low-frequency structural features to obtain a compensation feature map. Based on the compensation feature map, correct the uneven illumination of the initially enhanced image data to obtain the corrected image data, including: Construct a deep residual reconstruction network, input the enhancement coefficients into the deep residual reconstruction network, and reconstruct the initially enhanced image data through residual learning and skip connection structures; Construct a dual-branch attention diffusion network, which includes a spatial attention branch and a channel attention branch. Input the initially enhanced image data into the dual-branch attention diffusion network, extract a high-frequency detail feature map through the spatial attention branch, and extract a low-frequency structural feature map through the channel attention branch; Construct a hierarchical feature pyramid network, which includes a feature mapping module, a feature reconstruction module, and a feature optimization module. Input the high-frequency detail feature map and the low-frequency structural feature map into the feature mapping module to generate a set of feature maps at multiple scales. Input the set of feature maps into the feature reconstruction module, calculate the similarity matrix between features at different scales by combining the self-attention mechanism and perform feature reconstruction based on the similarity matrix. Input the reconstructed feature maps into the feature optimization module, and iteratively optimize the features through a recurrent neural network and a feedback mechanism to obtain an optimized compensation feature map; Construct an adaptive illumination compensation network. The adaptive illumination compensation network includes an illuminance analysis module, a reflectance analysis module, and an illumination equalization module. Input the compensation feature map into the illuminance analysis module, estimate the image illuminance component and the illuminance distribution map through a multi-scale convolutional neural network. Input the compensation feature map into the reflectance analysis module, and restore the image reflectance component and the texture detail map through a depthwise separable convolutional network. Input the illuminance distribution map and the texture detail map into the illumination equalization module, and perform illumination non-uniformity correction on the preliminarily enhanced image data through an adaptive exposure compensation algorithm and a local contrast enhancement algorithm, and output the corrected image data.
4. The method according to claim 1, wherein Construct a dynamic hybrid feature enhancement network based on the neural architecture search method. Add the corrected image data to the dynamic hybrid feature enhancement network and generate a multi-level feature pyramid. Perform cross-layer attention operations on the feature maps in adjacent layers of the multi-level feature pyramid, generate inter-layer correlation weights, and reconstruct and enhance the feature maps at each level to generate a fused feature map and add it to a pre-set multi-branch expert decision module. Among them, the multi-branch expert decision module contains multiple sub-networks for different types of defects. Each sub-network independently generates an independent detection result corresponding to the fused feature map, and combines the independent detection results through a dynamic integration strategy based on uncertainty to obtain the initial defect detection result, including: Construct a dynamic hybrid feature enhancement network based on the neural architecture search method. The neural architecture search method selects and combines various convolutional layers, pooling layers, and activation function components from a predefined network structure space through a search algorithm, generates candidate networks, evaluates the performance of the candidate networks, and selects the network architecture with the best performance as the structure of the dynamic hybrid feature enhancement network; Input the corrected image data into the dynamic hybrid feature enhancement network. Perform multiple downsampling operations through the encoder module in the dynamic hybrid feature enhancement network to convert the corrected image data into low-resolution features. Perform multiple upsampling operations through the decoder module in the dynamic hybrid feature enhancement network to restore the low-resolution features to the original resolution features. Extract multiple feature maps of different scales at different levels of the encoder module and the decoder module to construct a multi-level feature pyramid; Perform cross-layer attention operations on the feature maps in adjacent layers of the multi-level feature pyramid. Map the corrected feature maps in adjacent layers to the same number of channels through convolutional operations, calculate the query matrix and the key matrix of the corrected feature maps in adjacent layers, generate an attention weight matrix based on the query matrix and the key matrix, reconstruct the corrected feature maps based on the attention weight matrix, and perform the cross-layer attention operation and the feature reconstruction operation on the corrected feature maps in each layer of the multi-level feature pyramid to generate a fused pyramid feature map; Construct a multi-branch expert decision-making module, where the multi-branch expert decision-making module includes multiple independent sub-networks. Each sub-network adopts an encoder-decoder structure and is equipped with a classifier. Input the fused pyramid feature map into the multi-branch expert decision-making module, and detect different types of defects through multiple sub-networks respectively. Each sub-network outputs a corresponding independent detection result; Obtain the confidence levels corresponding to the independent detection results output by multiple sub-networks. Based on the confidence levels, perform weighted average fusion on the independent detection results through a dynamic integration strategy to obtain an initial defect detection result.
5. The method according to claim 1, wherein Based on the initial defect detection result, establish a spatio-temporal knowledge inference graph. Set the initial defect detection result as a graph node, calculate the morphological similarity and spatial position relationship between each initial defect detection result and construct the edge connection strength. Perform hierarchical clustering on the initial defect detection result based on the edge connection strength to generate similar defect groups. Perform graph attention network operations on the spatio-temporal knowledge inference graph to generate a defect correlation feature vector and perform temporal correlation analysis on the defect correlation feature vector and the process parameter sequence. Construct a process parameter influence model. Based on the process parameter influence model and the similar defect groups, perform reliability evaluation and optimization on the initial defect detection result to obtain the final defect detection result, including: Obtain the initial defect detection result of the chip image. The initial defect detection result contains the morphological features and spatial position coordinates of each defect. The morphological features include length, width, and area. The spatial position coordinates include the row number and column number of the center point of the defect; Set each defect in the initial defect detection result as a graph node and establish a spatio-temporal knowledge inference graph. Each graph node contains the morphological features and spatial position coordinates of the corresponding defect. For each pair of graph nodes in the spatio-temporal knowledge inference graph, obtain the morphological similarity by calculating the cosine similarity between the morphological features of different graph nodes, and obtain the spatial position distance by calculating the Euclidean distance between the spatial position coordinates of different graph nodes. Construct the edge connection strength based on the morphological similarity and the spatial position distance. If the morphological similarity between two graph nodes is greater than a pre-set similarity threshold and the spatial position distance is less than a pre-set distance threshold, calculate the edge connection strength based on the spatial position distance; Based on the edge connection strength, perform clustering on the graph nodes through a bottom-up hierarchical clustering method. Initially, set each graph node as an independent cluster, iteratively merge the two clusters with the maximum edge connection strength, and repeat the merging until the edge connection strength between all clusters is less than a pre-set connection strength threshold, generating multiple similar defect groups; Construct a graph attention network, which includes an attention layer and a feature propagation layer. For the spatio-temporal knowledge inference graph, calculate the attention weights between each graph node and its corresponding adjacent nodes of the current graph node through the attention layer. The attention weights are determined based on the similarity degree of node features. Based on the attention weights, update the features of the graph nodes through the feature propagation layer, so that the features of each graph node aggregate the weighted information of the corresponding adjacent nodes, and generate a node feature matrix; Extract the defect correlation feature vectors corresponding to each similar defect group in the node feature matrix, obtain the process parameter sequence within the corresponding time period of the similar defect group. The process parameter sequence includes temperature parameters and pressure parameters. Calculate the cross-correlation coefficient between the defect correlation feature vector and the process parameter sequence through the Granger causality test method. At the same time, construct a process parameter influence model according to the cross-correlation coefficient. Input each defect in the initial defect detection result into the process parameter influence model. Determine the corresponding similar defect group based on the morphological features in the initial defect detection result. Judge whether the current defect is caused by process abnormality according to the prediction result output by the process parameter influence model. Mark the defect determined to be caused by process abnormality as a reliable defect, filter out the defects not caused by process abnormality, and obtain the final defect detection result.
6. The method according to claim 5, wherein Calculating the cross-correlation coefficient between the defect correlation feature vector and the process parameter sequence through the Granger causality test method includes: Obtain the process parameter sequence and the defect correlation feature vector in the chip manufacturing process. Divide the process parameter sequence into a training data set and a test data set in chronological order. The training data set is used to construct a prediction model, and the test data set is used to evaluate the performance of the prediction model; Construct a first prediction model based on the training data set, and use the historical data of the process parameter sequence within a preset time window to predict the process parameter value at the current moment. Construct a second prediction model based on the training data set, and at the same time use the historical data of the process parameter sequence and the defect correlation feature vector within a preset time window to predict the process parameter value at the current moment; Input the test data set into the first prediction model and the second prediction model, calculate the sum of the squared residuals of the two prediction models by summing the squares of the differences between the predicted values and the actual values, and construct a Granger causality test statistic. The Granger causality test statistic is calculated based on the sum of the squared residuals of the first prediction model, the sum of the squared residuals of the second prediction model, the number of parameters of the first prediction model, the number of parameters of the second prediction model, and the number of samples in the test data set; Determine a critical value based on a preset significance level. When the Granger causality test statistic is greater than the critical value, it is determined that the defect-related feature vector is the Granger cause of the process parameter sequence. Calculate the cross-correlation coefficient between each feature in the defect-related feature vector and the process parameter sequence at different time lags respectively based on the covariance and variance of the defect-related feature vector and the process parameter sequence.
7. An automatic semiconductor chip defect detection system based on hyperspectral imaging, for implementing the method according to any one of the preceding claims 1-6, characterized in that, Comprising: A first unit for collecting hyperspectral image data of a semiconductor chip to be detected, adding the hyperspectral image data to an improved neuromorphic adaptive processing network and calculating a noise distribution density map and a texture complexity map of a local area, constructing an adaptive enhancement matrix based on the noise distribution density map and the texture complexity map, calculating wavelet decomposition coefficients of different scales based on the adaptive enhancement matrix, performing a non-linear mapping on the wavelet decomposition coefficients to obtain enhancement coefficients, and reconstructing through inverse wavelet transform to obtain preliminarily enhanced image data; Input the preliminarily enhanced image data into a dual-path diffusion model, extract high-frequency detail features through a first path, extract low-frequency structure features through a second path, and adaptively fuse the high-frequency detail features and the low-frequency structure features to obtain a compensation feature map, and perform illumination non-uniformity correction on the preliminarily enhanced image data based on the compensation feature map to obtain corrected image data; A second unit for constructing a dynamic hybrid feature enhancement network based on a neural architecture search method, adding the corrected image data to the dynamic hybrid feature enhancement network and generating a multi-level feature pyramid, performing cross-layer attention operations on the feature maps in adjacent layers in the multi-level feature pyramid, generating inter-layer association weights and reconstructing and enhancing the feature maps of each level, generating a fused feature map and adding it to a pre-set multi-branch expert decision module, where the multi-branch expert decision module includes multiple sub-networks for different types of defects, each sub-network independently generates an independent detection result corresponding to the fused feature map, and fuses the independent detection results in combination with a dynamic integration strategy based on uncertainty to obtain an initial defect detection result; A third unit for establishing a spatio-temporal knowledge inference graph based on the initial defect detection result, setting the initial defect detection result as a graph node, calculating the morphological similarity and spatial position relationship between each initial defect detection result and constructing an edge connection strength, performing hierarchical clustering on the initial defect detection result based on the edge connection strength to generate similar defect groups, performing graph attention network operations on the spatio-temporal knowledge inference graph, generating a defect-related feature vector and performing temporal correlation analysis on the defect-related feature vector and a process parameter sequence, constructing a process parameter influence model, and performing reliability evaluation and optimization on the initial defect detection result based on the process parameter influence model and the similar defect groups to obtain a final defect detection result.
8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.