Semiconductor defect detection and process optimization method based on deep learning
Through deep learning-based methods, semiconductor defect detection and process optimization are solved, and the problems of low defect detection accuracy, insufficient classification accuracy and lack of real-time optimization closed-loop control in the existing technology are solved, high-precision detection and process optimization are achieved, and production efficiency and product quality are improved.
Patent Information
- Application Number
- CN202510563392.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing semiconductor manufacturing technology, defect detection accuracy is low and cannot meet the needs of high-precision manufacturing; defect classification depends on manual experience, and the accuracy and real-timeness are insufficient; there is a lack of real-time process optimization closed-loop control mechanism based on defect detection results, which affects production efficiency and product quality.
Semiconductor defect detection and process optimization are used based on deep learning, data is collected through multi-spectral imaging, multi-modal fusion preprocessing is implemented, micro-nanoscale topological features are extracted, frequency domain-airspace joint modeling is carried out, defects are dynamically classified, process optimization schemes are generated, and closed-loop control is realized.
It realizes high-precision detection and classification of semiconductor wafer surface defects, dynamically optimizes semiconductor manufacturing processes, improves production efficiency and product quality, and realizes intelligent and refined process optimization.
Smart Images

Figure CN120107239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semiconductor manufacturing technology, and more specifically, to a semiconductor defect detection and process optimization method based on deep learning. Background Art
[0002] In the semiconductor manufacturing process, wafer surface defect detection and process optimization are key links to ensure product quality. Traditional defect detection methods mainly rely on optical inspection equipment, which uses single-band imaging technology to obtain wafer surface images, and then combines manual experience or simple image processing algorithms to identify defects. However, this method has many limitations. First, single-band imaging is difficult to fully reflect the complex defect characteristics on the wafer surface, especially for micro-nano scale defects, the detection accuracy is insufficient. Secondly, traditional methods cannot integrate process parameter data in real time for dynamic classification, resulting in poor accuracy and real-time performance of defect classification. In addition, the existing technology lacks an effective closed-loop control mechanism in process optimization, and it is difficult to quickly adjust manufacturing process parameters according to defect detection results, which affects production efficiency and product quality.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: the existing detection methods have low detection accuracy for micro-nano scale defects and cannot meet the needs of high-precision manufacturing; defect classification relies on manual experience and lacks accuracy and real-time performance; there is a lack of real-time process optimization closed-loop control mechanism based on defect detection results, which makes it difficult to effectively improve production efficiency and product quality. Summary of the invention
[0004] The present invention provides a semiconductor defect detection and process optimization method based on deep learning, comprising: S1. Collect multispectral image data of the semiconductor wafer surface through a multispectral imaging device, and simultaneously obtain the wafer reflectivity distribution data and thickness distribution data, wherein the multispectral image data includes a visible light band, a near infrared band, and a defect characteristic response band; S2. Implementing multimodal fusion preprocessing based on the multispectral image data to generate a defect enhanced image; S3. Extracting micro- and nanoscale topological features of wafer surface through lattice structure-guided deep learning network; S4. Implementing frequency domain-spatial domain joint modeling on the topological features to generate a defect feature energy distribution map; S5. Dynamically classify defect features by integrating real-time process parameter data, and output defect type and confidence level; S6. Generate process equipment control instructions based on the defect classification results; S7. Build a knowledge graph of defect causes through a causal reasoning engine and output a process optimization solution; S8. Feedback the optimized process parameters to the semiconductor manufacturing equipment to form a closed-loop control.
[0005] Furthermore, the step S2 comprises: S21. Perform sub-pixel pattern alignment based on the diffraction characteristics of visible light band images, specifically using a grating period matching algorithm; S22. Compensation correction is performed on the near-infrared band image based on the thickness distribution data. The correction formula is:
[0006] in, represents the corrected pixel value, represents the original pixel value, is the reference thickness value, is the thickness measurement value at the current position; S23. Adaptively enhance the defect characteristic response band image based on the band gap energy level of the semiconductor material.
[0007] Furthermore, the step S3 comprises: S31. Construct a convolution kernel with rotational symmetry constraints, and its weights are initialized by the following formula:
[0008] in, is the final convolution kernel weight, N is the rotational symmetry order, represents the rotation transformation operation, W is the initial weight matrix; S32. Dynamically adjust the shape of the pooling layer window based on the wafer cutting direction; S33. The lattice periodicity of the feature map is maintained by the following constraints:
[0009] in, Represents feature map coordinates The eigenvalue at is the lattice period parameter.
[0010] Furthermore, the step S5 comprises: S51. Encode the lithography process parameters into a process feature vector ; S52. Based on defect feature vector Process feature vector Calculate dynamic association weights:
[0011] in, is the association weight, is the trainable projection matrix, Represents element-wise multiplication; S53. Generate a fusion feature vector based on the association weight:
[0012] in, is the fusion feature vector of the input classifier.
[0013] Further, the step S6 comprises: S61. Calculate the defocus compensation value of the lithography machine based on the defect spatial distribution gradient:
[0014] in, is the focal plane compensation value, D is the defect density distribution function, is the equipment characteristic coefficient; S62. Dynamically adjust the exposure dose based on the line width roughness deviation, and the adjustment formula includes a feedback item of the current measurement value and the target value.
[0015] Furthermore, the step S7 comprises: S71. Construct a process parameter cause-and-effect graph, where the nodes include temperature, pressure, and gas flow parameters, and the edge weights represent the strength of the cause-and-effect relationship between the parameters; S72. Calculate the effect of parameter intervention using counterfactual reasoning algorithm:
[0016] in, represents the probability of the outcome after implementing intervention X, and Z is a set of confounding variables; S73. Generate an optimization proposal including a minimum set of intervention parameters.
[0017] Furthermore, the step S21 includes: S211. Extracting the main diffraction peak position of the grating based on Fourier spectrum analysis; S212. Calculate the sub-pixel offset using the following formula:
[0018] in, is the offset, is the spectral phase, is the imaging wavelength, is the frequency domain coordinate; S213. Compensate image deformation based on thin plate spline interpolation algorithm.
[0019] Further, the step S32 includes: S321. Activate the diamond pooling window when a universal crystal orientation is detected; S322. Activate the square pooling window when a cross-cutting crystal orientation is detected; S323. Dynamically switch the pooling mode through the crystal orientation marking channel.
[0020] Further, the step S62 includes: S621. Set exposure dose adjustment constraints:
[0021] in, For dose adjustment, is the predefined safety factor, is the baseline exposure dose; S622. When the constraint condition is exceeded, jointly adjust the exposure dose, mask offset and illumination aperture angle parameters.
[0022] Furthermore, the S8 further includes: S81. Update the lithography machine control instructions based on the optimized exposure dose parameters; S82. Adjust the mass flow controller setting value based on the new etching gas ratio parameter; S83. Update the process parameter association database of the causal reasoning engine after each batch of production.
[0023] The above-mentioned embodiments according to the present invention have at least the following beneficial effects: the present invention can realize high-precision detection of surface defects of semiconductor wafers. By collecting multispectral image data including visible light band, near-infrared band and defect characteristic response band through multispectral imaging equipment, and combining wafer reflectivity distribution data and thickness distribution data, the complex defect characteristics of the wafer surface can be fully reflected. The defect enhancement image is generated based on multimodal fusion preprocessing, and the micro-nano scale topological features are extracted through a deep learning network guided by a lattice structure, which can further improve the accuracy and reliability of defect detection. In addition, the defect feature energy distribution map generated by the frequency domain-spatial domain joint modeling can present the defect characteristics more intuitively, providing strong support for subsequent defect classification and process optimization.
[0024] The present invention can also realize dynamic optimization of semiconductor manufacturing processes. By integrating real-time process parameter data to dynamically classify defect features, the defect type and confidence can be output, providing an accurate basis for process optimization. According to the process equipment control instructions generated by the defect classification results, the manufacturing process parameters, such as the defocus compensation value and exposure dose of the lithography machine, can be adjusted in real time, thereby improving production efficiency and product quality. By constructing a defect cause knowledge graph through a causal reasoning engine and outputting optimization suggestions containing a minimum set of intervention parameters, the intelligent and refined process optimization can be realized. The optimized process parameters are fed back to the semiconductor manufacturing equipment to form a closed-loop control, which can further improve the effect and stability of process optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which: Figure 1 A schematic flow chart of a semiconductor defect detection and process optimization method based on deep learning provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0026] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0027] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, apparatus, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0028] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0029] Reference below Figure 1 , Figure 1 A schematic diagram of a process flow of a semiconductor defect detection and process optimization method based on deep learning provided by an embodiment of the present invention. Figure 1 As shown, a semiconductor defect detection and process optimization method 100 based on deep learning includes: S1. Collect multispectral image data of the semiconductor wafer surface through a multispectral imaging device, and simultaneously obtain the wafer reflectivity distribution data and thickness distribution data, wherein the multispectral image data includes a visible light band, a near infrared band, and a defect characteristic response band; S2. Implementing multimodal fusion preprocessing based on the multispectral image data to generate a defect enhanced image; S3. Extracting micro- and nanoscale topological features of wafer surface through lattice structure-guided deep learning network; S4. Implementing frequency domain-spatial domain joint modeling on the topological features to generate a defect feature energy distribution map; S5. Dynamically classify defect features by integrating real-time process parameter data, and output defect type and confidence level; S6. Generate process equipment control instructions based on the defect classification results; S7. Build a knowledge graph of defect causes through a causal reasoning engine and output a process optimization solution; S8. Feedback the optimized process parameters to the semiconductor manufacturing equipment to form a closed-loop control.
[0030] It should be noted that the present invention proposes a semiconductor defect detection and process optimization method based on deep learning. Among them, multispectral image data of the semiconductor wafer surface is collected by a multispectral imaging device. The multispectral image data here refers to image data including visible light band, near infrared band and defect characteristic response band. The visible light band refers to the range of light waves that the human eye can perceive, which is usually used to observe the macroscopic features of the wafer surface; the near infrared band can penetrate the wafer material to a certain depth and is used to detect internal defects; the defect characteristic response band refers to the light band that is sensitive to specific defects. By selecting a suitable band, the detectability of defects can be enhanced. At the same time, the method also synchronously obtains the wafer reflectivity distribution data and thickness distribution data. The reflectivity distribution data reflects the light reflection ability of different positions on the wafer surface, while the thickness distribution data provides the thickness information of the wafer at different positions. These data are crucial for subsequent defect detection and process optimization.
[0031] Specifically, the multispectral imaging device may include light sources of multiple different wavelength bands and corresponding imaging sensors. For example, the light source in the visible light band may be an ordinary white light source, and its wavelength range is usually between 400 nanometers and 700 nanometers; the wavelength range of the light source in the near-infrared band is between 700 nanometers and 2500 nanometers. The selection of defect characteristic response bands needs to be determined according to the specific semiconductor material and defect type. For example, for certain specific oxide layer defects, a specific near-infrared band will be selected for imaging. While collecting image data, the wafer reflectivity distribution data is obtained through optical measurement equipment. For example, equipment such as ellipsometers can be used. The measurement principle is to determine the optical properties of the material based on the polarization characteristics of light, and then obtain the reflectivity distribution. Thickness distribution data can be obtained through technologies such as optical coherence tomography (OCT), which uses the coherence principle of light to determine the thickness distribution of the material by measuring the phase difference of the reflected light. The acquisition of these data provides a basis for subsequent multimodal fusion preprocessing.
[0032] Preferably, when collecting multispectral image data, images of different bands can be time synchronized and spatially aligned to ensure the consistency of image data. For example, time synchronization can be achieved by timestamp marking, and spatial alignment can be achieved by image registration algorithm. In multimodal fusion preprocessing, machine learning-based algorithms can be used to fuse features of images of different bands, such as using dimensionality reduction techniques such as principal component analysis (PCA) to extract the main features of the image, and then weighted fusion of these features to generate defect enhanced images. In addition, for the processing of wafer reflectivity distribution data and thickness distribution data, data smoothing algorithms can be used to remove noise, such as using moving average method or Gaussian filtering to improve data quality and reliability.
[0033] In some embodiments, step S2 includes: S21. Perform sub-pixel pattern alignment based on the diffraction characteristics of visible light band images, specifically using a grating period matching algorithm; S22. Compensation correction is performed on the near-infrared band image based on the thickness distribution data. The correction formula is:
[0034] in, represents the corrected pixel value, represents the original pixel value, is the reference thickness value, is the thickness measurement value at the current position; S23. Adaptively enhance the defect characteristic response band image based on the band gap energy level of the semiconductor material.
[0035] It should be noted that in the multimodal fusion preprocessing stage, the present invention optimizes the collected multispectral image data through a series of image processing technologies. Specifically, first, sub-pixel pattern alignment is performed based on the diffraction characteristics of the visible light band image. This is because the visible light band image can clearly reflect the grating structure on the surface of the wafer, and the grating period matching algorithm can achieve accurate alignment of the image by identifying the periodic characteristics of the grating. Secondly, the near-infrared band image is compensated and corrected using thickness distribution data. This is because the near-infrared band image is easily affected by changes in wafer thickness. This effect can be eliminated through compensation correction, so that the image can more accurately reflect the true state of the wafer. Finally, the defect feature response band image is adaptively enhanced according to the band gap energy level of the semiconductor material. This is because different semiconductor materials have different absorption and reflection characteristics for light in specific bands. Adaptive enhancement can highlight defect features and improve detection sensitivity.
[0036] Specifically, sub-pixel pattern alignment is achieved through a grating period matching algorithm. The core of the grating period matching algorithm is to identify the periodic features of the grating in the visible light band image, and then calculate the displacement and rotation angle between the images to achieve high-precision alignment. The thickness distribution data is used to compensate and correct the near-infrared band image, and its compensation formula is derived based on the relationship between the wafer thickness and the image pixel value. In this formula, the corrected pixel value is calculated by multiplying the original pixel value by the thickness ratio, where the reference thickness value is a pre-set standard thickness value, and the current position thickness measurement value is obtained in real time through the thickness measurement device. Adaptive enhancement adjusts the contrast and brightness of the defect feature response band image according to the band gap energy level of the semiconductor material to make the defect feature more obvious. The band gap energy level refers to the energy required for electrons in semiconductor materials to transition from the valence band to the conduction band. It determines the absorption characteristics of the material for light of different wavelengths. Therefore, appropriate enhancement parameters can be selected according to the band gap energy level.
[0037] Preferably, in the process of sub-pixel pattern alignment, Fourier spectrum analysis can be used to extract the position of the main diffraction peak of the grating, and then the sub-pixel offset is determined by calculating the relationship between the spectrum phase and the imaging wavelength. This method can improve the accuracy of alignment, especially when processing high-resolution images. For compensation correction of near-infrared band images, a correction coefficient can be introduced to adjust the intensity of compensation, and this correction coefficient can be optimized according to the characteristics of the wafer material. For example, the correction coefficient will be different for different types of silicon wafers. In terms of adaptive enhancement, the enhancement parameters can be dynamically adjusted in combination with the local and global features of the image, such as adjusting the degree of contrast enhancement according to the grayscale distribution of the defective area, so as to better highlight the defect features.
[0038] In some embodiments, step S3 includes: S31. Construct a convolution kernel with rotational symmetry constraints, and its weights are initialized by the following formula:
[0039] in, is the final convolution kernel weight, N is the rotational symmetry order, represents the rotation transformation operation, W is the initial weight matrix; S32. Dynamically adjust the shape of the pooling layer window based on the wafer cutting direction; S33. The lattice periodicity of the feature map is maintained by the following constraints:
[0040] in, Represents feature map coordinates The eigenvalue at is the lattice period parameter.
[0041] It should be noted that in the deep learning network construction stage, the present invention uses a series of innovative technical means to extract the topological features of the wafer surface at the micro-nano scale. Specifically, first, a convolution kernel with rotational symmetry constraints is constructed, and its weight initialization formula is based on rotational symmetry to ensure that the convolution kernel can better capture the periodic features of the wafer surface. Secondly, the shape of the pooling layer window is dynamically adjusted based on the wafer cutting direction. This is because the cutting direction of the wafer will affect the distribution and feature performance of the defects. By dynamically adjusting the pooling window, it can better adapt to the defect features in different directions. Finally, the lattice periodicity of the feature map is maintained by the lattice periodicity constraint term, which helps to retain the periodic structural information of the wafer surface during the feature extraction process, thereby improving the accuracy of defect detection.
[0042] Specifically, the rotational symmetry constrained convolution kernel weight initialization formula is obtained by performing rotational symmetry operations on the initial weight matrix and averaging it. Among them, the rotational symmetry order N represents the number of times the convolution kernel needs to maintain symmetry during the rotation process. The rotation transformation operation rotates the initial weight matrix at a certain angle to generate multiple rotated weight matrices, and then averages the corresponding elements of these matrices to obtain the final convolution kernel weight. The shape of the pooling layer window is dynamically adjusted according to the wafer cutting direction to determine the shape of the pooling window according to the actual cutting direction of the wafer. For example, if the wafer is cut along the universal crystal direction, the diamond pooling window is activated; if it is cut along the cross-crystalline direction, the square pooling window is activated. The lattice periodicity constraint term is realized by constraining the eigenvalue difference between adjacent lattice points of the feature map, which can ensure that the feature map remains periodic in the lattice direction, thereby better reflecting the periodic structure of the wafer surface.
[0043] Preferably, when constructing a convolution kernel with rotational symmetry constraints, the initial weight matrix can be rotated multiple times, and the angle of each rotation can be determined according to the periodicity of the wafer surface features. For example, for a wafer with a hexagonal lattice structure, the rotation angle can be set to an integer multiple of 60 degrees. When dynamically adjusting the shape of the pooling layer window, the dynamic switching of the pooling mode can be achieved through the crystal orientation marking channel. For example, a marking channel can be set for each crystal orientation, and when a specific crystal orientation is detected, the corresponding pooling window is activated. In addition, an adaptive learning mechanism can be introduced to adjust the weight of the lattice periodic constraint term, for example, the strength of the constraint term is dynamically adjusted according to the error feedback during the training process, thereby further improving the accuracy and robustness of feature extraction.
[0044] In some embodiments, step S5 includes: S51. Encode the lithography process parameters into a process feature vector ; S52. Based on defect feature vector Process feature vector Calculate dynamic association weights:
[0045] in, is the association weight, is the trainable projection matrix, Represents element-wise multiplication; S53. Generate a fusion feature vector based on the association weight:
[0046] in, is the fusion feature vector of the input classifier.
[0047] It should be noted that in the stage of dynamic classification of defect features, the present invention realizes dynamic classification of defect types by fusing real-time process parameter data with defect feature vectors. Specifically, the lithography process parameters are first encoded into process feature vectors. This is because the lithography process parameters have a direct impact on the formation of defects. These parameters can be converted into vector forms that can be used for subsequent calculations through encoding. Secondly, dynamic association weights are calculated based on the defect feature vector and the process feature vector. This process is implemented through a trainable projection matrix, which can dynamically adjust the association strength between defect features and process parameters. Finally, a fused feature vector is generated based on the association weight, and the vector is passed to the classifier as input to output the defect type and confidence, thereby achieving high-precision defect classification.
[0048] Specifically, when the lithography process parameters are encoded as process feature vectors, key process parameters such as exposure dose, focal length, mask offset, etc. can be normalized and then combined into a vector. For example, the exposure dose can be normalized to a range of 0 to 1, the focal length can be converted to an offset relative to the standard focal length, etc. The dynamic association weight is calculated by projecting the defect feature vector and the process feature vector into a low-dimensional space respectively, and then multiplying them element by element. The trainable projection matrix is automatically learned through the training process of the deep learning network, and its purpose is to maximize the accuracy of defect classification. The generation of the fused feature vector is to combine the defect feature vector and the process feature vector by weighted summation, and the weight is determined by the dynamic association weight, thereby generating a feature vector that comprehensively considers the defect characteristics and process parameters.
[0049] Preferably, when calculating the dynamic association weight, a normalization step can be introduced to ensure that the value of the association weight is between 0 and 1, thereby avoiding numerical instability problems caused by excessive or small weights. For example, the calculated association weights can be normalized by the Softmax function. When generating a fused feature vector, it is possible to consider introducing an adjustable parameter to balance the contribution of the defect feature vector and the process feature vector, for example, by determining the optimal value of the parameter through methods such as cross-validation. In addition, other types of process parameters, such as etching time, gas flow rate, etc., can be combined to further enrich the process feature vector, thereby improving the accuracy of defect classification.
[0050] In some embodiments, step S6 includes: S61. Calculate the defocus compensation value of the lithography machine based on the defect spatial distribution gradient:
[0051] in, is the focal plane compensation value, D is the defect density distribution function, is the equipment characteristic coefficient; S62. Dynamically adjust the exposure dose based on the line width roughness deviation, and the adjustment formula includes a feedback item of the current measurement value and the target value.
[0052] It should be noted that in the process optimization stage, the present invention realizes dynamic optimization of lithography process parameters by calculating the defocus compensation value of the lithography machine based on the gradient of the spatial distribution of defects and dynamically adjusting the exposure dose based on the line width roughness deviation. Specifically, the calculation of the defocus compensation value of the lithography machine is based on the defect density distribution function, and the compensation value is determined by analyzing the spatial distribution gradient of the defects on the wafer surface, thereby adjusting the focal plane of the lithography machine to reduce defects caused by defocus. At the same time, the dynamic adjustment of the line width roughness deviation is to dynamically adjust the exposure dose by real-time monitoring the deviation between the actual measured value of the line width and the target value to ensure accurate control of the line width, thereby improving the quality of lithography.
[0053] Specifically, in the calculation of the defocus compensation value of the lithography machine, the defect density distribution function is obtained through statistical analysis of the defects on the wafer surface, which reflects the density changes of defects at different locations. The equipment characteristic coefficient is pre-set according to the specific parameters and performance characteristics of the lithography machine, and is used to adjust the size of the compensation value to adapt to the characteristics of different lithography machines. In the dynamic adjustment of the line width roughness deviation, the line width roughness deviation refers to the difference between the actual line width and the target line width. Through real-time measurement and feedback mechanism, the exposure dose is dynamically adjusted to reduce this deviation. The feedback item includes the difference between the current measurement value and the target value, which is used to calculate the adjustment amount to ensure accurate control of the line width.
[0054] Preferably, when calculating the defocus compensation value of the lithography machine, a smoothing factor can be introduced to reduce the fluctuation of the compensation value caused by the sudden change of the local defect density. For example, the defect density distribution function is smoothed by the moving average method to obtain a more stable compensation value. When dynamically adjusting the exposure dose, a safety threshold can be set. When the calculated dose adjustment exceeds the threshold, the joint adjustment mechanism is triggered to adjust the exposure dose, mask offset and illumination aperture angle parameters at the same time to ensure the safety and effectiveness of the process parameter adjustment. In addition, other process parameters such as etching time, gas flow rate, etc. can be combined to further optimize the lithography process and improve the overall manufacturing quality.
[0055] In some embodiments, step S7 includes: S71. Construct a process parameter cause-and-effect graph, where the nodes include temperature, pressure, and gas flow parameters, and the edge weights represent the strength of the cause-and-effect relationship between the parameters; S72. Calculate the effect of parameter intervention using counterfactual reasoning algorithm:
[0056] in, represents the probability of the outcome after implementing intervention X, and Z is a set of confounding variables; S73. Generate an optimization proposal including a minimum set of intervention parameters.
[0057] It should be noted that in the process optimization stage, the present invention realizes dynamic optimization of lithography process parameters by calculating the defocus compensation value of the lithography machine based on the gradient of the spatial distribution of defects and dynamically adjusting the exposure dose based on the line width roughness deviation. Specifically, the calculation of the defocus compensation value of the lithography machine is based on the defect density distribution function, and the compensation value is determined by analyzing the spatial distribution gradient of the defects on the wafer surface, thereby adjusting the focal plane of the lithography machine to reduce defects caused by defocus. At the same time, the dynamic adjustment of the line width roughness deviation is to dynamically adjust the exposure dose by real-time monitoring the deviation between the actual measured value of the line width and the target value to ensure accurate control of the line width, thereby improving the quality of lithography.
[0058] Specifically, in the calculation of the defocus compensation value of the lithography machine, the defect density distribution function is obtained through statistical analysis of the defects on the wafer surface, which reflects the density changes of defects at different locations. The equipment characteristic coefficient is pre-set according to the specific parameters and performance characteristics of the lithography machine, and is used to adjust the size of the compensation value to adapt to the characteristics of different lithography machines. In the dynamic adjustment of the line width roughness deviation, the line width roughness deviation refers to the difference between the actual line width and the target line width. Through real-time measurement and feedback mechanism, the exposure dose is dynamically adjusted to reduce this deviation. The feedback item includes the difference between the current measurement value and the target value, which is used to calculate the adjustment amount to ensure accurate control of the line width.
[0059] Preferably, when calculating the defocus compensation value of the lithography machine, a smoothing factor can be introduced to reduce the fluctuation of the compensation value caused by the sudden change of the local defect density. For example, the defect density distribution function is smoothed by the moving average method to obtain a more stable compensation value. When dynamically adjusting the exposure dose, a safety threshold can be set. When the calculated dose adjustment exceeds the threshold, the joint adjustment mechanism is triggered to adjust the exposure dose, mask offset and illumination aperture angle parameters at the same time to ensure the safety and effectiveness of the process parameter adjustment. In addition, other process parameters such as etching time, gas flow rate, etc. can be combined to further optimize the lithography process and improve the overall manufacturing quality.
[0060] In some embodiments, step S21 includes: S211. Extracting the main diffraction peak position of the grating based on Fourier spectrum analysis; S212. Calculate the sub-pixel offset using the following formula:
[0061] in, is the offset, is the spectral phase, is the imaging wavelength, is the frequency domain coordinate; S213. Compensate image deformation based on thin plate spline interpolation algorithm.
[0062] It should be noted that in the process optimization stage, the present invention realizes intelligent optimization of semiconductor manufacturing processes by constructing a process parameter causal graph and using a counterfactual reasoning algorithm to calculate the parameter intervention effect. Specifically, first, a process parameter causal graph is constructed, in which nodes represent key process parameters such as temperature, pressure, and gas flow, and edge weights represent the strength of the causal relationship between parameters, thereby clearly showing the mutual influence between the parameters. Secondly, a counterfactual reasoning algorithm is used to calculate the parameter intervention effect, and by analyzing the changes in each parameter after the implementation of a specific intervention, the impact of the intervention on the final result is predicted, and then an optimization suggestion containing a minimum set of intervention parameters is generated, providing a scientific basis for process optimization.
[0063] Specifically, in the construction of the process parameter causal graph, the nodes include but are not limited to key process parameters such as temperature, pressure, and gas flow rate, which are key factors that directly affect product quality and production efficiency in the semiconductor manufacturing process. The edge weight represents the strength of the causal relationship between parameters, which can be determined through historical data statistics or expert experience. For example, changes in temperature will affect the stability of gas flow, and this causal relationship can be quantified by edge weights. The core of the counterfactual reasoning algorithm is to calculate the probability change of the result after the intervention is implemented, where the confounding variable set Z represents other factors that may affect the result. Through this algorithm, the effects of different intervention measures can be evaluated, so as to select the optimal intervention strategy.
[0064] Preferably, when constructing a process parameter causal graph, a dynamic update mechanism can be introduced to dynamically adjust the node and edge weights according to real-time monitoring data to reflect the real-time changes in process parameters. For example, when an abnormal temperature fluctuation is detected, the edge weights related to the temperature are updated in a timely manner to more accurately reflect the causal relationship between process parameters. In the counterfactual reasoning algorithm, a machine learning model can be combined to improve the prediction accuracy of the intervention effect. For example, a deep learning network is used to learn historical data to more accurately predict the probability of the result after the intervention. In addition, a multi-objective optimization algorithm can be introduced to consider multiple optimization goals at the same time, such as improving product quality, reducing production costs, etc., to generate more comprehensive optimization suggestions.
[0065] In some embodiments, step S32 includes: S321. Activate the diamond pooling window when a universal crystal orientation is detected; S322. Activate the square pooling window when a cross-cutting crystal orientation is detected; S323. Dynamically switch the pooling mode through the crystal orientation marking channel.
[0066] It should be noted that in the preprocessing stage of multispectral image data, the present invention achieves high-precision alignment of images by extracting the main diffraction peak position of the grating based on Fourier spectrum analysis and calculating the sub-pixel offset. This method can effectively compensate for image deformation and improve image quality, thereby providing more accurate input data for subsequent defect detection. Specifically, Fourier spectrum analysis is an image processing technology based on the frequency domain. It extracts the main diffraction peak position of the grating by analyzing the spectral characteristics of the image, and then calculates the sub-pixel offset between the images. Finally, the thin plate spline interpolation algorithm is used to compensate for image deformation to ensure the accuracy of image alignment.
[0067] Specifically, Fourier spectrum analysis converts the image from the spatial domain to the frequency domain by performing a fast Fourier transform (FFT) on the image, thereby extracting the frequency components of the image. The position of the main diffraction peak of the grating refers to the position of the obvious peak caused by the grating structure in the spectrum diagram. These peaks reflect the periodic characteristics of the grating. By analyzing the position of these peaks, the sub-pixel offset between images can be calculated. The calculation of the sub-pixel offset is based on the relationship between the spectral phase and the imaging wavelength, and the offset of the image in the horizontal and vertical directions is calculated through a specific formula. The thin plate spline interpolation algorithm is an interpolation method based on spline functions, which is used to compensate for image deformation and ensure the accuracy of image alignment.
[0068] Preferably, in Fourier spectrum analysis, the spectrum graph can be high-pass filtered to remove the interference of low-frequency noise, so as to more accurately extract the main diffraction peak position of the grating. For example, a threshold can be set to filter out frequency components below the threshold. When calculating the sub-pixel offset, an error correction mechanism can be introduced to improve the accuracy of the offset through multiple iterative calculations. For example, the least squares method can be used to optimize the calculation results. In addition, the parameters of the thin plate spline interpolation algorithm can be adjusted according to the specific characteristics of the image, such as adjusting the smoothing parameters of the spline function to better adapt to different types of image deformation.
[0069] In some embodiments, step S62 includes: S621. Set exposure dose adjustment constraints:
[0070] in, For dose adjustment, is the predefined safety factor, is the baseline exposure dose; S622. When the constraint condition is exceeded, jointly adjust the exposure dose, mask offset and illumination aperture angle parameters.
[0071] It should be noted that in the deep learning network construction stage, the present invention dynamically adjusts the shape of the pooling layer window to adapt to the micro-nano scale characteristics of the wafer surface according to the detection requirements of different crystal directions. Specifically, the diamond pooling window is activated when the universal crystal direction is detected, and the square pooling window is activated when the cross-sectional crystal direction is detected, and the pooling mode is dynamically switched through the crystal direction marking channel. This design can better capture the defect characteristics of different crystal directions and improve the accuracy and robustness of defect detection.
[0072] Specifically, crystal orientation refers to the direction of the crystal structure on the surface of the wafer. The defect characteristics and distribution patterns of different crystal orientations may be different. Universal crystal orientation and cross-sectional crystal orientation are two common types of crystal orientations, which correspond to different crystal structures and defect manifestations. Diamond pooling windows and square pooling windows are two different shape configurations of the pooling layer. The diamond window is more suitable for capturing the defect characteristics of the universal crystal downward, while the square window is more suitable for the cross-sectional crystal orientation. The crystal orientation marking channel is a mechanism for indicating the current crystal orientation. By marking the crystal orientation information in the data stream, the network can dynamically select the appropriate pooling window shape based on the mark.
[0073] Preferably, when detecting the crystal orientation, a pre-trained crystal orientation classifier can be introduced, which automatically identifies the current crystal orientation by analyzing the features of the input image, thereby triggering the corresponding pooling window. For example, a convolutional neural network (CNN) can be used to classify the crystal orientation, with its input being multispectral image data and its output being the crystal orientation type. In the activation mechanism of the pooling window, a threshold can be set, and when the confidence of the crystal orientation classifier exceeds the threshold, the corresponding pooling window is activated to reduce misjudgment. In addition, other feature extraction techniques, such as local sensitive hashing (LSH), can be combined to further optimize the selection strategy of the pooling window to adapt to more complex crystal orientation features.
[0074] In some embodiments, the S8 further includes: S81. Update the lithography machine control instructions based on the optimized exposure dose parameters; S82. Adjust the mass flow controller setting value based on the new etching gas ratio parameter; S83. Update the process parameter association database of the causal reasoning engine after each batch of production.
[0075] It should be noted that in the closed-loop control stage, the present invention realizes closed-loop control of process optimization by feeding back the optimized process parameters to the semiconductor manufacturing equipment. Specifically, first, the photolithography machine control instructions are updated based on the optimized exposure dose parameters to ensure the accuracy of the photolithography process. Secondly, the mass flow controller setting value is adjusted based on the new etching gas ratio parameters to optimize the etching process. Finally, the process parameter association database of the causal reasoning engine is updated after each batch of production to accumulate production data and continuously optimize process parameters, thereby improving production efficiency and product quality.
[0076] Specifically, the update of the lithography machine control instructions is based on the optimized exposure dose parameters, which are calculated in the previous steps and used to adjust the exposure dose of the lithography machine to reduce the line width roughness deviation and improve the lithography quality. The adjustment of the mass flow controller setting value is based on the new etching gas ratio parameters, which are also obtained through the optimization process and are used to control the gas flow rate during the etching process to ensure the uniformity and accuracy of etching. The process parameter association database of the causal reasoning engine is a data structure that stores process parameters and their relationships. Through updates after each batch of production, more production data can be accumulated, thereby providing a more accurate basis for subsequent process optimization.
[0077] Preferably, when updating the lithography machine control instructions, a buffer mechanism can be introduced to avoid equipment instability caused by too fast parameter adjustment. For example, a transition interval can be set between the old and new parameters, and the parameters can be gradually adjusted to ensure the smooth operation of the equipment. When adjusting the set value of the mass flow controller, dynamic adjustments can be made in combination with real-time monitoring data to cope with fluctuations in the production process. For example, the actual flow rate of the etching gas is monitored in real time by a sensor, and fine-tuned as needed. In addition, when updating the process parameter association database, a data cleaning and verification mechanism can be introduced to ensure the accuracy and reliability of the data. For example, by comparing historical data with new data, outliers can be eliminated, thereby improving the quality of the database and the optimization effect.
[0078] The above-mentioned embodiments of the present invention have the following beneficial effects: the present invention can realize high-precision detection and classification of surface defects of semiconductor wafers. By collecting multi-band image data and wafer reflectivity and thickness distribution data through multi-spectral imaging equipment, and combining multi-modal fusion preprocessing to generate defect enhanced images, the complex defect features on the wafer surface can be presented comprehensively and accurately. The deep learning network guided by the lattice structure extracts micro-nano scale topological features, and implements frequency domain-spatial domain joint modeling to generate defect feature energy distribution maps, which can further improve the accuracy and reliability of defect detection. In addition, the defect features are dynamically classified by fusing real-time process parameter data, and the defect type and confidence can be output, providing an accurate basis for subsequent process optimization. At the same time, preprocessing methods for multi-spectral image data, such as sub-pixel pattern alignment, near-infrared band image compensation correction, and defect feature response band image adaptive enhancement, can effectively improve image quality and enhance the recognizability of defect features, thereby further improving the accuracy of defect detection.
[0079] The present invention can also realize dynamic optimization and closed-loop control of semiconductor manufacturing process. According to the defect classification results, process equipment control instructions are generated, and manufacturing process parameters, such as the defocus compensation value and exposure dose of the lithography machine, can be adjusted in real time, thereby improving production efficiency and product quality. By constructing a defect cause knowledge map through a causal reasoning engine, and using a counterfactual reasoning algorithm to calculate the parameter intervention effect, an optimization suggestion containing a minimum set of intervention parameters can be generated to achieve intelligent and refined process optimization. The optimized process parameters are fed back to the semiconductor manufacturing equipment to form a closed-loop control, which can further improve the effect and stability of process optimization. In addition, the optimization method of process parameters, such as calculating the defocus compensation value of the lithography machine based on the defect spatial distribution gradient, dynamically adjusting the exposure dose based on the line width roughness deviation, and setting the exposure dose adjustment constraint conditions, can effectively ensure the rationality and safety of process parameter adjustment, and avoid production problems caused by improper parameter adjustment.
[0080] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0081] The above descriptions are only some preferred embodiments of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A semiconductor defect detection and process optimization method based on deep learning, characterized in that: The following steps are involved: S1. Collect multispectral image data of the semiconductor wafer surface through a multispectral imaging device, and simultaneously obtain the wafer reflectivity distribution data and thickness distribution data, wherein the multispectral image data includes a visible light band, a near infrared band, and a defect characteristic response band; S2. Implementing multimodal fusion preprocessing based on the multispectral image data to generate a defect enhanced image; S3. Extracting micro- and nanoscale topological features of wafer surface through lattice structure-guided deep learning network; S4. Implementing frequency domain-spatial domain joint modeling on the topological features to generate a defect feature energy distribution map; S5. Dynamically classify defect features by integrating real-time process parameter data, and output defect type and confidence level; S6. Generate process equipment control instructions based on the defect classification results; S7. Build a knowledge graph of defect causes through a causal reasoning engine and output a process optimization solution; S8. Feedback the optimized process parameters to the semiconductor manufacturing equipment to form a closed-loop control.
2. The method according to claim 1, characterized in that The step S2 comprises: S21. Based on the diffraction characteristics of visible light band images, a grating period matching algorithm is used to perform sub-pixel pattern alignment; S22. Compensation correction is performed on the near-infrared band image based on the thickness distribution data. The correction formula is: ; in, represents the corrected pixel value, represents the original pixel value, is the reference thickness value, is the thickness measurement value at the current position; S23. Adaptively enhance the defect characteristic response band image based on the band gap energy level of the semiconductor material.
3. The method according to claim 1, characterized in that The step S3 comprises: S31. Construct a convolution kernel with rotational symmetry constraints, and its weights are initialized by the following formula: ; in, is the final convolution kernel weight, N is the rotational symmetry order, represents the rotation transformation operation, W is the initial weight matrix; S32. Dynamically adjust the shape of the pooling layer window based on the wafer cutting direction; S33. The lattice periodicity of the feature map is maintained by the following constraints: ; in, Represents feature map coordinates The eigenvalue at is the lattice period parameter.
4. The method according to claim 1, characterized in that: The step S5 comprises: S51. Encode the lithography process parameters into a process feature vector ; S52. Based on defect feature vector Process feature vector Calculate dynamic association weights: ; in, is the association weight, is the trainable projection matrix, Represents element-wise multiplication; S53. Generate a fusion feature vector based on the association weight: ; in, is the fusion feature vector of the input classifier.
5. The method according to claim 1, characterized in that The step S6 comprises: S61. Calculate the defocus compensation value of the lithography machine based on the defect spatial distribution gradient: ; in, is the focal plane compensation value, D is the defect density distribution function, is the equipment characteristic coefficient; S62. Dynamically adjust the exposure dose based on the line width roughness deviation, and the adjustment formula includes a feedback item of the current measurement value and the target value.
6. The method according to claim 1, characterized in that The step S7 comprises: S71. Construct a process parameter cause-and-effect graph, where the nodes include temperature, pressure, and gas flow parameters, and the edge weights represent the strength of the cause-and-effect relationship between the parameters; S72. Calculate the effect of parameter intervention using counterfactual reasoning algorithm: ; in, represents the probability of the outcome after implementing intervention X, and Z is a set of confounding variables; S73. Generate an optimization proposal including a minimum set of intervention parameters.
7. The method according to claim 2, characterized in that The step S21 comprises: S211. Extracting the main diffraction peak position of the grating based on Fourier spectrum analysis; S212. Calculate the sub-pixel offset using the following formula: ; in, is the offset, is the spectral phase, is the imaging wavelength, is the frequency domain coordinate; S213. Compensate image deformation based on thin plate spline interpolation algorithm.
8. The method according to claim 3, characterized in that The step S32 comprises: S321. Activate the diamond pooling window when a universal crystal orientation is detected; S322. Activate the square pooling window when a cross-cutting crystal orientation is detected; S323. Dynamically switch the pooling mode through the crystal orientation marking channel.
9. The method according to claim 5, characterized in that The step S62 comprises: S621. Set exposure dose adjustment constraints: ; in, For dose adjustment, is the predefined safety factor, is the baseline exposure dose; S622. When the constraint condition is exceeded, jointly adjust the exposure dose, mask offset and illumination aperture angle parameters.
10. The method according to claim 1, characterized in that The S8 further comprises: S81. Update the lithography machine control instructions based on the optimized exposure dose parameters; S82. Adjust the mass flow controller setting value based on the new etching gas ratio parameter; S83. Update the process parameter association database of the causal reasoning engine after each batch of production.
Citation Information
Patent Citations
Product surface defect detection system based on multi-spectral imaging
CN110441312A
Image classification method for generalized equality convolution network model based on partial differential operator
CN112257753A
Wafer surface defect detection method and system based on deep learning and storage medium
CN117710378A
Automatic control method and system for battery film production based on visual inspection
CN119229372A
Cited By
PCB (Printed Circuit Board) defect detection method and system
CN120446164A
Wafer boat defect monitoring method and system based on image processing
CN120525885A
A wafer boat defect monitoring method and system based on image processing
CN120525885B
Automobile wire harness precision fastener surface flaw detection system based on visual identification
CN120672732A
Automobile wire harness precision fastener surface flaw detection system based on visual recognition
CN120672732B