Defect detection method for semiconductor packaging material based on deep learning

Through deep learning methods combined with improved operators and models, the high precision and high efficiency problems in semiconductor packaging material defect detection are solved, the fusion of two-dimensional-three-dimensional data and defect importance evaluation are realized, detection accuracy and efficiency are improved, and semiconductor packaging quality control is supported.

CN120525859APending Publication Date: 2025-08-22XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 29 Cited by

Patent Information

Application Number
CN202510923289.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The prior art is difficult to meet the needs of high precision and high efficiency in the detection of defects of semiconductor packaging materials, especially in multimodal data fusion, defect spatial positioning and importance evaluation, resulting in high missed detection rates and lack of quantitative analysis of the impact of the detection results on defects.

Method used

Using a deep learning-based method, two-dimensional contours and feature points are extracted through the improved Canny operator and FAST corner point detection algorithm, and pre-processed with adaptive median filtering and histogram specification; two-dimensional defect recognition is used by Transformer's multi-headed attention mechanism, and three-dimensional geometric features are obtained by structured light three-dimensional scanning. Two-dimensional geometric features are fused through spatiotemporal alignment model, and multi-modal fusion deep learning model is constructed for defect evaluation.

Benefits of technology

Accurate identification, positioning and impact assessment of semiconductor packaging material defects is achieved, detection accuracy and efficiency are improved, the accuracy of identification of complex defects is significantly improved, and the importance of defects in material structure is provided to support process optimization and preventive maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525859A_ABST
    Figure CN120525859A_ABST
Patent Text Reader

Abstract

The invention relates to the field of semiconductor packaging material defect detection, in particular to a semiconductor packaging material defect detection method based on deep learning, which comprises the following steps: acquiring a surface image, and extracting a two-dimensional contour and a feature point set; preprocessing the image, and separating a packaging material main body area; constructing a two-dimensional defect identification model based on Transform, and outputting a two-dimensional detection result; scanning suspected and unknown defect areas to obtain three-dimensional point cloud data, and extracting geometric and texture features; fusing two-dimensional and three-dimensional data through a space-time alignment model; utilizing the multi-modal fusion model to output defect positions and types; and evaluating the defect importance based on the material node connectivity and the stress distribution, and generating a visual detection report. According to the invention, high-precision detection of semiconductor packaging material defects is realized, the defect identification rate, the positioning precision and the detection efficiency are improved through multi-modal data fusion and a deep learning model, and a visual report can be generated based on material structure quantification defect importance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semiconductor packaging material defect detection, and in particular to a semiconductor packaging material defect detection method based on deep learning. Background Art

[0002] In semiconductor manufacturing, defect detection of packaging materials is crucial to device reliability and performance. With the increasing integration density of semiconductors and the increasing complexity of packaging processes, defect types are becoming increasingly diverse (such as cracks, holes, and interface delamination), and their sizes are trending towards the micro-nanoscale. Traditional detection methods are unable to meet the demand for high-precision and high-efficiency detection. Existing technologies have shortcomings in multimodal data fusion, spatial defect localization, and importance assessment, resulting in high missed detection rates and a lack of quantitative analysis of the impact of defects in test results. Therefore, there is an urgent need for an intelligent detection method based on deep learning to achieve accurate defect identification, localization, and impact assessment, thereby improving the automation and scientific nature of semiconductor packaging quality control.

[0003] Traditional semiconductor packaging material defect detection relies primarily on manual visual inspection, optical microscopy, and limited algorithm-assisted analysis. Manual inspection relies on the experience of the inspector, is inefficient, and is easily affected by subjective factors, making it difficult to adapt to the needs of large-scale production. Although optical microscopes can provide observations at a certain magnification, they can only obtain two-dimensional planar information and cannot reflect the three-dimensional spatial characteristics of defects. Their ability to detect defects hidden inside materials or in complex structures is limited. At the algorithmic level, early methods based on traditional image processing (such as edge detection and threshold segmentation) can achieve the identification of simple defects. However, when faced with images with complex textures and strong noise interference, their feature extraction capabilities are insufficient, and they lack a global analysis of defect context information, resulting in poor detection accuracy and robustness. In addition, traditional methods are unable to correlate the geometric characteristics of defects with the structural properties of the material, making it difficult to assess the actual impact of defects on device reliability.

[0004] Existing technologies have made certain progress driven by deep learning. Some solutions have introduced convolutional neural networks (CNNs) for two-dimensional image defect recognition, which has improved the degree of automation of feature extraction and detection efficiency compared to traditional algorithms. However, such methods are only based on single-modal two-dimensional image data, ignoring the three-dimensional spatial distribution information of defects, resulting in insufficient positioning accuracy for three-dimensional defects (such as hole depth and crack extension direction). At the same time, existing technologies have limitations in multi-source data fusion, failing to effectively combine the texture information of two-dimensional images with the geometric information of three-dimensional point clouds, making it difficult to construct a complete defect characterization model. In the defect assessment link, existing solutions mostly stay at the identification of defect types and locations, lack analysis of the importance of defects in the material structure, and are unable to quantify the potential risks of defects based on the connectivity and stress distribution of key material nodes. It is difficult to meet the needs of defect classification management and priority repair in semiconductor packaging processes. Summary of the Invention

[0005] Based on the above technical problems, this application discloses a semiconductor packaging material defect detection method based on deep learning, including: S1. Collecting a surface image of a semiconductor packaging material, extracting two-dimensional contour information of the image, and obtaining a two-dimensional feature point set using a feature point extraction algorithm; S2. Preprocessing the image containing the two-dimensional contour information and the two-dimensional feature point set, removing image noise using an adaptive median filter, enhancing the image grayscale through histogram normalization, and separating the main area of ​​the packaging material using an image segmentation algorithm; S3. Build a Transformer-based 2D defect recognition model. It takes the pre-processed 2D image of the main area, its contours, and feature point information as input, analyzes the global features of the image through a multi-head attention mechanism, and outputs 2D detection results that include suspected defect areas and unknown feature areas. S4. For semiconductor packaging materials with suspected defects and unknown defect areas in the two-dimensional inspection results, use structured light three-dimensional scanning technology to obtain three-dimensional point cloud data of the corresponding areas and extract three-dimensional geometric features and texture features; S5. Register and associate the feature points in the 3D point cloud data with the feature point set of the 2D image. Through the spatiotemporal alignment model constructed by deep learning, the 2D image information is fused with the 3D point cloud data to form a 2D-3D multimodal data volume. S6. Build a multimodal fusion deep learning defect assessment model, input the multimodal data volume into the model, and output the precise location and type of the defect; S7. Through the constructed defect importance assessment model, the importance of defects in semiconductor packaging materials is evaluated based on the connectivity and stress distribution of key material nodes, and a defect detection report is generated in the form of visual markings.

[0006] Preferably, the two-dimensional contour information of the image is extracted in S1 to obtain a two-dimensional feature point set, specifically: the edge detection is performed on the collected surface image of the semiconductor packaging material by using an improved Canny operator, and the edge is determined by using an adaptive double threshold calculation method, wherein the high threshold and low threshold The calculation formula is: , , where is the image grayscale mean, is the image grayscale standard deviation, 、 are adjustable coefficients to obtain two-dimensional contour information; obtain a two-dimensional feature point set, and use the improved FAST corner detection algorithm and the grayscale gradient direction constraint to detect each pixel. , calculate its surrounding radius Weighted sum of grayscale differences of pixels in the neighborhood , the formula is: ,in, Pixel radius Neighborhood, is the pixel point in the neighborhood The weight coefficient of Exceeding the set threshold When the pixel is the feature point, and a two-dimensional feature point set is obtained.

[0007] Preferably, the image of the two-dimensional contour information and the two-dimensional feature point set is preprocessed in S2, specifically: noise is removed from the image by multi-scale guided filtering, and the similarity weight matrix of the image pixels at different scales is calculated. , for the image For filtering, the formula is: ,in Pixel Neighborhood, is the mean value in the neighborhood; in the process of histogram specification, the target histogram based on the defect feature distribution is used , by calculating the cumulative distribution function Perform grayscale mapping to convert the original image grayscale Mapping to target grayscale , the formula is: ,in is the cumulative distribution function of the image, is the cumulative distribution function of the target image.

[0008] Preferably, the image segmentation algorithm is used in S2 to separate the main area of ​​the packaging material, specifically: combining the two-dimensional contour information to construct the regional growth seed point, by calculating the grayscale similarity and geometric structure similarity between the pixel point and the seed point, and using the weighted fusion strategy to segment the image into the main area of ​​the packaging material, wherein the similarity calculation formula is , 、 are the grayscale values ​​of the pixel to be segmented and the seed point pixel, is the geometric distance from the pixel to be segmented to the seed point, 、 are weight coefficients respectively.

[0009] Preferably, the method for constructing the two-dimensional defect recognition model in S3 is: encoding the two-dimensional image data, two-dimensional contour information and two-dimensional feature point set into feature vectors respectively, calculating the self-attention weights of different modal features in parallel through a multi-head attention mechanism, and using the formula Calculate the attention score, where 、 、 are query, key, and value vectors respectively, is the key vector dimension; using the adaptive gating fusion unit, through the gating function Calculate the fusion weight of each modality feature, where 、 、 are the feature vectors of the two-dimensional image, contour and feature point set respectively, 、 、 is the corresponding weight matrix, is an activation function that passes through the Transformer encoder layer and outputs a two-dimensional detection result containing suspected defect areas and unknown feature areas.

[0010] Preferably, the method for extracting three-dimensional geometric features and texture features in S4 is: for the acquired three-dimensional point cloud data, the local projection entropy LPE algorithm is used to extract the three-dimensional geometric features, and the LPE algorithm is used to calculate the LPE of each point. In its The local projection entropy within the neighborhood is: ,in Yes At the point The probability density of the projection position in the local coordinate system is used to map the three-dimensional point cloud to the two-dimensional image. By establishing a corresponding relationship between the point cloud and the image pixels, the texture information on the two-dimensional image is associated with each three-dimensional point. The multi-scale texture feature extraction network of deep learning is used. The network contains multiple convolution layers with different convolution kernel sizes. The associated two-dimensional image blocks are subjected to feature extraction and the feature vector containing high-frequency and low-frequency texture information is output. The texture feature vectors of different scales are spliced ​​with the three-dimensional geometric feature vector through the feature fusion module to obtain a composite feature vector containing three-dimensional geometric features and texture features.

[0011] Preferably, the construction and operation method of the spatiotemporal alignment model constructed by deep learning in S5 is as follows: a network structure including a dual-branch encoder is constructed, wherein the first branch encoder uses a graph convolutional network to encode the three-dimensional point cloud data, and the three-dimensional point cloud is constructed as an undirected graph. , where the vertex is a point cloud data point, edge For the connection relationship between points, through graph convolution operation Update node features, Vertex Neighborhood, is the normalization constant, For nodes exist The second branch encoder uses a residual convolutional network to encode the two-dimensional image; a cross-modal attention mechanism is introduced to construct a three-dimensional-two-dimensional feature interaction matrix ,in 、 They are the 3D point cloud coding features and the 2D image coding features, and the features are weighted fused by attention weights; in the alignment stage, the spatiotemporal joint loss function is used Perform model training, where is the feature similarity loss, is the spatial position loss, is the feature change consistency loss, 、 、 , are the corresponding weight coefficients respectively. By minimizing the loss function, the spatiotemporal alignment of the two-dimensional image information and the three-dimensional point cloud data is achieved to form a multimodal data body containing a two-dimensional-three-dimensional correspondence relationship.

[0012] Preferably, the multimodal fusion deep learning defect assessment model in S6 is constructed as follows: a dual-path fusion network containing a multi-scale feature pyramid is built. The first path uses a 3D convolutional neural network to perform multi-level extraction of three-dimensional point cloud features in the multimodal data volume to generate a three-dimensional feature pyramid containing different spatial resolutions; the second path uses a void convolutional neural network to densely sample two-dimensional image features to form a two-dimensional feature pyramid; through a cross-modal feature interaction module, a bidirectional attention mechanism is used to calculate the attention weight between three-dimensional and two-dimensional features, and the formula is: , ,in 、 are three-dimensional and two-dimensional feature vectors respectively, 、 Respectively weight matrices; through the gated fusion unit Adaptive fusion of multimodal features is achieved, and the fused features are input into a multi-layer perceptron to output preliminary defect location and type prediction results.

[0013] Preferably, the defect importance assessment model in S7 is constructed as follows: the structural topology of the semiconductor packaging material is converted into a weighted undirected graph, where the nodes represent the key structural units of the material, and the edge weights are composed of the connection strength and stress transfer coefficient between the nodes; the attention weight of each node is calculated by the graph attention mechanism, and the formula is: ,in 、 is the adjacent node feature vector, To query the weight matrix, the stress conduction loss function is used ,in is the measured stress value, Predict stress values ​​for the model by minimizing the loss of structural connectivity , stress distribution loss The joint loss function Training model, where are the corresponding weight coefficients respectively.

[0014] Preferably, the operation method of the defect importance evaluation model in S7 is: calculating the structural connectivity contribution value of the node where the defect is located based on the trained model and stress concentration factor , through the normalized weighted formula Generate defect importance assessment values, where are the corresponding weight coefficients respectively, which realize the evaluation of the importance of defects in semiconductor packaging materials.

[0015] Compared with the prior art, the technical solution of this application has the following technical effects: This paper uses an improved Canny operator and the FAST corner detection algorithm to accurately extract two-dimensional contours and feature points from images of semiconductor packaging material surfaces. High-threshold and low-threshold adaptive calculation methods (based on image grayscale mean and standard deviation) dynamically adapt to images with varying lighting and texture conditions, significantly improving contour integrity and noise immunity compared to traditional fixed-threshold edge detection. Multi-scale guided filtering combined with an adaptive weight matrix removes noise while preserving image detail. Histogram regularization customizes target grayscale mapping based on defect feature distribution, enhancing defect-background contrast. A region growing algorithm, combining grayscale and geometric structure similarity, accurately segments the bulk of the packaging material, reduces background interference, and provides a high-quality data foundation for subsequent defect identification, reducing edge localization error to sub-pixel levels.

[0016] The Transformer-based two-dimensional defect recognition model of the present invention analyzes the global features of the image through a multi-head attention mechanism, breaking through the local receptive field limitation of traditional convolutional neural networks, capturing the long-distance dependency between defects and surrounding structures, and effectively detecting tiny defects in complex textures (such as micron-level cracks); the adaptive gated fusion unit realizes cross-modal information interaction of two-dimensional images, contours, and feature points, improving the detection rate of suspected defective areas; the combination of structured light three-dimensional scanning and the local projection entropy algorithm can obtain the three-dimensional geometric features (such as curvature, convexity and concavity) and multi-scale texture features of the defective area; the spatiotemporal alignment model solves the spatial misalignment problem of two-dimensional and three-dimensional data through a dual-branch encoder and a cross-modal attention mechanism, with a registration accuracy of millimeter level. The formed multimodal data body can fully characterize the three-dimensional morphology and texture details of the defect, providing multi-dimensional information support for accurate defect assessment.

[0017] The multimodal fusion deep learning model proposed in this paper constructs a dual-path feature pyramid using a 3D convolutional neural network and a dilated convolutional neural network. This model, combined with a bidirectional attention mechanism, enables deep interaction between 3D geometric and 2D texture features. A gated fusion unit adaptively assigns weights based on feature importance, minimizing defect location errors and improving type classification accuracy. Compared to single-modal models, this multimodal fusion approach improves the recognition accuracy of complex defects (such as mixed delamination and crack defects) by 40%. The defect importance assessment model transforms the material structure topology into a weighted undirected graph. Using a graph attention mechanism and a stress conduction loss function, it quantifies the connectivity contribution and stress concentration coefficient of the defect node. The resulting defect importance assessment reflects the actual impact of defects on the material's mechanical properties, providing a basis for prioritization for process optimization. This shifts defect management from qualitative identification to quantitative assessment, significantly improving semiconductor packaging reliability prediction capabilities.

[0018] The inspection process constructed by this invention is fully automated, from image acquisition and feature extraction to defect assessment, meeting the real-time inspection requirements of semiconductor production lines. Through a spatiotemporal alignment model and multimodal fusion technology, a seamless integration of two-dimensional images and three-dimensional point cloud data is achieved, generating visually labeled reports containing defect location, type, and importance. The inspection results are intuitive and easy to understand, allowing process personnel to quickly locate problem areas. The defect importance assessment module, combined with the stress distribution of key material nodes, can predict the potential risks of defects in long-term service and assist in the development of preventive maintenance strategies.

[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application so that it can be implemented in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following is a detailed description of the preferred embodiment of the present application in conjunction with the accompanying drawings.

[0020] Based on the detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings below, those skilled in the art will become more aware of the above and other objects, advantages and features of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without inventive work. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.

[0022] Figure 1This is a flow chart of a semiconductor packaging material defect detection method based on deep learning in the present invention; Figure 2 This is a dual-branch network architecture diagram of the spatiotemporal alignment model for semiconductor packaging material defect detection according to the present invention; Figure 3 A histogram comparing the corner detection performance of the improved Canny operator and FAST in an embodiment of the present invention; Figure 4 This is a comparison curve of the crack defect IoU distribution between the Transformer model in an embodiment of the present invention and the traditional CNN; Figure 5 A comparison diagram of the registration error and positioning error probability density between the spatiotemporal alignment model in the embodiment of the present invention and the prior art; Figure 6 This is a bar chart comparing the recognition rates of mixed defects (cracks + delamination) in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. In the following description, specific details such as specific configurations and components are provided only to help fully understand the embodiments of the present application. Therefore, it should be clear to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, for clarity and brevity, the description of known functions and structures has been omitted in the embodiments.

[0024] It should be understood that references throughout this specification to "one embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearance of "one embodiment" or "this embodiment" throughout this specification does not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0025] In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0026] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" in this article describes another type of association object relationship, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0027] The term "at least one" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0028] It should also be noted that, in this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include," "comprises," or any other variations thereof are intended to cover non-exclusive inclusion.

[0029] Example 1

[0030] This embodiment mainly describes a semiconductor packaging material defect detection method based on deep learning, such as Figure 1 Shown, including: S1. Collecting a surface image of a semiconductor packaging material, extracting two-dimensional contour information of the image, and obtaining a two-dimensional feature point set using a feature point extraction algorithm; S2. Preprocessing the image containing the two-dimensional contour information and the two-dimensional feature point set, removing image noise using an adaptive median filter, enhancing the image grayscale through histogram normalization, and separating the main area of ​​the packaging material using an image segmentation algorithm; S3. Build a Transformer-based 2D defect recognition model. It takes the pre-processed 2D image of the main area, its contours, and feature point information as input, analyzes the global features of the image through a multi-head attention mechanism, and outputs 2D detection results that include suspected defect areas and unknown feature areas. S4. For semiconductor packaging materials with suspected defects and unknown defect areas in the two-dimensional inspection results, use structured light three-dimensional scanning technology to obtain three-dimensional point cloud data of the corresponding areas and extract three-dimensional geometric features and texture features; S5. Register and associate the feature points in the 3D point cloud data with the feature point set of the 2D image. Through the spatiotemporal alignment model constructed by deep learning, the 2D image information is fused with the 3D point cloud data to form a 2D-3D multimodal data volume. S6. Build a multimodal fusion deep learning defect assessment model, input the multimodal data volume into the model, and output the precise location and type of the defect; S7. Through the constructed defect importance assessment model, the importance of defects in semiconductor packaging materials is evaluated based on the connectivity and stress distribution of key material nodes, and a defect detection report is generated in the form of visual markings.

[0031] Furthermore, S1 extracts the two-dimensional contour information of the image and obtains a two-dimensional feature point set. Specifically, the surface image of the collected semiconductor packaging material is edge detected by the improved Canny operator, and the edge is determined by the adaptive double threshold calculation method, where the high threshold and low threshold The calculation formula is: , , where is the image grayscale mean, is the image grayscale standard deviation, 、 are adjustable coefficients to obtain two-dimensional contour information; obtain a two-dimensional feature point set, and use the improved FAST corner detection algorithm and the grayscale gradient direction constraint to detect each pixel. , calculate its surrounding radius Weighted sum of grayscale differences of pixels in the neighborhood , the formula is: ,in, Pixel radius Neighborhood, is the pixel point in the neighborhood The weight coefficient of Exceeding the set threshold When the pixel is the feature point, and a two-dimensional feature point set is obtained.

[0032] Furthermore, the image of the two-dimensional contour information and the two-dimensional feature point set is preprocessed in S2, specifically: the image is denoised by multi-scale guided filtering, and the similarity weight matrix of the image pixels at different scales is calculated. , for the image For filtering, the formula is: ,in Pixel Neighborhood, is the mean value in the neighborhood; in the process of histogram specification, the target histogram based on the defect feature distribution is used , by calculating the cumulative distribution function Perform grayscale mapping to convert the original image grayscale Mapping to target grayscale , the formula is: ,in is the cumulative distribution function of the image, is the cumulative distribution function of the target image.

[0033] Furthermore, S2 uses an image segmentation algorithm to separate the main area of ​​the packaging material. Specifically, the two-dimensional contour information is combined to construct the regional growth seed point, and the grayscale similarity and geometric structure similarity between the pixel point and the seed point are calculated. The weighted fusion strategy is used to segment the image into the main area of ​​the packaging material. The similarity calculation formula is: , 、 are the grayscale values ​​of the pixel to be segmented and the seed point pixel, is the geometric distance from the pixel to be segmented to the seed point, 、 are weight coefficients respectively.

[0034] Furthermore, the two-dimensional defect recognition model in S3 is constructed by encoding the two-dimensional image data, two-dimensional contour information and two-dimensional feature point set into feature vectors respectively, and calculating the self-attention weights of different modal features in parallel through the multi-head attention mechanism, and using the formula Calculate the attention score, where 、 、 are query, key, and value vectors respectively, is the key vector dimension; using the adaptive gating fusion unit, through the gating function Calculate the fusion weight of each modality feature, where 、 、 are the feature vectors of the two-dimensional image, contour and feature point set respectively, 、 、 is the corresponding weight matrix, As the activation function, the Transformer encoder layer is used to output the two-dimensional detection results containing suspected defect areas and unknown feature areas. Furthermore, the method of extracting three-dimensional geometric features and texture features in S4 is as follows: for the acquired three-dimensional point cloud data, the local projection entropy LPE algorithm is used to extract three-dimensional geometric features, and the LPE algorithm is used to calculate the LPE of each point. In its The local projection entropy within the neighborhood is: ,in Yes At the point The probability density of the projection position in the local coordinate system is used to map the three-dimensional point cloud to the two-dimensional image. By establishing a corresponding relationship between the point cloud and the image pixels, the texture information on the two-dimensional image is associated with each three-dimensional point. The multi-scale texture feature extraction network of deep learning is used. The network contains multiple convolution layers with different convolution kernel sizes. The associated two-dimensional image blocks are subjected to feature extraction and the feature vector containing high-frequency and low-frequency texture information is output. The texture feature vectors of different scales are spliced ​​with the three-dimensional geometric feature vector through the feature fusion module to obtain a composite feature vector containing three-dimensional geometric features and texture features.

[0035] Further, if Figure 2 As shown in Figure 5, the construction and operation method of the spatiotemporal alignment model constructed by deep learning in S5 is as follows: a network structure including a dual-branch encoder is constructed. The first branch encoder uses a graph convolutional network to encode the 3D point cloud data and constructs the 3D point cloud into an undirected graph. , where the vertex is a point cloud data point, edge For the connection relationship between points, through graph convolution operation Update node features, Vertex Neighborhood, is the normalization constant, For nodes exist The second branch encoder uses a residual convolutional network to encode the two-dimensional image; a cross-modal attention mechanism is introduced to construct a three-dimensional-two-dimensional feature interaction matrix ,in 、 They are the 3D point cloud coding features and the 2D image coding features, and the features are weighted fused by attention weights; in the alignment stage, the spatiotemporal joint loss function is used Perform model training, where is the feature similarity loss, is the spatial position loss, is the feature change consistency loss, 、 、 are the corresponding weight coefficients respectively, and the spatiotemporal alignment of two-dimensional image information and three-dimensional point cloud data is achieved by minimizing the loss function to form a multimodal data body containing two-dimensional-three-dimensional correspondence.

[0036] Furthermore, the multimodal fusion deep learning defect assessment model in S6 is constructed as follows: a dual-path fusion network containing a multi-scale feature pyramid is built. The first path uses a 3D convolutional neural network to perform multi-level extraction of three-dimensional point cloud features in the multimodal data volume to generate a three-dimensional feature pyramid with different spatial resolutions; the second path uses a dilated convolutional neural network to densely sample two-dimensional image features to form a two-dimensional feature pyramid; through the cross-modal feature interaction module, the bidirectional attention mechanism is used to calculate the attention weight between three-dimensional and two-dimensional features. The formula is: , ,in 、 are three-dimensional and two-dimensional feature vectors respectively, 、 Respectively weight matrices; through the gated fusion unit Adaptive fusion of multimodal features is achieved, and the fused features are input into a multi-layer perceptron to output preliminary defect location and type prediction results.

[0037] Furthermore, the defect importance assessment model in S7 is constructed as follows: the structural topology of the semiconductor packaging material is converted into a weighted undirected graph, where the nodes represent the key structural units of the material and the edge weights are composed of the connection strength and stress transfer coefficient between nodes; the attention weight of each node is calculated through the graph attention mechanism, and the formula is: ,in 、 is the adjacent node feature vector, To query the weight matrix, the stress conduction loss function is used ,in is the measured stress value, Predict stress values ​​for the model by minimizing the loss of structural connectivity , stress distribution loss The joint loss function Training model, where are the corresponding weight coefficients respectively.

[0038] Furthermore, the operation method of the defect importance evaluation model in S7 is: based on the trained model, the structural connectivity contribution value of the node where the defect is located is calculated. and stress concentration factor , through the normalized weighted formula Generate defect importance assessment values, where are the corresponding weight coefficients respectively, which realize the evaluation of the importance of defects in semiconductor packaging materials.

[0039] This embodiment describes in detail the present invention, which combines an improved Canny operator with the FAST corner detection algorithm to achieve high-precision two-dimensional contour and feature point extraction. It utilizes a two-dimensional defect recognition model based on the Transformer architecture to capture global image features through a multi-head attention mechanism, effectively improving the detection rate of suspected defect areas. In terms of three-dimensional information acquisition, it introduces structured light three-dimensional scanning technology and the local projection entropy (LPE) algorithm to achieve composite extraction of three-dimensional geometric features and texture features of defect areas, and solves the spatial registration problem of two-dimensional and three-dimensional data through a spatiotemporal alignment model. Through a multimodal fusion deep learning model, it utilizes a bidirectional attention mechanism and a gated fusion unit to achieve adaptive interaction of cross-modal features, significantly improving the accuracy of defect localization and classification. The constructed defect importance assessment model transforms the material structure topology into a weighted undirected graph, comprehensively considers node connectivity and stress distribution, and achieves quantitative assessment of defect impact, providing a scientific basis for defect management in semiconductor packaging processes and filling the gap in the existing technology in the field of dynamic defect risk assessment.

[0040] Based on Example 1, this implementation describes in detail the specific implementation effects of this application, specifically: The FC-BGA package substrate produced by a semiconductor company was selected as the experimental object. The substrate contains a multi-layer structure including epoxy molding compound (EMC), copper interconnect layer and silicon chip. Common defect types include molding compound cracks (width 0.5-3 ), copper layer voids (diameter 2-10 )、Interface stratification (area 10-50 The experiment collected 1,200 images of the package substrate surface (resolution 2,048 × 2,048 pixels, including 400 normal samples, 400 cracked samples, and 400 void / delamination samples). A structured light 3D scanner was used to obtain 3D point cloud data of 300 defective areas (point cloud density 500 points / mm²).

[0041] Experimental environment and parameter settings, hardware platform: NVIDIA RTX4090 GPU (24GB video memory), Intel i9-13900K CPU, 64GB RAM; software tools: Python 3.9, PyTorch 2.1, Open3D 0.18.0; key parameters: Improved Canny operator: =1.2, =0.5, grayscale mean A and standard deviation B are dynamically calculated; FAST corner detection: radius =3, threshold =40, weight coefficient Using Gaussian distribution; spatiotemporal alignment model: feature similarity loss Weight =0.6, spatial position loss Weight =0.3, feature change consistency loss Weight =0.1; This application performs edge detection and corner extraction on the collected two-dimensional image. The contour integrity detected by the improved Canny operator reaches 98.7%, which is higher than the traditional Canny operator (fixed threshold =150, =50) increased by 12.8%; after the FAST corner detection algorithm introduced the grayscale gradient direction constraint, the feature point positioning error was reduced from 2.3 pixels to 0.8 pixels, and a total of about 1500-2000 corner points were extracted per image (such as Figure 3 shown); In the preprocessing stage, multi-scale guided filtering (with scale factors of 1, 3, and 5) was used. The noise standard deviation was reduced from 25.6 to 4.2, and the peak signal-to-noise ratio (PSNR) was increased from 28.3 dB to 36.5 dB. After histogram normalization, the grayscale entropy of the defect area increased from 6.8 bits to 8.9 bits, and the contrast was improved by 45% (see Table 1). Table 1 Data on the impact of preprocessing steps on image quality Preprocessing steps Noise standard deviation PSNR (dB) Grayscale entropy (bits) Contrast improvement Original image 25.6 28.3 6.8 - Multiscale filtering 4.2 36.5 7.5 22% Histogram Normalization - 35.8 8.9 45% For 2D defect recognition and 3D data acquisition, the Transformer-based 2D defect recognition model in this application achieved a detection rate of 96.2% for suspected defect areas and a recall rate of 95.8% after 200 epochs of training. Taking cracked samples as an example, the intersection over union (IoU) between the suspected defect areas output by the model and the actual crack locations reached 0.89±0.04, an improvement of 23.6% compared to the traditional CNN model (IoU=0.72±0.06). Figure 4 shown); Structured light 3D scanning was performed on 180 suspected defect areas (including 60 cracks, 60 voids, and 60 delaminations) detected in 2D detection to obtain 3D point cloud data. The local projection entropy (LPE) algorithm was used to extract geometric features. The LPE value for normal areas was 1.2±0.1, for crack areas LPE=2.8±0.3, for void areas LPE=3.5±0.4, and for delamination areas LPE=2.1±0.2. The geometric features of different defect types were significantly different (see Table 2).

[0042] Table 2: Statistics of 3D geometric and texture features of different defect types

[0043] For multimodal data fusion and defect assessment, after training the spatiotemporal alignment model, the registration error between 2D feature points and 3D point clouds is 0.08±0.02mm, which is 68% lower than the traditional ICP algorithm (error 0.25±0.05mm). The multimodal fusion model has a positioning error of 42.3±5.6mm for defect locations. The accuracy of type classification reached 93.1%, among which the recognition accuracy of mixed defects (such as cracks + delamination) was improved from 52% of the existing technology to 89% (such as Figure 5 shown).

[0044] The defect importance assessment model is trained based on stress distribution data of 100 key nodes. The root mean square error (RMSE) between the predicted stress value and the measured value is 4.2MPa, and the structural connectivity contribution value is and stress concentration factor The normalized weighted evaluation value of It can effectively distinguish the risk level of defects; for example, the risk level of a crack defect =0.78, =1.65, =0.89, which is a high-risk defect; a certain void defect =0.32, =1.12, =0.51, which is a medium-risk defect. Four defects are randomly selected to judge the level of risk (see Table 3).

[0045] Table 3: Example of output results of defect importance assessment model

[0046] The proposed solution takes 1.28 ± 0.15 seconds per sample, representing a 700-fold improvement in efficiency compared to manual inspection (average 15 minutes per sample). The generated visual report, including a 2D defect marker map, a 3D point cloud rendering, and a risk heat map, demonstrates a defect location annotation error of less than 0.1 mm, and a correlation of 0.91 between importance level and process failure probability.

[0047] A traditional 2D inspection solution (CNN-based) using a ResNet-50 network tested the same sample, achieving an 82.5% detection rate for suspected defect areas, a 78.9% recall rate, and an Intersection of Union (IoU) of 0.72 ± 0.06. Classification accuracy for 300 defect samples was 75.3%, with a 28% false positive rate for void defects (often misclassified as delamination) and a mere 52% accuracy rate for mixed defects. Single-sample inspection took 2.15 ± 0.23 seconds, resulting in a PSNR of 32.1 dB and a noise standard deviation of 7.8 for the preprocessed image.

[0048] A single 3D inspection solution (based on structured light scanning) uses only 3D point cloud data for defect detection. It employs traditional geometric features (such as curvature and height difference) for classification, with an accuracy rate of 68.7%. This approach is unable to detect surface texture defects (such as microcracks). The 3D coordinate error is 0.25±0.05mm, making it impossible to establish a precise correspondence with the 2D image. Inspection reports, which only contain point cloud data, lack intuitiveness.

[0049] Multimodal fusion without importance evaluation scheme, using early fusion strategy (directly splicing 2D features and 3D features), defect classification accuracy is 81.2%, positioning error is 68.5±8.2 , it is impossible to output the defect risk level, and the missed rate of high-risk defects is as high as 35%.

[0050] The core performance comparison between the prior art and the present application is shown in Table 4. Table 4: Comparison of core performance between this application solution and existing technology index This application plan Traditional CNN 2D detection Single 3D detection Early multimodal fusion Defect detection rate 96.2% 82.5% 75.0% 88.3% Classification accuracy 93.1% 75.3% 68.7% 81.2% Positioning error (μm) 42.3±5.6 120.5±15.2 250.0±30.0 68.5±8.2 Mixed defect recognition rate 89% 52% - 65% Single sample time (seconds) 1.28±0.15 2.15±0.23 3.87±0.41 2.79±0.32 Importance Assessment Correlation 0.91 - - - Table 4 shows the differences in core performance indicators between this application solution and traditional CNN two-dimensional detection, single three-dimensional detection, and early multimodal fusion solutions. Data shows that the defect detection rate of this application solution reached 96.2%, significantly higher than the 82.5% of traditional CNN two-dimensional detection, 75.0% of single three-dimensional detection, and 88.3% of early multimodal fusion; the classification accuracy was 93.1%, far exceeding other solutions (75.3%, 68.7%, and 81.2%, respectively); the positioning error was only 42.3±5.6μm, significantly lower than the 120.5±15.2μm of traditional CNN two-dimensional detection and 250.0±30.0μm of single three-dimensional detection. Compared to the 68.5±8.2μm achieved by earlier multimodal fusion methods, the proposed solution achieved an 89% hybrid defect recognition rate, significantly higher than the 52% achieved by traditional CNN two-dimensional detection and the 65% achieved by earlier multimodal fusion methods. The single-sample detection time was only 1.28±0.15 seconds, outperforming other solutions (2.15±0.23 seconds, 3.87±0.41 seconds, and 2.79±0.32 seconds, respectively). Furthermore, the proposed solution achieved a defect importance assessment correlation of 0.91, while other solutions did not involve importance assessment. Overall, the proposed solution demonstrated significant advantages across all key metrics, demonstrating its advancement and effectiveness in semiconductor packaging material defect detection.

[0051] This embodiment details the technology of this application that improves image quality through multi-dimensional feature extraction and preprocessing, accurately locates defects through two-dimensional defect recognition and multimodal fusion, realizes stereoscopic characterization by combining three-dimensional data, and improves classification accuracy to 93.1% through multimodal model. Defect importance assessment quantifies risks, and full process automation increases efficiency by about 700 times, which is significantly better than existing technologies.

[0052] The above are only preferred embodiments of the present invention, which do not limit the scope of protection of the present invention. For those skilled in the art, the present invention can be modified and varied in various ways. Any changes, modifications, replacements, integrations and parameter changes to these embodiments through conventional substitutions or that can achieve the same functions without departing from the principles and spirit of the present invention fall within the scope of protection of the present invention.

Claims

1. A semiconductor packaging material defect detection method based on deep learning, characterized in that: include: S1. Collecting a surface image of a semiconductor packaging material, extracting two-dimensional contour information of the image, and obtaining a two-dimensional feature point set using a feature point extraction algorithm; S2. Preprocessing the image containing the two-dimensional contour information and the two-dimensional feature point set, using adaptive median filtering to remove image noise, achieving image grayscale enhancement through histogram normalization, and using an image segmentation algorithm to separate the main area of ​​the packaging material; S3. Build a Transformer-based 2D defect recognition model. This model takes the pre-processed 2D image of the main area, its contours, and feature point information as input, analyzes the global features of the image through a multi-head attention mechanism, and outputs 2D detection results that include suspected defect areas and unknown feature areas. S4. For semiconductor packaging materials with suspected defects and unknown defect areas in the two-dimensional inspection results, use structured light three-dimensional scanning technology to obtain three-dimensional point cloud data of the corresponding areas and extract three-dimensional geometric features and texture features; S5. Register and associate the feature points in the 3D point cloud data with the feature point set of the 2D image. Through the spatiotemporal alignment model constructed by deep learning, the 2D image information is fused with the 3D point cloud data to form a 2D-3D multimodal data volume. S6. Build a multimodal fusion deep learning defect assessment model, input the multimodal data volume into the model, and output the precise location and type of the defect; S7. Through the constructed defect importance assessment model, the importance of defects in semiconductor packaging materials is evaluated based on the connectivity and stress distribution of key material nodes, and a defect detection report is generated in the form of visual markings.

2. The semiconductor packaging material defect detection method based on deep learning according to claim 1, characterized in that: In the step S1, the two-dimensional contour information of the image is extracted to obtain a two-dimensional feature point set. Specifically, the edge detection is performed on the surface image of the semiconductor packaging material collected by the improved Canny operator, and the edge is determined by the adaptive double-threshold calculation method, wherein the high threshold and low threshold The calculation formula is: , , where is the image grayscale mean, is the image grayscale standard deviation, 、 are adjustable coefficients to obtain two-dimensional contour information; obtain a two-dimensional feature point set, and use the improved FAST corner detection algorithm and the grayscale gradient direction constraint to detect each pixel. , calculate its surrounding radius Weighted sum of grayscale differences of pixels in the neighborhood , the formula is: ,in, Pixel radius Neighborhood, is the pixel point in the neighborhood The weight coefficient of Pixel The gray value of Pixel radius Pixels in the neighborhood The gray value of Exceeding the set threshold When the pixel is the feature point, and a two-dimensional feature point set is obtained.

3. The semiconductor packaging material defect detection method based on deep learning according to claim 1, characterized in that: The two-dimensional contour information and the image of the two-dimensional feature point set in S2 are preprocessed, specifically: the image is noise-removed by multi-scale guided filtering, and the similarity weight matrix of the image pixels at different scales is calculated. , for the image For filtering, the formula is: ,in is the output image pixel value, Pixel Neighborhood, is the pixel point in the original image Neighborhood Pixels within The gray value of is the mean value in the neighborhood; In the process of histogram specification, the target histogram based on the defect feature distribution is used , by calculating the cumulative distribution function Perform grayscale mapping to convert the original image grayscale Mapping to target grayscale , the formula is: ,in is the cumulative distribution function of the image, is the cumulative distribution function of the target image.

4. A semiconductor packaging material defect detection method based on deep learning according to claim 1 or 3, characterized in that: In S2, an image segmentation algorithm is used to separate the main area of ​​the packaging material. Specifically, the following steps are used: combining two-dimensional contour information to construct regional growth seed points, calculating the grayscale similarity and geometric structure similarity between the pixel point and the seed point, and using a weighted fusion strategy to segment the image into the main area of ​​the packaging material, wherein the similarity calculation formula is: , 、 are the grayscale values ​​of the pixel to be segmented and the seed point pixel, is the maximum grayscale value of all pixels in the image. is the geometric distance from the pixel to be segmented to the seed point, is the maximum geometric distance from the pixel to be segmented to the seed point, 、 are weight coefficients respectively.

5. The method for detecting semiconductor packaging material defects based on deep learning according to claim 1, wherein: The method for constructing the two-dimensional defect recognition model in S3 is as follows: the two-dimensional image data, the two-dimensional contour information and the two-dimensional feature point set are respectively encoded into feature vectors, and the self-attention weights of different modal features are calculated in parallel through the multi-head attention mechanism. Calculate the attention score, where 、 、 are query, key, and value vectors respectively, is the key vector dimension; using the adaptive gating fusion unit, through the gating function Calculate the fusion weight of each modality feature, where 、 、 are the feature vectors of the two-dimensional image, contour and feature point set respectively, 、 、 are the corresponding weight matrices, is an activation function that passes through the Transformer encoder layer and outputs a two-dimensional detection result containing suspected defect areas and unknown feature areas.

6. The method for detecting semiconductor packaging material defects based on deep learning according to claim 1, wherein: The method of extracting 3D geometric features and texture features in S4 is as follows: for the acquired 3D point cloud data, the local projection entropy LPE algorithm is used to extract 3D geometric features, and the LPE algorithm is used to calculate the LPE of each point. In its The local projection entropy within the neighborhood is: ,in Yes At the point The probability density of the projected position in the local coordinate system, For each point in the 3D point cloud data The number of neighborhood points is calculated, and the three-dimensional point cloud and the two-dimensional image are mapped to coordinates. By establishing a corresponding relationship between the point cloud and the image pixels, the texture information on the two-dimensional image is associated with each three-dimensional point. The multi-scale texture feature extraction network of deep learning is used. The network contains multiple convolution layers with different convolution kernel sizes. The associated two-dimensional image blocks are extracted and a feature vector containing high-frequency and low-frequency texture information is output. The texture feature vectors of different scales are spliced ​​with the three-dimensional geometric feature vector through the feature fusion module to obtain a composite feature vector containing three-dimensional geometric features and texture features.

7. The method for semiconductor packaging material defect detection based on deep learning according to claim 1, characterized in that: In S5, a spatiotemporal alignment model is constructed through deep learning. Specifically, a network structure including a dual-branch encoder is constructed. The first branch encoder uses a graph convolutional network to encode the 3D point cloud data and constructs the 3D point cloud into an undirected graph. , where the vertex is a point cloud data point, edge For the connection relationship between points, through graph convolution operation Update node features, Vertex Neighborhood, is the normalization constant, For nodes exist Layer characteristics, For the The bias term of the graph convolution layer, is the activation function; the second branch encoder uses a residual convolutional network to encode the two-dimensional image; a cross-modal attention mechanism is introduced to construct a three-dimensional-two-dimensional feature interaction matrix ,in 、 They are the first The feature vector of the point and the The feature vector of each pixel, Represents the 3D point cloud points and two-dimensional images The feature vector of each pixel is weighted and fused by attention weights; In the alignment stage, the spatiotemporal joint loss function is used Perform model training, where is the feature similarity loss, is the spatial position loss, is the feature change consistency loss, 、 、 are the corresponding weight coefficients respectively, and the spatiotemporal alignment of two-dimensional image information and three-dimensional point cloud data is achieved by minimizing the loss function to form a multimodal data body containing two-dimensional-three-dimensional correspondence.

8. The method for semiconductor packaging material defect detection based on deep learning according to claim 1, characterized in that: The multimodal fusion deep learning defect assessment model in S6 is constructed specifically as follows: a dual-path fusion network containing a multi-scale feature pyramid is built. The first path uses a 3D convolutional neural network to perform multi-level extraction of three-dimensional point cloud features in the multimodal data volume to generate a three-dimensional feature pyramid with different spatial resolutions; the second path uses a void convolutional neural network to densely sample two-dimensional image features to form a two-dimensional feature pyramid; through a cross-modal feature interaction module, a bidirectional attention mechanism is used to calculate the attention weight between three-dimensional and two-dimensional features. The formula is: , ,in 、 are three-dimensional and two-dimensional feature vectors respectively, 、 Respectively weight matrices; through the gated fusion unit Adaptive fusion of multimodal features is achieved, and the fused features are input into a multi-layer perceptron to output preliminary defect location and type prediction results.

9. The method for semiconductor packaging material defect detection based on deep learning according to claim 1, characterized in that: The defect importance assessment model in S7 is constructed as follows: the structural topology of the semiconductor packaging material is converted into a weighted undirected graph, where the nodes represent the key structural units of the material and the edge weights are composed of the connection strength and stress transfer coefficient between nodes; the attention weight of each node is calculated through the graph attention mechanism, and the formula is: ,in 、 is the adjacent node feature vector, To query the weight matrix, the stress conduction loss function is used ,in is the measured stress value, Predict stress values ​​for the model by minimizing the loss of structural connectivity , stress distribution loss The joint loss function Training model, where are the corresponding weight coefficients respectively.

10. The semiconductor packaging material defect detection method based on deep learning according to claim 9, characterized in that: The operation method of the defect importance evaluation model in S7 is: based on the trained model, the structural connectivity contribution value of the node where the defect is located is calculated. and stress concentration factor , through the normalized weighted formula Generate defect importance assessment values, where are the corresponding weight coefficients respectively, which realize the evaluation of the importance of defects in semiconductor packaging materials.

Citation Information

Cited By

  • Surface anomaly detection method based on multi-modal dynamic fusion and double-path cooperative detection

    CN120726054A

  • Surface anomaly detection method based on multi-modal dynamic fusion and double-path collaborative detection

    CN120726054B

  • Quality detection system and method for automobile bumper

    CN120741789A

  • Medical catheter coating quality detection method based on machine vision

    CN120833524A

  • Projection welding workpiece surface defect detection method and system

    CN120852419A