Chip packaging condition nondestructive analysis system and method based on microscopic detection

By acquiring and fusing visible light and infrared image sequences, using deep learning models to extract and enhance the chip surface and internal features, the problem of difficulty in comprehensively evaluating the chip packaging status in the prior art is solved, and lossless and accurate packaging quality evaluation is achieved.

CN120259279AActive Publication Date: 2025-07-04SHANGHAI MAGIC PHOTOELECTRIC TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510688195.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-04
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing chip packaging detection methods are difficult to achieve lossless and accurate evaluation of the surface and internal structure of chip packaging, especially when complex and diverse packaging defects, the robustness and accuracy of traditional image processing or single-modal analysis methods are insufficient, and multi-dimensional information cannot be fully obtained.

Method used

By acquiring visible light and infrared image sequences, image preprocessing and spatial registration are performed, and visual features of the chip surface state and internal structure are extracted and enhanced by using a multi-input deep learning model, and organically fuse it, and finally the packaging state is judged through an intelligent classifier.

Benefits of technology

It realizes efficient and accurate evaluation of the packaging quality without destroying the chip, and can automatically identify and judge the packaging status, improving the robustness and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259279A_ABST
    Figure CN120259279A_ABST
Patent Text Reader

Abstract

The invention provides a chip packaging condition nondestructive analysis system and method based on microscopic detection, and relates to the field of intelligent analysis, and the method comprises the steps: firstly obtaining an original visible light image and an original infrared image sequence, and carrying out the image preprocessing and spatial registration, so as to ensure that information from different light sources and different depths can be mutually correlated in space; secondly, key visual features representing the surface state and the internal structure of the chip are extracted, refined and enhanced, and therefore the distinguishing capacity and robustness are improved; and finally, organically fusing the processed and enhanced surface and internal features, automatically identifying and judging the packaging state of the chip through a trained intelligent classifier, and giving specific categories (such as normality, cavities, cracks and the like) and the reliability degree thereof, so that the packaging quality of the chip is evaluated on the premise that the chip is not damaged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent analysis, and more particularly, in an embodiment of the present application, to a non-destructive analysis system and method for chip packaging conditions based on microscopic inspection. Background Art

[0002] As the core of modern electronic devices, the quality and reliability of integrated circuit (IC) chips are of vital importance. The packaging of chips is a key link in forming the final product, protecting the internal circuit from the external environment and providing electrical connections. Any defects or abnormalities in the packaging process may cause the chip performance to deteriorate or even fail. Therefore, effective detection of chip packaging conditions is a necessary means to ensure product quality, improve production yield and reduce after-sales costs. Traditional chip packaging detection methods often rely on manual visual inspection, destructive physical analysis (such as slice analysis) or special detection equipment. Manual visual inspection is inefficient, subjective and difficult to detect deep or subtle defects. Although destructive analysis can provide detailed internal structure information, it cannot be used for comprehensive inspection of batch products and is time-consuming and labor-intensive. With the miniaturization of chip size and the complexity of packaging structure, the demand for efficient, non-destructive, accurate and multi-dimensional information detection technology is becoming increasingly urgent. Especially for application scenarios that pursue high reliability, the ability to quickly and batch evaluate the chip packaging health, including its surface integrity and potential problems of internal structure, has become a technical problem that needs to be solved.

[0003] Existing chip packaging condition analysis technology based on microscopic inspection, if it only relies on visible light microscopy, its detection range is usually limited to the surface or shallow structure of the chip package. It is powerless to detect defects under the packaging material and internal solder joints, wiring and other structures, and cannot achieve true non-destructive internal inspection. Although some non-destructive testing methods such as ultrasonic microscopy or X-ray detection can detect internal structures, they often require specific equipment and imaging principles, and may be difficult to combine with high-resolution visible light microscopy to provide comprehensive information. In addition, the image data generated by these methods may have their own characteristics. How to effectively integrate data from different modalities to achieve a comprehensive assessment of the chip packaging condition is a major challenge facing existing technologies. Simply analyzing different modal images independently or shallowly fusing them is difficult to fully explore the complementary information hidden in multi-source data. Especially when faced with complex and diverse packaging defects, the robustness and accuracy of traditional image processing or single modality analysis methods are often insufficient to meet the needs of industrial-level inspection.

[0004] Therefore, an optimized non-destructive analysis system of chip packaging status based on microscopic inspection is desired. Summary of the invention

[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a non-destructive analysis system and method for chip packaging conditions based on microscopic detection. First, the original visible light image and the original infrared image sequence are acquired and image preprocessing and spatial registration are performed to ensure that information from different light sources and different depths can be spatially correlated. Then, the key visual features representing the surface state and internal structure of the chip are extracted, refined, and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically fused and passed through a trained intelligent classifier to automatically identify and judge the packaging state of the chip, give the specific category (such as normal, void, crack, etc.) and its reliability, so as to evaluate the packaging quality of the chip without damaging the chip.

[0006] According to one aspect of the present application, there is provided a non-destructive analysis method for chip packaging conditions based on microscopic detection, which includes: Acquire the original visible light image and the original infrared image sequence of the target chip sample; After performing image preprocessing on the original visible light image and the original infrared image sequence respectively, perform spatial registration on them to obtain the registered chip sample visible light image and the registered chip sample infrared volume data; Extract features and perform same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced visual features of the chip sample surface state and enhanced visual features of the chip sample internal structure; Based on the multi-modal fusion representation information of the chip sample state between the enhanced visual features of the chip sample surface state and the enhanced visual features of the chip sample internal structure, determine the packaging state detection result, and the packaging state detection result includes pre-defined packaging state category labels and their confidence levels.

[0007] In the above non-destructive analysis method for chip packaging status based on microscopic detection, feature extraction and homogeneous paradigm feature enhancement are performed on the registered visible light image of the chip sample and the registered infrared volume data of the chip sample through a multi-input deep learning model with two branches, including: inputting the registered visible light image of the chip sample and the registered infrared volume data of the chip sample into the multi-input deep learning model including a first feature extraction branch and a second feature extraction branch respectively to obtain a visual feature encoding map of the chip sample surface state and a visual feature encoding map of the chip sample internal structure; performing homogeneous paradigm feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure respectively to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced visual feature of the chip sample surface state and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced visual feature of the chip sample internal structure.

[0008] In the above non-destructive analysis method for the chip packaging status based on microscopic detection, the same-paradigm feature enhancement is respectively performed on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure to obtain the enhanced visual feature encoding map of the chip sample surface state as the enhanced visual feature of the chip sample surface state and the enhanced visual feature encoding map of the chip sample internal structure as the enhanced visual feature of the chip sample internal structure, including: extracting the channel feature vector at the (h, w) pixel position from the visual feature encoding map of the chip sample surface state as the initial feature vector to be enhanced for the visual feature of the chip sample surface state; in the visual feature encoding map of the chip sample surface state, using the initial feature vector to be enhanced for the visual feature of the chip sample surface state as the center to perform feature sampling along multiple directions to obtain a set of first-direction context feature vectors for the visual feature of the chip sample surface state, a set of second-direction context feature vectors for the visual feature of the chip sample surface state, and a set of third-direction context feature vectors for the visual feature of the chip sample surface state; respectively inputting the set of first-direction context feature vectors for the visual feature of the chip sample surface state, the set of second-direction context feature vectors for the visual feature of the chip sample surface state, and the set of third-direction context feature vectors for the visual feature of the chip sample surface state into the direction context-aware perceptron based on the transformer structure to obtain the first-direction context-aware hidden encoding vector for the object to be enhanced for the visual feature of the chip sample surface state, the second-direction context-aware hidden encoding vector for the object to be enhanced for the visual feature of the chip sample surface state, and the third-direction context-aware hidden encoding vector for the object to be enhanced for the visual feature of the chip sample surface state; performing direction-sensitive attention analysis on the first-direction context-aware hidden encoding vector for the object to be enhanced for the visual feature of the chip sample surface state, the second-direction context-aware hidden encoding vector for the object to be enhanced for the visual feature of the chip sample surface state, and the third-direction context-aware hidden encoding vector for the object to be enhanced for the visual feature of the chip sample surface state to obtain the enhanced visual feature encoding vector of the chip sample surface state, where the enhanced visual feature encoding vector of the chip sample surface state is the channel feature vector at the (h, w) pixel position of the enhanced visual feature encoding map of the chip sample surface state.

[0009] In the above non-destructive analysis method for chip packaging status based on microscopic detection, in the visual feature encoding map of the chip sample surface status, taking the initial feature vector of the visual feature of the chip sample surface status to be enhanced as the center, feature sampling is performed along multiple directions to obtain a set of visual first-direction context feature vectors of the chip sample surface status, a set of visual second-direction context feature vectors of the chip sample surface status, and a set of visual third-direction context feature vectors of the chip sample surface status, including: in the visual feature encoding map of the chip sample surface status, taking the initial feature vector of the visual feature of the chip sample surface status to be enhanced as the center, feature sampling is performed along the first direction to obtain a set of visual first-direction context feature vectors of the chip sample surface status, and the initial feature vector of the visual feature of the chip sample surface status to be enhanced is located at the center position of the set of visual first-direction context feature vectors of the chip sample surface status; in the visual feature encoding map of the chip sample surface status, taking the initial feature vector of the visual feature of the chip sample surface status to be enhanced as the center, feature sampling is performed along the second direction to obtain a set of visual second-direction context feature vectors of the chip sample surface status, and the initial feature vector of the visual feature of the chip sample surface status to be enhanced is located at the center position of the set of visual second-direction context feature vectors of the chip sample surface status; in the visual feature encoding map of the chip sample surface status, taking the initial feature vector of the visual feature of the chip sample surface status to be enhanced as the center, feature sampling is performed along the third direction to obtain a set of visual third-direction context feature vectors of the chip sample surface status, and the initial feature vector of the visual feature of the chip sample surface status to be enhanced is located at the center position of the set of visual third-direction context feature vectors of the chip sample surface status.

[0010] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic inspection, a direction-sensitive attention analysis is performed on the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced to obtain an enhanced chip sample surface state visual feature coding vector, wherein the enhanced chip sample surface state visual feature coding vector is a channel feature vector of the (h, w)th pixel position of the enhanced chip sample surface state visual feature coding map, including: constructing a direction-sensitive attention module based on the first direction, the second direction and the third direction, and inputting the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0011] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic detection, a direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector, including: respectively inputting the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced The coding vector is anisotropically corrected to obtain a first-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, a second-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, and a third-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced; a direction-sensitive attention module is constructed based on the first direction, the second direction, and the third direction, and the first-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, the second-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, and the third-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0012] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic inspection, the packaging status detection result is determined based on the chip sample status multimodal fusion representation information between the enhanced chip sample surface status visual features and the enhanced chip sample internal structure visual features, and the packaging status detection result includes a predefined packaging status category label and its confidence, including: fusing the chip sample surface status visual feature coding map and the chip sample internal structure visual feature coding map to obtain a chip sample status multimodal fusion representation map as the chip sample status multimodal fusion representation information; inputting the chip sample status multimodal fusion representation map into a trained classifier to obtain a packaging status detection result, and the packaging status detection result includes a predefined packaging status category label and its confidence.

[0013] According to another aspect of the present application, a non-destructive analysis system for chip packaging conditions based on microscopic detection is provided, comprising: A target chip sample data acquisition module is used to acquire the original visible light image and original infrared image sequence of the target chip sample; A target chip sample data processing module, which is used to perform image preprocessing on the original visible light image and the original infrared image sequence respectively, and then perform spatial registration on them to obtain a registered chip sample visible light image and registered chip sample infrared volume data; A target chip sample data same-paradigm feature enhancement module, which is used to perform feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a multi-input deep learning model with two branches to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features; An encapsulation state detection result determination module, which is used to determine the encapsulation state detection result based on the chip sample state multi-modal fusion characterization information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, and the encapsulation state detection result includes predefined encapsulation state category labels and their confidence levels.

[0014] In the above non-destructive analysis system for chip packaging status based on microscopic detection, the target chip sample data same-paradigm feature enhancement module is used to: respectively input the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a chip sample surface state visual feature encoded map and a chip sample internal structure visual feature encoded map; respectively perform same-paradigm feature enhancement on the chip sample surface state visual feature encoded map and the chip sample internal structure visual feature encoded map to obtain an enhanced chip sample surface state visual feature encoded map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature encoded map as the enhanced chip sample internal structure visual feature.

[0015] Compared with the prior art, the non-destructive analysis system and method for chip packaging status provided by the present application first acquire the original visible light image and the original infrared image sequence and perform image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Then, key visual features representing the chip surface state and internal structure are extracted, refined and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically fused, and through a trained intelligent classifier, the packaging state of the chip is automatically identified and judged, and specific categories (such as normal, void, crack, etc.) and their reliability are given, so as to evaluate the packaging quality of the chip without damaging the chip. Description of the Drawings

[0016] The above - mentioned and other objects, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, but do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 It is a flowchart of a non - destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of the present application.

[0018] Figure 2 It is a schematic diagram of data flow of a non - destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of the present application.

[0019] Figure 3 It is a flowchart of feature extraction and same - paradigm feature enhancement for the visible - light image of the registered chip sample and the infrared volume data of the registered chip sample through a multi - input deep - learning model with double branches in a non - destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of the present application to obtain enhanced visual features of the chip sample surface state and enhanced visual features of the chip sample internal structure.

[0020] Figure 4 It is a flowchart of respectively performing same - paradigm feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure in a non - destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of the present application to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced visual features of the chip sample surface state and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced visual features of the chip sample internal structure.

[0021] Figure 5 It is a system block diagram of a non - destructive analysis system for chip packaging status based on microscopic detection according to an embodiment of the present application. Detailed implementation manners

[0022] The following will detail various exemplary embodiments, features, and aspects of the present application with reference to the accompanying drawings. The same reference numerals in the drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, unless otherwise specified, the drawings do not have to be drawn to scale.

[0023] As the core of modern electronic devices, the quality and reliability of integrated circuit (IC) chips are of utmost importance. Chip packaging is a crucial step in protecting internal circuits, providing electrical connections, and forming the final product. Any packaging defect may lead to performance degradation or even failure. Therefore, efficient and accurate detection of packaging quality is of great significance for ensuring product reliability, improving production yield, and reducing after-sales costs. Traditional detection methods such as manual visual inspection are inefficient and subjective, while destructive analysis is not applicable to batch detection and is time-consuming and labor-intensive. With the miniaturization of chip size and the complexity of packaging structures, there is an urgent need for a non-destructive, efficient, and multi-dimensional information acquisition detection technology. Existing methods based on visible light microscopes can only detect surface or shallow defects and are difficult to evaluate internal structures. Although non-destructive detection technologies such as ultrasonic and X-ray can detect internal problems, the equipment is specialized and the imaging principles are different, making it difficult to fuse and analyze with visible light images. How to effectively integrate multi-modal image data, mine complementary information, and achieve a comprehensive and accurate assessment of the packaging status remains the main challenge in current industrial detection.

[0024] To address the above technical problems, in the technical solution of this application, a non-destructive analysis method for chip packaging status based on microscopic detection is proposed, which uses multi-modal microscopic imaging technology to obtain comprehensive information of chip samples. Specifically, the system captures the surface details of the chip with visible light images, and at the same time uses infrared imaging sequences to penetrate the packaging material to obtain internal structure information. After these original optical detection data of different modalities are input into the system, they first go through a series of automatic processing steps for calibration and alignment to ensure that the information from different light sources and different depths can be spatially correlated. Subsequently, the system will intelligently extract the key visual features representing the chip surface state and internal structure from these aligned multi-modal data. To improve the discriminability and robustness of these features, the system refines and enhances the extracted features. Finally, the system organically fuses these processed and enhanced surface and internal features to form a unified and high-dimensional representation covering the overall packaging status of the chip. Based on this comprehensive representation integrating multi-modal information, the trained intelligent classifier can automatically identify and judge the packaging state of the chip, give the specific category (such as normal, void, crack, etc.) and its reliability, so as to achieve the evaluation of the packaging quality of the chip without damaging the chip.

[0025] This application proposes a non-destructive analysis method for chip packaging status based on microscopic detection. Figure 1 It is a flowchart of the non-destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of this application. Figure 2 It is a schematic diagram of data flow of the non-destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of this application. As Figure 1 and Figure 2As shown, the non-destructive analysis method for chip packaging status based on microscopic detection according to an embodiment of the present application includes: S110, obtaining an original visible light image and an original infrared image sequence of a target chip sample; S120, respectively performing image preprocessing on the original visible light image and the original infrared image sequence, and then performing spatial registration on them to obtain a registered chip sample visible light image and registered chip sample infrared volume data; S130, extracting features and enhancing features in the same paradigm from the registered chip sample visible light image and the registered chip sample infrared volume data through a multi-input deep learning model with two branches to obtain enhanced visual features of the chip sample surface state and enhanced visual features of the chip sample internal structure; S140, determining a packaging status detection result based on the chip sample state multi-modal fusion characterization information between the enhanced visual features of the chip sample surface state and the enhanced visual features of the chip sample internal structure, where the packaging status detection result includes predefined packaging status category labels and their confidence levels.

[0026] In the above non-destructive analysis method for chip packaging status based on microscopic inspection, in step S110, the original visible light image and the original infrared image sequence of the target chip sample are obtained. It should be understood that for the non-destructive inspection of the packaging status of integrated circuit (IC) chips, it is first necessary to obtain the original visible light image and the original infrared image sequence of the target chip sample. This step is the basis of the entire analysis process, and its purpose is to comprehensively capture the detailed information of the surface and internal structure of the chip packaging through multimodal imaging technology. Specifically, a high-resolution microscope system is used to photograph the chip sample. First, the visible light images are obtained, which can clearly reflect the detailed features on the surface of the chip packaging, such as surface defects like scratches, cracks, and contamination, as well as physical properties such as the color and texture of the packaging material. At the same time, in order to further understand the internal structure of the chip, especially those potential problems that may affect its performance and reliability, such as voids, delamination, and soldering defects, it is crucial to use infrared imaging technology. Infrared imaging can penetrate opaque packaging materials and reveal the internal thermal distribution or absorption characteristics, thereby helping to identify the above internal defects. When obtaining the original visible light image, an optical microscope equipped with a high-precision lens is usually used. This microscope can provide sufficient magnification to observe features as fine as a few micrometers on the surface of the chip packaging. During the operation, it is necessary to ensure that the light source evenly illuminates the sample to avoid misjudgment caused by the shadow effect. In addition, the working distance and focal length of the microscope need to be adjusted so that the image focus accurately falls on the surface of the chip packaging, thereby obtaining a clear and sharp image. During the acquisition process, multiple exposures or different filter combinations may be used to record various visual features on the chip surface under different lighting conditions and improve the accuracy of subsequent analysis. As for the acquisition of the original infrared image sequence, it depends on a specially designed infrared microscope or infrared imaging system. Such systems usually have high sensitivity and spatial resolution and can detect temperature changes or infrared radiation absorption differences inside the chip due to material property differences. In actual operation, first, it is necessary to select an appropriate infrared wavelength range according to the specific packaging type of the chip to be measured because different materials have different transmittances to infrared rays, and choosing the appropriate wavelength helps to optimize the imaging effect. Then, the chip is placed in a stable and controlled environmental condition to reduce the influence of external factors on the infrared imaging result. Subsequently, by gradually changing the focus depth of the infrared imaging system or moving the sample along a specific direction, a series of infrared images are continuously taken to form a three-dimensional data set containing multiple slices, that is, the so-called infrared volume data. This not only provides spatial information about the internal structure of the chip but also reflects to a certain extent the distribution of thermal or electrical properties at different depth levels.It should be noted that when acquiring these two types of images, the consistency of experimental conditions must be strictly controlled, including but not limited to environmental temperature, humidity, illumination intensity, and microscope parameter settings, etc., to ensure that the obtained data has good comparability and repeatability. At the same time, considering that there may be various complex background noise interferences in practical applications, appropriate preprocessing measures, such as background correction, noise reduction processing, etc., are also required to lay a solid foundation for subsequent steps such as image registration and feature extraction.

[0027] In the above non-destructive analysis method for chip packaging status based on microscopic detection, in step S120, after respectively performing image preprocessing on the original visible light image and the original infrared image sequence, spatial registration is performed on them to obtain the registered chip sample visible light image and the registered chip sample infrared volume data. It should be understood that although both visible light and infrared microscopic imaging systems can acquire chip information, they are based on different imaging principles and wavelengths, capturing different aspects of the chip: visible light mainly provides surface details, while infrared light can penetrate the packaging material to detect internal structures. The image data generated by these different modality data may have their own characteristics, and in the actual acquisition process, due to differences in optical path, focal plane position, field of view center, or scaling ratio, etc., the originally acquired visible light image and the infrared image sequence are often not precisely aligned in spatial position. Simple processing or shallow fusion is difficult to fully exploit complementary information. In order to overcome the spatial inconsistency of different modality data and prepare for subsequent analysis, in the technical solution of this application, after respectively performing image preprocessing on the original visible light image and the original infrared image sequence, spatial registration is performed on them to obtain the registered chip sample visible light image and the registered chip sample infrared volume data. That is to say, the infrared image sequence represents different depth or thermal state information inside the chip. Integrating it into volume data and aligning it with the planar visible light image through registration can ensure that subsequent multi-modal fusion is carried out under the same spatial reference system, avoiding information confusion caused by spatial misalignment, and laying a foundation for accurately extracting and fusing features of different modalities. In this way, it is ensured that information from different light sources and different depths can be spatially correlated, that is, an accurate correspondence relationship is established in space between the visible light image and the infrared volume data (generated or represented by the infrared sequence) representing the internal structure. By preprocessing (such as correcting distortion, noise reduction, enhancing contrast, etc.), the image quality of each modality can be optimized, and spatial registration maps these processed data into the same spatial coordinate system.

[0028] Figure 3The flowchart shows the method for non-destructively analyzing the chip packaging status based on microscopic detection according to the embodiments of the present application. In this method, a multi-input deep learning model with a dual-branch structure is used to extract features and perform homogeneous feature enhancement on the registered visible light image of the chip sample and the registered infrared volume data of the chip sample, so as to obtain the enhanced visual features of the chip sample surface state and the enhanced visual features of the chip sample internal structure. As Figure 3 shown, in the embodiment of the present application, step S130 includes: S131, respectively inputting the registered visible light image of the chip sample and the registered infrared volume data of the chip sample into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a visual feature encoding map of the chip sample surface state and a visual feature encoding map of the chip sample internal structure; S132, respectively performing homogeneous feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced visual features of the chip sample surface state and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced visual features of the chip sample internal structure.

[0029] Specifically, in step S131, the registered visible light image of the chip sample and the registered infrared volume data of the chip sample are respectively input into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a visual feature encoding map of the surface state of the chip sample and a visual feature encoding map of the internal structure of the chip sample. It should be understood that since the evaluation of the chip packaging condition requires obtaining and understanding multi-dimensional information from the surface to the inside, it is difficult for single-modal or traditional methods to fully capture and interpret these complex and interrelated features. Due to different imaging principles, the visible light image and the infrared volume data respectively carry key information on different aspects of the chip packaging state - visible light focuses on visual features such as surface defects, textures, and colors, while infrared can penetrate the packaging material to reveal defects related to thermal or absorption characteristics such as internal structures, voids, and delaminations. Simple manual observation or traditional image processing methods based on thresholds are difficult to effectively extract deep features that are truly discriminative for fault classification from high-dimensional and complex image data, especially when facing diverse defect types. Based on this, the registered visible light image of the chip sample and the registered infrared volume data of the chip sample are further respectively input into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a visual feature encoding map of the surface state of the chip sample and a visual feature encoding map of the internal structure of the chip sample. The purpose of performing this step is precisely to utilize the powerful non-linear feature learning ability of the deep learning model to automatically and intelligently extract high-level abstract features that respectively characterize the surface state and internal structure of the chip from the spatially aligned multi-modal raw data. Inputting the two registered modal data into a model with independent feature extraction branches respectively can ensure that each modal data can perform optimal feature learning through a dedicated network structure, avoiding interference from the differences in data attributes of different modalities during the feature extraction process, and at the same time providing clear-source and modality-independent feature descriptions for subsequent feature fusion. The visual feature encoding map of the surface state of the chip sample and the visual feature encoding map of the internal structure of the chip sample generated by this process are no longer the original pixel matrices, but compressed and high-semantic-density feature vectors or maps that accurately encode the key visual information on the surface and inside of the chip.

[0030] Figure 4 The flowchart shows the process of performing the same-paradigm feature enhancement on the visual feature encoding map of the surface state of the chip sample and the visual feature encoding map of the internal structure of the chip sample according to the embodiment of the present application to obtain the enhanced visual feature encoding map of the surface state of the chip sample as the enhanced visual feature of the surface state of the chip sample and the enhanced visual feature encoding map of the internal structure of the chip sample as the enhanced visual feature of the internal structure of the chip sample. As Figure 4As shown, in the embodiment of the present application, the step S132 includes: S1321, extracting the channel feature vector at the (h, w) pixel position from the visual feature encoding map of the chip sample surface state as the initial feature vector of the visual feature of the chip sample surface state to be enhanced; S1322, in the visual feature encoding map of the chip sample surface state, using the initial feature vector of the visual feature of the chip sample surface state to be enhanced as the center to perform feature sampling along multiple directions to obtain a set of first-direction context feature vectors of the visual feature of the chip sample surface state, a set of second-direction context feature vectors of the visual feature of the chip sample surface state, and a set of third-direction context feature vectors of the visual feature of the chip sample surface state; S1323, respectively inputting the set of first-direction context feature vectors of the visual feature of the chip sample surface state, the set of second-direction context feature vectors of the visual feature of the chip sample surface state, and the set of third-direction context feature vectors of the visual feature of the chip sample surface state into the direction context perceptron based on the transformer structure to obtain the first-direction context-aware implicit encoding vector of the visual feature of the chip sample surface state to be enhanced, the second-direction context-aware implicit encoding vector of the visual feature of the chip sample surface state to be enhanced, and the third-direction context-aware implicit encoding vector of the visual feature of the chip sample surface state to be enhanced; S1324, performing direction-sensitive attention analysis on the first-direction context-aware implicit encoding vector of the visual feature of the chip sample surface state to be enhanced, the second-direction context-aware implicit encoding vector of the visual feature of the chip sample surface state to be enhanced, and the third-direction context-aware implicit encoding vector of the visual feature of the chip sample surface state to be enhanced to obtain the enhanced visual feature encoding vector of the chip sample surface state, where the enhanced visual feature encoding vector of the chip sample surface state is the channel feature vector at the (h, w) pixel position of the enhanced visual feature encoding map of the chip sample surface state. It should be understood that although the multi-input deep learning model in the previous step has extracted the preliminary visual feature encoding map from the registered multi-modal images, the standard convolution operation (commonly used in the feature extraction layer of deep learning models) has inherent limitations, that is, its receptive field expands relatively slowly and usually homogenizes the local neighborhoods in all directions (isotropic). However, the defects or features contained in the surface state and internal structure of chip packaging often have obvious spatial directionality (such as cracks, solder pad arrays, wiring routes) and long-range dependencies (such as material delamination or structural deformation extending along a specific direction), which belong to anisotropic features. Standard convolution is difficult to efficiently and accurately capture these long-range dependencies or structural information extending along specific directions, thus affecting the sufficiency and discriminability of the extracted features.That is to say, in the standard convolutional neural network feature extraction process, although local patterns can be captured, its receptive field expands gradually and the processing of neighborhoods in all directions is relatively homogeneous (isotropic), making it difficult to efficiently capture long-range spatial dependencies or structural information that extends along specific directions. For example, to identify a fine but long surface crack or judge a large-area internal delamination defect, it is necessary to understand the continuity or change pattern of features along a specific direction in a wider range. If the key information with directionality and long-range dependencies in the feature map is not fully extracted and enhanced, subsequent analysis and classification may miss or misjudge these defects.

[0031] Based on this, in the technical solution of this application, the same-paradigm feature enhancement is respectively performed on the visual feature encoding map of the surface state of the chip sample and the visual feature encoding map of the internal structure of the chip sample to obtain the enhanced visual feature encoding map of the surface state of the chip sample and the enhanced visual feature encoding map of the internal structure of the chip sample. Through the processing of the same-paradigm feature enhancement, it is possible to explicitly model and integrate the context information collected along multiple predefined or adaptively learned spatial directions around each target feature vector in the visual feature encoding map of the surface state of the chip sample or the visual feature encoding map of the internal structure of the chip sample to refine and enhance the representation of the input feature map. Specifically, the way of this feature enhancement processing aims to sample and process the context information explicitly along multiple directions (such as horizontal, vertical, diagonal, etc.) and use powerful modeling capabilities (such as the Transformer structure mentioned in the corpus) to capture the long-range dependencies within the sequence, so as to more accurately understand the association between the feature at each pixel position and its surrounding (even far-distance) along a specific direction. In the chip detection scenario, this means that the network can better understand a linear crack, a planar delamination, or a chain of voids arranged along the solder joint array direction, and how these defects interact with their environmental features. By intelligently fusing and optimizing the context in different directions, an enhanced visual feature of the surface state of the chip sample containing richer and more discriminative multi-directional structured environmental knowledge is generated.

[0032] In an actual production line, the system needs to quickly and accurately identify various types of defects, such as distinguishing between simple scratches and through cracks, isolated bubbles and dangerous delaminations. Through the enhancement of visible light and infrared features in a homogeneous paradigm (i.e., using the same or similar network structures and enhancement logics) and based on multi-directional context awareness, the system has obtained a more powerful expression ability for the anisotropic features of the chip surface and internal structure. This means that the enhanced visual feature encoding map of the chip sample surface state can more clearly highlight and encode information such as directional cracks and abnormal packaging material textures; while the enhanced visual feature encoding map of the chip sample internal structure can more accurately capture and characterize defects distributed along specific directions or regions, such as internal voids, delaminations, and solder connection abnormalities. These enhanced feature maps with high quality and high semantic information provide more reliable inputs for subsequent multi-modal fusion and final intelligent classification, enabling the classifier to more confidently judge the packaging state of the chip (such as normal, cracked, voided, delaminated, etc.), reducing the false detection rate and missed detection rate. In a high-speed batch detection scenario, this enhanced feature expression ability means a more refined and robust assessment of the chip packaging health, directly improving the accuracy and reliability of the entire non-destructive analysis system, which is crucial for ensuring the quality of high-reliability chips.

[0033] Specifically, in step S1321, the channel feature vector at the (h, w) pixel position is extracted from the visual feature encoding map of the chip sample surface state as the initial feature vector to be enhanced for the visual features of the chip sample surface state, which is expressed by the channel feature extraction formula as:

[0034] Where is the visual feature encoding map of the chip sample surface state, the set of real numbers, are respectively the height, width, and number of channels of is in the channel feature vector of the channel, that is, the initial feature vector to be enhanced for the visual features of the chip sample surface state, is the Pixel position. It should be understood that due to chip packaging defects (such as cracks, delamination, solder joint anomalies, etc.), they often have significant directional characteristics (such as the linear extension direction of cracks, the periodic arrangement of solder joint arrays) and long-range dependencies (such as the extension of material delamination along a specific path). However, the standard convolution operation is limited by its fixed-size receptive field and isotropic neighborhood sampling pattern, making it difficult to effectively capture these direction-sensitive defect patterns. Especially in high-density packaging processes, micron-scale subtle defects may be distributed along specific geometric directions, and the homogeneous neighborhood aggregation of traditional convolution will lead to the weakening or even loss of directional characteristics. By extracting the channel feature vectors of each pixel point from the preprocessed multi-modal aligned images, it aims to construct a direction-sensitive feature enhancement basis. Each channel feature vector is essentially a compressed representation of the multi-spectral and multi-scale information within the local receptive field of the convolutional network at a specific spatial position (h, w), containing the semantic information of key defect criteria such as surface texture gradients and infrared thermal radiation distributions learned by the underlying network. By performing direction-sensitive context-aware enhancement on these high-dimensional feature vectors, it can break through the limitations of traditional convolution, explicitly model the anisotropic propagation law of defects in space, and provide a more discriminative feature representation for subsequent multi-modal fusion and intelligent classification. This operation essentially makes up for the representational ability gap of traditional convolution in directional defect detection through the directional enhancement of the feature space.

[0035] In the embodiment of the present application, the step S1322 includes: S1322-1, in the visual feature encoding map of the chip sample surface state, taking the initial feature vector to be enhanced of the visual feature of the chip sample surface state as the center, performing feature sampling along a first direction to obtain a set of visual first-direction context feature vectors of the chip sample surface state, and the initial feature vector to be enhanced of the visual feature of the chip sample surface state is located at the center position of the set of visual first-direction context feature vectors of the chip sample surface state; S1322-2, in the visual feature encoding map of the chip sample surface state, taking the initial feature vector to be enhanced of the visual feature of the chip sample surface state as the center, performing feature sampling along a second direction to obtain a set of visual second-direction context feature vectors of the chip sample surface state, and the initial feature vector to be enhanced of the visual feature of the chip sample surface state is located at the center position of the set of visual second-direction context feature vectors of the chip sample surface state; S1322-3, in the visual feature encoding map of the chip sample surface state, taking the initial feature vector to be enhanced of the visual feature of the chip sample surface state as the center, performing feature sampling along a third direction to obtain a set of visual third-direction context feature vectors of the chip sample surface state, and the initial feature vector to be enhanced of the visual feature of the chip sample surface state is located at the center position of the set of visual third-direction context feature vectors of the chip sample surface state.

[0036] Specifically, the step S1322 is represented by the feature sampling formula as follows:

[0037] Wherein, and are the step vectors in the directions respectively. For each direction , is the channel feature vector obtained by sampling adjacent positions in the direction (symmetric sampling, with the center being ). If the coordinates are out of bounds, zero values are used for supplementation. is the sampling neighborhood value, is the sampling boundary, and are the sets of the first - direction context feature vectors of the chip sample surface state vision, the second - direction context feature vectors of the chip sample surface state vision, and the third - direction context feature vectors of the chip sample surface state vision respectively. It should be understood that to address the deficiency of standard convolution in direction sensitivity and effectively represent the asymmetric features of spatial context, a multi - direction context feature acquisition mechanism is employed. Specifically, based on the multi - dimensional feature vector of the target pixel point in the pre - processed initial feature map, feature trajectory tracking is performed along a preset direction set (including typical geometric directions such as orthogonal axes and oblique directions). In this way, not only the aggregation of local neighborhood information is achieved, but also the construction of the directional feature sequence reveals the direction - related structural evolution law - the generated directional feature sequence contains the spatial topological features, the feature dynamic evolution map, and the long - range correlation information along a specific direction, such as the extension directionality of lattice defects or the periodic characteristics of metallization layer textures. This direction - constrained feature sampling strategy essentially realizes the precise decoupling of the anisotropic features of the material microstructure through directional trajectory modeling.

[0038] Specifically, in the step S1323, the set of the first - direction context feature vectors of the chip sample surface state vision, the set of the second - direction context feature vectors of the chip sample surface state vision, and the set of the third - direction context feature vectors of the chip sample surface state vision are respectively input into the direction context perceptron based on the transformer structure to obtain the first - direction context - aware implicit encoding vector of the chip sample surface state vision to - be - enhanced object, the second - direction context - aware implicit encoding vector of the chip sample surface state vision to - be - enhanced object, and the third - direction context - aware implicit encoding vector of the chip sample surface state vision to - be - enhanced object. It is represented by the direction context perception formula as follows:

[0039] Wherein, is a direction context aware device based on a transformer structure, are respectively a sequence of the first direction context aware hidden encoding vectors of the chip sample surface state vision, a sequence of the second direction context aware hidden encoding vectors of the chip sample surface state vision, and a sequence of the third direction context aware hidden encoding vectors of the chip sample surface state vision, is to extract the central position in (the th context aware hidden encoding vector) of the hidden encoding feature, are respectively the first direction context aware hidden encoding vector of the chip sample surface state vision object to be enhanced, the second direction context aware hidden encoding vector of the chip sample surface state vision object to be enhanced, and the third direction context aware hidden encoding vector of the chip sample surface state vision object to be enhanced. It should be understood that, to address the deficiency of standard convolution in direction sensitivity and effectively represent the asymmetric features of spatial context, a multi-direction context feature acquisition mechanism is employed. Specifically, based on the multi-dimensional feature vector of the target pixel point in the pre-processed initial feature map, feature trajectory tracking is performed along a preset direction set (including typical geometric directions such as orthogonal axes and oblique directions). In this way, not only is the aggregation of local neighborhood information achieved, but also the structural evolution law related to direction is revealed through the construction of the directional feature sequence - the generated directional feature sequence contains the spatial topological features, the feature dynamic evolution map, and the long-range correlation information along a specific direction, such as the extension directionality of lattice defects or the periodic characteristics of metallization layer textures. This direction-constrained feature sampling strategy essentially achieves the precise decoupling of the anisotropic features of the material microstructure through directional trajectory modeling.

[0040] In the embodiment of the present application, the step S1324 includes: constructing a direction-sensitive attention module based on the first direction, the second direction, and the third direction, and inputting the first direction context aware hidden encoding vector of the chip sample surface state vision object to be enhanced, the second direction context aware hidden encoding vector of the chip sample surface state vision object to be enhanced, and the third direction context aware hidden encoding vector of the chip sample surface state vision object to be enhanced into the direction-sensitive attention module to obtain the enhanced chip sample surface state vision feature encoding vector.

[0041] In an embodiment of the present application, the step S1324 includes: S1324-1, respectively performing anisotropy correction on the first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced to obtain the corrected first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the corrected second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the corrected third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced; S1324-2, constructing a direction-sensitive attention module based on the first direction, the second direction, and the third direction, and inputting the corrected first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the corrected second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the corrected third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0042] Specifically, the step S1324 is expressed as a direction-sensitive attention encoding formula:

[0043] in, for The transposed vector of and Direction and direction The embedding code vector of for The transposed vector of for and The correlation factor between The natural constant The exponential function value with base , For direction Relative to direction The attention score, and All represent the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, Not equal to , To enhance the visual feature encoding vector of the chip sample surface state, the enhanced visual feature encoding vector of the chip sample surface state is the channel feature vector at the pixel position of the enhanced visual feature encoding map of the chip sample surface state. It should be understood that for the fusion problem of the multi-direction context-aware implicit encoding vector, the network realizes intelligent integration by introducing a direction-sensitive attention mechanism. The core of this mechanism is to break the static weighting mode of traditional feature fusion and dynamically learn the contribution weights of different direction contexts in specific defect scenarios. Specifically, first, the context-aware implicit encoding vectors in three orthogonal directions (such as horizontal, vertical, diagonal) are input into a shared Transformer encoder, and the multi-head self-attention mechanism is used to capture the global dependencies of the feature sequences in each direction. On this basis, the direction-sensitive attention module further introduces a direction identity embedding, so that the attention weight calculation is not only based on the content similarity of the encoding vectors, but also explicitly associates with the direction attributes. For example, when processing periodic defects along the wafer scribe line, the module will dynamically enhance the weight of the horizontal direction context and suppress the interference features in the vertical direction; for randomly distributed microcracks, the attention weights show a multi-directional balanced distribution. Through this dynamic direction weighting mechanism, the system can adaptively screen out the context information most relevant to the current defect type and spatial position. The finally generated enhanced visual feature encoding vector of the chip sample surface state not only retains the texture specificity of local details but also enhances the cross-scale directional defect representation ability, providing a more discriminative input representation for the subsequent classifier.

[0044] In view of the fact that the direction context-aware of the Transformer architecture independently encodes the features in each direction, although it strengthens the local feature distribution, the attention mechanism exacerbates the direction-sensitive differences, resulting in feature coupling distortion due to the anisotropy effect during multi-direction feature fusion, thereby weakening the discrimination accuracy of the encoding vector. Therefore, this application introduces a high-order anisotropic polarization correction mechanism before feature aggregation to non-linearly reconstruct the weighted multi-direction feature vectors, compensates for the anisotropic perturbation through the direction-sensitive polarization response, eliminates the cross-direction feature coupling interference, and finally generates an enhanced encoding vector with high-fidelity direction features. This process effectively suppresses the feature distortion caused by the difference in the spatial field polarization rate by dynamically correcting the direction weight distribution, and significantly improves the robustness of defect representation.

[0045] Therefore, before summation, each weighted context-aware implicit encoding vector in each direction after weighting , performs a high-order anisotropic polarization representation:

[0046] Among them, is the Weighted context-aware implicit coding vector in the direction, is the weighted context-aware implicit coding vector in the direction, and is the context-aware high-order cross-directional polarization representation vector of.

[0047] Thus, the high-order field polarization rate analysis in the predetermined direction is carried out in the form of bilinear responses to the other two directions outside the predetermined direction.

[0048] Then, under the covariant partial derivative representation of cross-directional analysis, respectively, with respect to take the partial derivative:

[0049] where is the corrected context-aware high-order cross-directional polarization representation vector.

[0050] That is, since the context-aware implicit coding vectors of the other two directions are substantially symmetric in the denominator position, their partial derivatives are the same, which means that the overall field polarization response has anisotropic correction through partial derivative calculation.

[0051] In this way, by and performing dot product weighted correction to obtain then calculate:

[0052] In this way, by introducing a high-order cross-directional polarization correction mechanism, non-linearly reconstruct the weighted multi-directional context-aware implicit coding vector, use the direction-sensitive polarization response to compensate for anisotropic field distortion, effectively suppress cross-directional feature coupling interference, and finally generate an enhanced coding vector with high-fidelity direction features. This process significantly improves the robustness of defect characterization by dynamically correcting the direction weight distribution, enabling the enhanced chip sample surface state visual feature coding vector after fusion to strengthen the characterization ability of cross-scale directional defects while retaining the local detail texture specificity.

[0053] In the above non-destructive analysis method for chip packaging status based on microscopic detection, in step S140, based on the multi-modal fusion characterization information of the chip sample status between the enhanced visual features of the chip sample surface state and the enhanced visual features of the chip sample internal structure, the packaging status detection result is determined. The packaging status detection result includes pre-defined packaging status category labels and their confidence levels, and includes: S141, fusing the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure to obtain a multi-modal fusion characterization map of the chip sample status as the multi-modal fusion characterization information of the chip sample status; S142, inputting the multi-modal fusion characterization map of the chip sample status into a trained classifier to obtain the packaging status detection result, and the packaging status detection result includes pre-defined packaging status category labels and their confidence levels.

[0054] Specifically, in step S141, the visual feature encoding map of the surface state of the chip sample and the visual feature encoding map of the internal structure of the chip sample are fused to obtain a multi-modal fusion representation map of the chip sample state as the multi-modal fusion representation information of the chip sample state. It should be understood that the integrity and reliability of chip packaging need to be evaluated from multiple dimensions. A visible light microscope can mainly capture information such as the surface details, texture, color, etc. of chip packaging, and has strong detection capabilities for surface cracks, scratches, abnormal appearance of packaging materials, etc.; while infrared imaging, especially the volume data combined with sequence construction, can penetrate the packaging materials and reveal internal structure defects such as voids, delamination, abnormal internal connections, etc. The information obtained by these two modalities is naturally complementary, and no single modality can provide a comprehensive understanding of the overall packaging condition of the chip. Therefore, in order to overcome the limitations of a single modality and make full use of the advantages of multi-source information, the high-level features extracted from different modalities must be deeply integrated. In the technical solution of this application, the visual feature encoding map of the surface state of the chip sample and the visual feature encoding map of the internal structure of the chip sample are further fused to obtain a multi-modal fusion representation map of the chip sample state. This fusion is not simply data superposition, but is carried out in the high-semantic feature space learned by the deep learning model, aiming to capture the internal correlation and complex interaction patterns between surface features and internal structure features. For example, some internal defects may cause minor deformations or texture changes on the surface, while surface defects may indicate deeper structural problems. Through effective feature fusion, the system can comprehensively consider this cross-modal correlation information and generate a more discriminative representation than any single-modal feature, providing the richest and most effective information basis for subsequent defect classification. Specifically, considering that a visible light microscope mainly captures surface details of chip packaging such as cracks, scratches or abnormal appearance of materials, while infrared imaging can penetrate the packaging materials to reveal internal structure defects such as voids, delamination or soldering defects, the information provided by these two modalities is naturally complementary. Therefore, in the fusion process, a mechanism needs to be designed that can not only effectively combine these data from different dimensions, but also mine and utilize the potential correlations between them. For the fusion of the two encoding maps, a deep learning-based method may be adopted, especially those architectures with strong feature expression capabilities, such as convolutional neural networks (CNNs) or their variants. By using a pre-trained deep model, the two modalities of input data can be encoded to obtain high-dimensional abstract feature representations. On this basis, in order to further enhance the fusion effect, an attention mechanism may be introduced, which can dynamically adjust the importance weights of different features, making the final fusion result pay more attention to the most discriminative information part.For example, when building a multi-modal fusion model, the cross-modal attention module can be used to calculate the correlation scores of each local region in each modality with respect to the corresponding regions in other modalities, so as to guide the selection and combination of features. The purpose of doing this is to ensure that even in a complex background, the key feature points that reflect both surface characteristics and internal structures can be accurately identified, providing strong support for subsequent classification tasks. In addition, in addition to directly based on feature-level fusion, higher-level semantic fusion strategies can also be considered. This means not only considering pixel-level or feature vector-level data merging, but also exploring how to combine the high-level semantic information contained in different modalities. To this end, some technical means in the field of natural language processing, such as the Transformer architecture, can be used. By means of sequence modeling, the feature sequences from two encoded images are used as inputs and processed through multiple self-attention mechanisms. The advantage of this method is that it can automatically learn the optimal combination method of features between different modalities without manually defining specific fusion rules, thus improving the flexibility and generalization ability of the model. In the whole fusion process, another important consideration is how to solve the problem of inconsistent feature distributions between different modalities caused by differences in imaging principles. Since visible light images and infrared images respectively reflect different physical properties of the object surface and internal, there may be significant statistical differences between them. To solve this problem, the two encoded images can be normalized before fusion to make them have similar scales and distribution ranges, facilitating subsequent operations. At the same time, some advanced regularization techniques, such as adversarial training, can also be tried to let the model learn to minimize the differences between cross-modalities while maintaining their respective uniqueness, thus promoting more effective fusion.

[0055] Specifically, in step S142, the multi-modal fusion representation map of the chip sample state is input into the trained classifier to obtain the package state detection result, and the package state detection result includes predefined package state category labels and their confidence levels. It should be understood that the multi-modal fusion representation map of the chip sample state is input into the trained classifier to obtain the package state detection result, and the package state detection result includes predefined package state category labels and their confidence levels. In this way, the trained classifier can be used to automatically and accurately judge the package state of the chip sample. Specifically, it is to assign the high-dimensional multi-modal fusion representation map of the chip sample state to the most likely package state category according to the patterns learned by the model, and at the same time give the confidence level of the judgment.

[0056] In summary, the non-destructive analysis method for chip packaging status based on microscopic detection according to the embodiments of the present application is elucidated. First, the original visible light image and the original infrared image sequence are acquired and subjected to image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Then, the key visual features representing the surface state and internal structure of the chip are extracted, refined, and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically fused, and through a trained intelligent classifier, the packaging status of the chip is automatically identified and judged, and the specific category (such as normal, void, crack, etc.) and its reliability are given, so as to realize the evaluation of the packaging quality of the chip without damaging the chip.

[0057] Figure 5 FIG. is a system block diagram of a non-destructive analysis system for chip packaging status based on microscopic detection according to an embodiment of the present application. As Figure 5 shown, the non-destructive analysis system 100 for chip packaging status based on microscopic detection according to an embodiment of the present application includes: a target chip sample data acquisition module 110 for acquiring the original visible light image and the original infrared image sequence of the target chip sample; a target chip sample data processing module 120 for respectively performing image preprocessing on the original visible light image and the original infrared image sequence and then performing spatial registration on them to obtain a registered chip sample visible light image and a registered chip sample infrared volume data; a target chip sample data homomorphic feature enhancement module 130 for performing feature extraction and homomorphic feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features; a packaging status detection result determination module 140 for determining the packaging status detection result based on the chip sample state multi-modal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, where the packaging status detection result includes a predefined packaging status category label and its confidence level.

[0058] In the system of the above-mentioned non-destructive analysis system for chip packaging status based on microscopic detection, the target chip sample data homomorphic feature enhancement module is used to: input the registered visible light image of the chip sample and the registered infrared volume data of the chip sample into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch respectively to obtain a visual feature encoding map of the chip sample surface state and a visual feature encoding map of the chip sample internal structure; perform homomorphic feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure respectively to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced visual feature of the chip sample surface state and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced visual feature of the chip sample internal structure.

[0059] Here, those skilled in the art can understand that the specific operations of each step in the above-mentioned non-destructive analysis system for chip packaging status based on microscopic detection have been introduced in detail in the description of the non-destructive analysis method for chip packaging status based on microscopic detection above with reference to Figures 1 to 4 and thus, its repeated description will be omitted.

[0060] In summary, the non-destructive analysis system for chip packaging status based on the embodiments of the present application is elucidated. It first acquires the original visible light image and the original infrared image sequence and performs image preprocessing and spatial registration to ensure that the information from different light sources and different depths can be spatially correlated. Then, the key visual features representing the chip surface state and internal structure are extracted, refined, and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically fused, and through the trained intelligent classifier, it automatically identifies and judges the packaging state of the chip, gives the specific category (such as normal, void, crack, etc.) and its reliability, so as to realize the evaluation of the packaging quality of the chip without damaging the chip.

Claims

1. A non-destructive analysis method for the chip packaging condition based on microscopic detection, characterized in that Including: Obtaining an original visible light image and an original infrared image sequence of a target chip sample; After respectively performing image preprocessing on the original visible light image and the original infrared image sequence, performing spatial registration on them to obtain a registered chip sample visible light image and registered chip sample infrared volume data; Performing feature extraction and homogeneous paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a multi-input deep learning model with a dual-branch structure to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features; Based on the chip sample state multi-modal fusion characterization information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, determining a packaging state detection result, where the packaging state detection result includes a pre-defined packaging state category label and its confidence level.

2. The non-destructive analysis method for chip packaging status based on microscopic detection according to claim 1, wherein Performing feature extraction and homogeneous paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a multi-input deep learning model with a dual-branch structure to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features, including: Respectively inputting the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a visual feature encoding map of the chip sample surface state and a visual feature encoding map of the chip sample internal structure; Respectively performing homogeneous paradigm feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced chip sample surface state visual feature and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced chip sample internal structure visual feature.

3. The non-destructive analysis method for chip packaging status based on microscopic detection according to claim 2, wherein Respectively performing homogeneous paradigm feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced chip sample surface state visual feature and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced chip sample internal structure visual feature, including: Extracting a channel feature vector at the (h, w) pixel position from the visual feature encoding map of the chip sample surface state as an initial feature vector to be enhanced for the chip sample surface state visual feature; In the visual feature encoding map of the chip sample surface state, using the initial feature vector to be enhanced for the chip sample surface state visual feature as the center, performing feature sampling along multiple directions to obtain a set of first-direction context feature vectors for the chip sample surface state visual feature, a set of second-direction context feature vectors for the chip sample surface state visual feature, and a set of third-direction context feature vectors for the chip sample surface state visual feature; Input the set of visual first-direction context feature vectors of the chip sample surface state, the set of visual second-direction context feature vectors of the chip sample surface state, and the set of visual third-direction context feature vectors of the chip sample surface state into a direction context-aware processor based on a transformer structure to obtain a first-direction context-aware hidden encoding vector of the visual object to be enhanced for the chip sample surface state, a second-direction context-aware hidden encoding vector of the visual object to be enhanced for the chip sample surface state, and a third-direction context-aware hidden encoding vector of the visual object to be enhanced for the chip sample surface state; Perform direction-sensitivity attention analysis on the first-direction context-aware hidden encoding vector of the visual object to be enhanced for the chip sample surface state, the second-direction context-aware hidden encoding vector of the visual object to be enhanced for the chip sample surface state, and the third-direction context-aware hidden encoding vector of the visual object to be enhanced for the chip sample surface state to obtain an enhanced visual feature encoding vector for the chip sample surface state, where the enhanced visual feature encoding vector for the chip sample surface state is the channel feature vector at the (h, w) pixel position of the enhanced visual feature encoding map for the chip sample surface state.

4. The non-destructive analysis method for chip packaging status based on microscopic detection according to claim 3, characterized in that In the visual feature encoding map for the chip sample surface state, perform feature sampling along multiple directions with the initial feature vector to be enhanced for the visual feature of the chip sample surface state as the center to obtain a set of visual first-direction context feature vectors of the chip sample surface state, a set of visual second-direction context feature vectors of the chip sample surface state, and a set of visual third-direction context feature vectors of the chip sample surface state, including: In the visual feature encoding map for the chip sample surface state, perform feature sampling along the first direction with the initial feature vector to be enhanced for the visual feature of the chip sample surface state as the center to obtain a set of visual first-direction context feature vectors of the chip sample surface state, and the initial feature vector to be enhanced for the visual feature of the chip sample surface state is located at the center position of the set of visual first-direction context feature vectors of the chip sample surface state; In the visual feature encoding map for the chip sample surface state, perform feature sampling along the second direction with the initial feature vector to be enhanced for the visual feature of the chip sample surface state as the center to obtain a set of visual second-direction context feature vectors of the chip sample surface state, and the initial feature vector to be enhanced for the visual feature of the chip sample surface state is located at the center position of the set of visual second-direction context feature vectors of the chip sample surface state; In the visual feature encoding map for the chip sample surface state, perform feature sampling along the third direction with the initial feature vector to be enhanced for the visual feature of the chip sample surface state as the center to obtain a set of visual third-direction context feature vectors of the chip sample surface state, and the initial feature vector to be enhanced for the visual feature of the chip sample surface state is located at the center position of the set of visual third-direction context feature vectors of the chip sample surface state.

5. The non-destructive analysis method for chip packaging status based on microscopic detection according to claim 4, wherein, Direction-sensitive attention analysis is performed on the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced to obtain an enhanced chip sample surface state visual feature coding vector, wherein the enhanced chip sample surface state visual feature coding vector is a channel feature vector at the (h, w)th pixel position of the enhanced chip sample surface state visual feature coding map, including: A direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

6. The non-destructive analysis method for chip packaging status based on microscopic detection according to claim 5, wherein A direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector, including: Respectively perform heterogeneity correction on the first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced to obtain the corrected first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the corrected second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the corrected third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced; A direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

7. The non-destructive analysis method for chip packaging status based on microscopic detection according to claim 6, characterized in that, Based on the chip sample state multimodal fusion characterization information between the visual features of the enhanced chip sample surface state and the visual features of the enhanced chip sample internal structure, determine the packaging state detection result, where the packaging state detection result includes predefined packaging state category labels and their confidence levels, including: Fuse the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure to obtain a chip sample state multimodal fusion characterization map as the chip sample state multimodal fusion characterization information; Input the chip sample state multimodal fusion characterization map into a trained classifier to obtain the packaging state detection result, where the packaging state detection result includes predefined packaging state category labels and their confidence levels.

8. A non-destructive analysis system for chip packaging status based on microscopic detection, characterized in that, Including: A target chip sample data acquisition module, configured to acquire the original visible light image and the original infrared image sequence of the target chip sample; A target chip sample data processing module, configured to perform image preprocessing on the original visible light image and the original infrared image sequence respectively, and then perform spatial registration on them to obtain the registered chip sample visible light image and the registered chip sample infrared volume data; A target chip sample data same-paradigm feature enhancement module, configured to perform feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a multi-input deep learning model with two branches to obtain the enhanced visual features of the chip sample surface state and the enhanced visual features of the chip sample internal structure; A packaging state detection result determination module, configured to determine the packaging state detection result based on the chip sample state multimodal fusion characterization information between the enhanced visual features of the chip sample surface state and the enhanced visual features of the chip sample internal structure, where the packaging state detection result includes predefined packaging state category labels and their confidence levels.

9. The non-destructive analysis system for chip packaging status based on microscopic detection according to claim 8, wherein, The target chip sample data same-paradigm feature enhancement module is configured to: Input the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch respectively to obtain a visual feature encoding map of the chip sample surface state and a visual feature encoding map of the chip sample internal structure; Perform same-paradigm feature enhancement on the visual feature encoding map of the chip sample surface state and the visual feature encoding map of the chip sample internal structure respectively to obtain an enhanced visual feature encoding map of the chip sample surface state as the enhanced visual features of the chip sample surface state and an enhanced visual feature encoding map of the chip sample internal structure as the enhanced visual features of the chip sample internal structure.

Citation Information

Patent Citations

  • Substation equipment defect identification method and device

    CN112734692A

  • Infrared and visible light visual information fusion method based on gradient transformation prior

    CN117173063A

  • Infrared visible light visual fusion method based on feature mapping and self-attention mechanism

    CN117496314A

  • Equipment defect detection method

    CN118967689A

  • Wood water paint surface defect detection method based on computer vision

    CN119831958A