Chip packaging condition non-destructive analysis system and method based on microscopic inspection

By acquiring and fusing visible light and infrared image sequences, using multi-input deep learning models to extract and enhance the chip surface and internal features, the lossless and comprehensive multi-dimensional information acquisition problem in traditional detection methods is solved, and the automated and accurate evaluation of chip packaging status is achieved.

CN120259279BActive Publication Date: 2025-08-19SHANGHAI MAGIC PHOTOELECTRIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510688195.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-19
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing chip packaging detection methods are difficult to achieve lossless and comprehensive multi-dimensional information acquisition, especially efficient and accurate evaluation of the chip surface and internal structure. Traditional microscopy detection is limited to surface defects, other non-destructive detection methods are difficult to combine with visible light microscopes, and multi-modal image data is difficult to effectively integrate.

Method used

By acquiring visible light and infrared image sequences, image preprocessing and spatial registration, multi-input deep learning models are used to extract and enhance the visual features of the chip surface and internal structure, and organically fusion is performed, and package state evaluation is performed using a trained intelligent classifier.

Benefits of technology

It realizes automated and accurate evaluation of chip packaging quality under lossless conditions, and can identify packaging states such as normal, hollows, cracks, etc. and their reliability, improving the efficiency and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259279B_ABST
    Figure CN120259279B_ABST
Patent Text Reader

Abstract

The present application provides a non-destructive analysis system and method for chip packaging conditions based on microscopic inspection, which relates to the field of intelligent analysis. It first obtains the original visible light image and original infrared image sequence and performs image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Then, the key visual features representing the surface state and internal structure of the chip are extracted, refined and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically integrated, and a trained intelligent classifier is used to automatically identify and judge the packaging status of the chip, giving a specific category (such as normal, void, crack, etc.) and its reliability, thereby achieving the evaluation of its packaging quality without destroying the chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent analysis, and more particularly, in an embodiment of the present application, relates to a non-destructive analysis system and method for chip packaging conditions based on microscopic inspection. Background Art

[0002] Integrated circuit (IC) chips are at the heart of modern electronic devices, and their quality and reliability are paramount. Chip packaging is a critical step in the final product, protecting internal circuitry from environmental influences and providing electrical connections. Any defects or anomalies in the packaging process can lead to chip performance degradation or even failure. Therefore, effective inspection of chip packaging conditions is essential to ensure product quality, improve production yield, and reduce after-sales costs. Traditional chip packaging inspection methods often rely on manual visual inspection, destructive physical analysis (such as cross-sectioning), or specialized inspection equipment. Manual visual inspection is inefficient, subjective, and difficult to detect deep or subtle defects. While destructive analysis can provide detailed internal structural information, it cannot be used for comprehensive inspection of batches of products and is time-consuming and labor-intensive. With the miniaturization of chip sizes and the increasing complexity of packaging structures, the demand for efficient, non-destructive, accurate, and multi-dimensional inspection technologies is becoming increasingly urgent. In particular, for applications requiring high reliability, the ability to rapidly and comprehensively assess chip packaging health, including potential issues with surface integrity and internal structure, has become a pressing technical challenge.

[0003] Existing microscopic inspection-based chip packaging condition analysis technologies, if they rely solely on visible light microscopy, are typically limited to the surface or shallow structures of the chip package. They are unable to detect defects beneath the packaging material or within internal structures such as solder joints and wiring, making true nondestructive internal inspection impossible. While some nondestructive testing methods, such as ultrasonic microscopy or X-ray testing, can detect internal structures, they often require specific equipment and imaging principles and may be difficult to combine with high-resolution visible light microscopy to provide comprehensive information. Furthermore, the image data generated by these methods may each have unique characteristics. Effectively integrating data from different modalities to achieve a comprehensive assessment of chip packaging conditions is a major challenge facing existing technologies. Simply analyzing images from different modalities independently or performing shallow fusion makes it difficult to fully exploit the complementary information hidden in multi-source data. This is especially true when faced with complex and diverse packaging defects. The robustness and accuracy of traditional image processing or single-modality analysis methods are often insufficient to meet the needs of industrial-grade inspection.

[0004] Therefore, an optimized non-destructive analysis system for chip packaging conditions based on microscopic inspection is desired. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a non-destructive analysis system and method for chip packaging status based on microscopic detection, which first obtains the original visible light image and the original infrared image sequence and performs image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Then, the key visual features representing the surface state and internal structure of the chip are extracted, and refined and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically integrated, and a trained intelligent classifier is used to automatically identify and judge the packaging status of the chip, and give specific categories (such as normal, voids, cracks, etc.) and their reliability, so as to achieve the evaluation of its packaging quality without destroying the chip.

[0006] According to one aspect of the present application, a non-destructive analysis method for chip packaging conditions based on microscopic inspection is provided, which includes:

[0007] Acquire the original visible light image and original infrared image sequence of the target chip sample;

[0008] After performing image preprocessing on the original visible light image and the original infrared image sequence, spatial registration is performed on them to obtain registered chip sample visible light images and registered chip sample infrared volume data;

[0009] Performing feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features;

[0010] Based on the chip sample state multimodal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, a packaging state detection result is determined, and the packaging state detection result includes a predefined packaging state category label and its confidence.

[0011] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic inspection, feature extraction and same-paradigm feature enhancement are performed on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features, including: inputting the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch respectively to obtain a chip sample surface state visual feature coding map and a chip sample internal structure visual feature coding map; performing same-paradigm feature enhancement on the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map respectively to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature.

[0012] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic inspection, the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map are respectively subjected to same-paradigm feature enhancement to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature, including: extracting the channel feature vector at the (h, w)th pixel position from the chip sample surface state visual feature coding map as the initial feature vector of the chip sample surface state visual feature to be enhanced; in the chip sample surface state visual feature coding map, taking the initial feature vector of the chip sample surface state visual feature to be enhanced as the center, performing feature sampling along multiple directions to obtain a set of chip sample surface state visual first direction context feature vectors, a set of chip sample surface state visual second direction context feature vectors and a set of chip sample surface state visual third direction context feature vectors; and taking the chip sample surface state visual first direction context feature vector as the initial feature vector of the chip sample surface state visual feature to be enhanced. The set of feature vectors, the set of the chip sample surface state visual second direction context feature vectors and the set of the chip sample surface state visual third direction context feature vectors are respectively input into the direction context perception based on the converter structure to obtain the chip sample surface state visual first direction context perception implicit coding vector of the object to be enhanced, the chip sample surface state visual second direction context perception implicit coding vector of the object to be enhanced and the chip sample surface state visual third direction context perception implicit coding vector of the object to be enhanced; the direction-sensitive attention analysis is performed on the chip sample surface state visual first direction context perception implicit coding vector of the object to be enhanced, the chip sample surface state visual second direction context perception implicit coding vector of the object to be enhanced and the chip sample surface state visual third direction context perception implicit coding vector of the object to be enhanced to obtain the enhanced chip sample surface state visual feature coding vector, wherein the enhanced chip sample surface state visual feature coding vector is the channel feature vector of the (h, w)th pixel position of the enhanced chip sample surface state visual feature coding map.

[0013] In the above-mentioned non-destructive analysis method of chip packaging condition based on microscopic inspection, in the chip sample surface state visual feature coding map, feature sampling is performed along multiple directions with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual first direction context feature vectors, a set of chip sample surface state visual second direction context feature vectors and a set of chip sample surface state visual third direction context feature vectors, including: in the chip sample surface state visual feature coding map, feature sampling is performed along the first direction with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual first direction context feature vectors, the chip sample surface state visual feature initial feature vector to be enhanced is located in the chip sample surface state visual first direction context feature direction the center position of a set of quantities; in the chip sample surface state visual feature coding map, feature sampling is performed along the second direction with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual second direction context feature vectors, and the initial feature vector of the chip sample surface state visual feature to be enhanced is located at the center position of the set of chip sample surface state visual second direction context feature vectors; in the chip sample surface state visual feature coding map, feature sampling is performed along the third direction with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual third direction context feature vectors, and the initial feature vector of the chip sample surface state visual feature to be enhanced is located at the center position of the set of chip sample surface state visual third direction context feature vectors.

[0014] In the above-mentioned non-destructive analysis method of chip packaging condition based on microscopic inspection, direction-sensitive attention analysis is performed on the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced to obtain an enhanced chip sample surface state visual feature coding vector, wherein the enhanced chip sample surface state visual feature coding vector is a channel feature vector of the (h, w)th pixel position of the enhanced chip sample surface state visual feature coding map, including: constructing a direction-sensitive attention module based on the first direction, the second direction, and the third direction, and inputting the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0015] In the above-mentioned non-destructive analysis method of chip packaging condition based on microscopic inspection, a direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector, including: respectively inputting the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced into the direction-sensitive attention module. The coding vector is anisotropically corrected to obtain a first-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, a second-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, and a third-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced; a direction-sensitive attention module is constructed based on the first direction, the second direction, and the third direction, and the first-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, the second-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, and the third-direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0016] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic inspection, the packaging status detection result is determined based on the chip sample status multimodal fusion representation information between the enhanced chip sample surface status visual features and the enhanced chip sample internal structure visual features, and the packaging status detection result includes a predefined packaging status category label and its confidence level, including: fusing the chip sample surface status visual feature coding map and the chip sample internal structure visual feature coding map to obtain a chip sample status multimodal fusion representation map as the chip sample status multimodal fusion representation information; inputting the chip sample status multimodal fusion representation map into a trained classifier to obtain a packaging status detection result, and the packaging status detection result includes a predefined packaging status category label and its confidence level.

[0017] According to another aspect of the present application, a microscopic inspection-based non-destructive analysis system for chip packaging conditions is provided, comprising:

[0018] A target chip sample data acquisition module is used to acquire the original visible light image and original infrared image sequence of the target chip sample;

[0019] a target chip sample data processing module, configured to perform image preprocessing on the original visible light image and the original infrared image sequence, and then spatially register them to obtain registered chip sample visible light images and registered chip sample infrared volume data;

[0020] A target chip sample data same-paradigm feature enhancement module is used to perform feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features;

[0021] The packaging state detection result determination module is used to determine the packaging state detection result based on the chip sample state multimodal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features. The packaging state detection result includes a predefined packaging state category label and its confidence.

[0022] In the above-mentioned non-destructive analysis system of chip packaging status based on microscopic inspection, the target chip sample data same-paradigm feature enhancement module is used to: input the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a chip sample surface state visual feature coding map and a chip sample internal structure visual feature coding map; perform same-paradigm feature enhancement on the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map respectively to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature.

[0023] Compared with the existing technology, the present application provides a non-destructive analysis system and method for chip packaging conditions based on microscopic inspection. It first obtains the original visible light image and original infrared image sequence and performs image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Then, the key visual features representing the surface state and internal structure of the chip are extracted, and refined and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically integrated, and a trained intelligent classifier is used to automatically identify and judge the packaging status of the chip, giving specific categories (such as normal, voids, cracks, etc.) and their reliability, thereby achieving the evaluation of its packaging quality without destroying the chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0025] Figure 1 Flowchart of a non-destructive analysis method of chip packaging status based on microscopic inspection according to an embodiment of the present application.

[0026] Figure 2 Schematic diagram of data flow of a non-destructive analysis method of chip packaging status based on microscopic inspection according to an embodiment of the present application.

[0027] Figure 3 This is a flowchart of a non-destructive analysis method of chip packaging status based on microscopic inspection according to an embodiment of the present application, which uses a dual-branch multi-input deep learning model to perform feature extraction and same-paradigm feature enhancement on the visible light image of the registered chip sample and the infrared volume data of the registered chip sample to obtain enhanced visual features of the chip sample surface state and enhanced visual features of the chip sample internal structure.

[0028] Figure 4 This is a flowchart of performing same-paradigm feature enhancement on the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map in the non-destructive analysis method of chip packaging status based on microscopic inspection according to an embodiment of the present application to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature.

[0029] Figure 5 4 is a system block diagram of a chip packaging condition non-destructive analysis system based on microscopic inspection according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0031] Integrated circuit (IC) chips are the core of modern electronic devices, and their quality and reliability are paramount. Chip packaging is a critical step in protecting internal circuitry, providing electrical connections, and forming the final product. Any packaging defect can lead to performance degradation or even failure. Therefore, efficient and accurate packaging quality inspection is crucial for ensuring product reliability, improving production yield, and reducing after-sales costs. Traditional inspection methods, such as manual visual inspection, are inefficient and highly subjective, while destructive analysis is unsuitable for mass production and is time-consuming and labor-intensive. With the miniaturization of chip sizes and the increasing complexity of packaging structures, there is an urgent need for non-destructive, efficient, and multi-dimensional inspection technologies. Existing methods based on visible light microscopy can only detect surface or shallow defects, making it difficult to assess internal structures. While non-destructive testing techniques such as ultrasound and X-rays can detect internal problems, their specialized equipment and different imaging principles make them difficult to fuse with visible light images for analysis. Effectively integrating multimodal image data, mining complementary information, and achieving a comprehensive and accurate assessment of packaging conditions remains a major challenge facing industrial inspection.

[0032] In response to the above technical problems, in the technical solution of the present application, a non-destructive analysis method of chip packaging conditions based on microscopic inspection is proposed, which uses multimodal microscopic imaging technology to obtain comprehensive information of chip samples. Specifically, the system captures the surface details of the chip with visible light images, and uses infrared imaging sequences to penetrate the packaging material to obtain internal structure information. After these original, different modal optical detection data are input into the system, they are first calibrated and aligned through a series of automatic processing steps to ensure that information from different light sources and different depths can be spatially correlated. Subsequently, the system intelligently extracts key visual features representing the surface state and internal structure of the chip from these aligned multimodal data. In order to improve the discrimination and robustness of these features, the system refines and enhances the extracted features. Finally, the system organically integrates these processed and enhanced surface and internal features to form a unified, high-dimensional representation covering the overall packaging condition of the chip. Based on this comprehensive representation that integrates multimodal information, the trained intelligent classifier can automatically identify and judge the packaging status of the chip, give specific categories (such as normal, voids, cracks, etc.) and their reliability, thereby realizing the evaluation of its packaging quality without destroying the chip.

[0033] This application proposes a non-destructive analysis method for chip packaging conditions based on microscopic inspection. Figure 1 Flowchart of a non-destructive analysis method of chip packaging status based on microscopic inspection according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the chip packaging condition non-destructive analysis method based on microscopic inspection according to an embodiment of the present application. Figure 1 and Figure 2As shown, the non-destructive analysis method of chip packaging status based on microscopic inspection according to the embodiment of the present application includes: S110, obtaining the original visible light image and the original infrared image sequence of the target chip sample; S120, performing image preprocessing on the original visible light image and the original infrared image sequence respectively, and then performing spatial registration on them to obtain the registered chip sample visible light image and the registered chip sample infrared volume data; S130, performing feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features; S140, determining the packaging state detection result based on the chip sample state multimodal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, and the packaging state detection result includes a predefined packaging state category label and its confidence.

[0034] In the aforementioned microscopically inspected nondestructive chip packaging condition analysis method, step S110 involves acquiring a sequence of raw visible light images and raw infrared images of a target chip sample. It should be understood that nondestructive inspection of integrated circuit (IC) chip packaging conditions first requires acquiring a sequence of raw visible light images and raw infrared images of the target chip sample. This step forms the foundation of the entire analysis process, aiming to comprehensively capture detailed information about the chip package surface and internal structure using multimodal imaging technology. Specifically, a high-resolution microscope system is used to image the chip sample, initially acquiring visible light images. These images clearly reveal detailed features of the chip package surface, such as surface defects like scratches, cracks, and contamination, as well as physical properties of the packaging material, such as color and texture. Furthermore, infrared imaging technology is crucial for further understanding the chip's internal structure, particularly potential issues that may affect its performance and reliability, such as voids, delamination, and solder defects. Infrared imaging can penetrate opaque packaging materials, revealing their internal thermal distribution and absorption characteristics, thereby assisting in identifying these internal defects. Acquiring raw visible light images typically involves using an optical microscope equipped with a high-precision lens. This microscope provides sufficient magnification to observe features as fine as a few microns on the chip package surface. During operation, the light source must be uniformly illuminated to avoid distorted image quality caused by shadowing. Furthermore, the microscope's working distance and focal length must be adjusted to ensure the image is accurately focused on the chip package surface, resulting in a clear, sharp image. Multiple exposures or different filter combinations may be used during acquisition to capture various visual features on the chip surface under varying lighting conditions, improving the accuracy of subsequent analysis. Acquiring raw infrared image sequences relies on a specially designed infrared microscope or infrared imaging system. These systems typically offer high sensitivity and spatial resolution, enabling detection of temperature changes or differences in infrared radiation absorption within the chip due to variations in material properties. In practice, the appropriate infrared wavelength range must be selected based on the specific package type of the chip being tested. Because different materials have varying infrared transmittances, choosing the appropriate wavelength helps optimize imaging. Next, the chip must be placed in a stable and controlled environment to minimize the impact of external factors on the infrared imaging results. Subsequently, by gradually changing the focus depth of the infrared imaging system or moving the sample in a specific direction, a series of infrared images are continuously captured, forming a three-dimensional dataset containing multiple slices, known as infrared volume data. This not only provides spatial information about the internal structure of the chip, but also, to a certain extent, reflects the distribution of thermal or electrical properties at different depth levels.It is worth noting that when acquiring these two types of images, the consistency of experimental conditions must be strictly controlled, including but not limited to ambient temperature, humidity, lighting intensity, and microscope parameter settings, to ensure that the obtained data are highly comparable and reproducible. At the same time, considering the various complex background noise interferences that may exist in practical applications, appropriate preprocessing measures such as background correction and noise reduction are also required to lay a solid foundation for subsequent steps such as image registration and feature extraction.

[0035] In the aforementioned microscopic inspection-based nondestructive analysis method for chip packaging conditions, step S120 preprocesses the original visible light image and the original infrared image sequence, and then spatially registers them to obtain registered visible light images and registered infrared volume data of the chip sample. It should be understood that while both visible light and infrared microscopic imaging systems can acquire chip information, they are based on different imaging principles and wavelengths, capturing different aspects of the chip: visible light primarily provides surface details, while infrared light can penetrate packaging materials to detect internal structures. The image data generated by these different modalities may each have unique characteristics, and during the actual acquisition process, due to differences in optical paths, focal plane positions, field of view centers, or scaling, the original visible light images and infrared image sequences are often not precisely aligned in terms of spatial position. Simple processing or shallow fusion makes it difficult to fully exploit complementary information. To overcome spatial inconsistencies in data from different modalities and prepare for subsequent analysis, the technical solution of this application performs image preprocessing on the original visible light image and the original infrared image sequence, followed by spatial registration to produce registered chip sample visible light images and registered chip sample infrared volume data. Specifically, the infrared image sequence represents information about different depths or thermal states within the chip. Integrating this into volume data and aligning it with the planar visible light image through registration ensures that subsequent multimodal fusion is performed within the same spatial reference system, avoiding information confusion caused by spatial misalignment and laying the foundation for accurate feature extraction and fusion of different modalities. This ensures that information from different light sources and depths can be spatially correlated, establishing an accurate spatial correspondence between the visible light image and the infrared volume data representing internal structures (generated or represented by the infrared sequence). Preprocessing (such as distortion correction, noise reduction, and contrast enhancement) optimizes the image quality of each modality, while spatial registration maps these processed data to the same spatial coordinate system.

[0036] Figure 3This is a flowchart of a method for non-destructive analysis of chip packaging conditions based on microscopic inspection according to an embodiment of the present application, in which a dual-branch multi-input deep learning model is used to extract features from the registered chip sample visible light image and the registered chip sample infrared volume data, and enhance the same-paradigm features to obtain enhanced visual features of the chip sample surface state and enhanced visual features of the chip sample internal structure. Figure 3 As shown, in an embodiment of the present application, the step S130 includes: S131, respectively inputting the chip sample visible light image after registration and the chip sample infrared volume data after registration into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a chip sample surface state visual feature coding map and a chip sample internal structure visual feature coding map; S132, respectively performing same-paradigm feature enhancement on the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature.

[0037] Specifically, step S131 involves inputting the registered visible light image and infrared volume data of the chip sample into a multi-input deep learning model comprising a first feature extraction branch and a second feature extraction branch, respectively, to obtain a visual feature encoding map of the chip sample's surface condition and a visual feature encoding map of the chip sample's internal structure. It should be understood that assessing chip packaging conditions requires acquiring and understanding multi-dimensional information from the surface to the interior, and single-modality or traditional methods struggle to fully capture and interpret these complex and interrelated features. Due to their different imaging principles, visible light images and infrared volume data each carry key information on different aspects of chip packaging conditions. Visible light focuses on visual features such as surface defects, texture, and color, while infrared can penetrate packaging materials to reveal defects related to thermal or absorption properties, such as internal structure, voids, and delamination. Simple manual observation or traditional threshold-based image processing methods struggle to effectively extract truly discriminative deep features from high-dimensional, complex image data, especially when faced with diverse defect types. Based on this, the registered chip sample visible light image and the registered chip sample infrared volume data are further input into a multi-input deep learning model comprising a first feature extraction branch and a second feature extraction branch, respectively, to generate a visual feature encoding map of the chip sample's surface state and a visual feature encoding map of the chip sample's internal structure. This step aims to leverage the powerful nonlinear feature learning capabilities of the deep learning model to automatically and intelligently extract high-level, abstract features representing the chip's surface state and internal structure, respectively, from the spatially aligned multimodal raw data. Inputting the registered two modal data into a model with independent feature extraction branches ensures optimal feature learning for each modality through a dedicated network architecture, preventing interference from differences in modal data attributes on the feature extraction process. This also provides a clearly sourced, modality-independent feature description for subsequent feature fusion. The resulting visual feature encoding map of the chip sample's surface state and internal structure is no longer a raw pixel matrix, but rather a compressed, highly semantically dense feature vector or map that accurately encodes key visual information about the chip's surface and internal structure.

[0038] Figure 4 This is a flowchart of performing the same-paradigm feature enhancement on the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map in the non-destructive analysis method of chip packaging status based on microscopic inspection according to an embodiment of the present application to obtain the enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and the enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature. Figure 4As shown, in an embodiment of the present application, the step S132 includes: S1321, extracting the channel feature vector of the (h, w)th pixel position from the chip sample surface state visual feature coding map as the initial feature vector of the chip sample surface state visual feature to be enhanced; S1322, in the chip sample surface state visual feature coding map, taking the initial feature vector of the chip sample surface state visual feature to be enhanced as the center, performing feature sampling along multiple directions to obtain a set of chip sample surface state visual first direction context feature vectors, a set of chip sample surface state visual second direction context feature vectors and a set of chip sample surface state visual third direction context feature vectors; S1323, combining the set of chip sample surface state visual first direction context feature vectors, the set of chip sample surface state visual second direction context feature vectors and the set of chip sample surface state visual third direction context feature vectors. The sets of quantities are respectively input into the direction context perception based on the converter structure to obtain the first direction context perception implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context perception implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context perception implicit coding vector of the chip sample surface state visual object to be enhanced; S1324, the first direction context perception implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context perception implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context perception implicit coding vector of the chip sample surface state visual object to be enhanced are subjected to direction-sensitive attention analysis to obtain an enhanced chip sample surface state visual feature coding vector, wherein the enhanced chip sample surface state visual feature coding vector is the channel feature vector of the (h, w)th pixel position of the enhanced chip sample surface state visual feature coding map. It's understandable that while the multi-input deep learning model in the previous step has extracted a preliminary visual feature encoding map from the registered multimodal images, standard convolution operations (commonly used in the feature extraction layer of deep learning models) have inherent limitations: their receptive field expands relatively slowly and they typically homogenize local neighborhoods in all directions (isotropic processing). However, defects or features contained in the surface state and internal structure of chip packages often exhibit significant spatial directionality (such as cracks, solder joint arrays, and wiring routing) and long-range dependencies (such as material delamination or structural deformation extending along specific directions), which are considered anisotropic features. Standard convolutions struggle to efficiently and accurately capture these long-range dependencies or structural information extending along specific directions, thus affecting the adequacy and discriminative power of the extracted features.In other words, while the standard convolutional neural network feature extraction process can capture local patterns, its receptive field gradually expands and its treatment of neighborhoods in all directions is relatively homogeneous (isotropic). This makes it difficult to effectively capture long-range spatial dependencies or structural information extending along specific directions. For example, identifying a subtle but long surface crack or determining a large internal delamination defect requires understanding the continuity or variation of features along a specific direction over a wider range. If these key directional and long-range dependencies are not fully extracted and enhanced in the feature map, subsequent analysis and classification may miss or misidentify these defects.

[0039] Based on this, in the technical solution of the present application, the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map are respectively subjected to same-paradigm feature enhancement to obtain an enhanced chip sample surface state visual feature coding map and an enhanced chip sample internal structure visual feature coding map. Through the same-paradigm feature enhancement process, it is possible to explicitly model and integrate each target feature vector in the chip sample surface state visual feature coding map or the chip sample internal structure visual feature coding map, and to refine and enhance the representation of the input feature map by collecting contextual information along multiple predefined or adaptively learned spatial directions. Specifically, this feature enhancement processing method aims to capture long-distance dependencies within the sequence by explicitly sampling and processing contextual information along multiple directions (such as horizontal, vertical, diagonal, etc.), and utilizing powerful modeling capabilities (such as the Transformer structure mentioned in the corpus), so as to more accurately understand the association between the features of each pixel position and its surroundings (even at a long distance) along a specific direction. In the context of chip inspection, this means the network can better understand a linear crack, a planar delamination, or a chain of voids along an array of solder joints, as well as how these defects interact with their surrounding features. By intelligently fusing and optimizing contexts from different directions, it generates enhanced visual features of the chip sample surface state that contain richer and more discriminative multi-directional structured environmental knowledge.

[0040] On actual production lines, systems must quickly and accurately identify various types of defects, such as simple scratches and through-the-lens cracks, isolated bubbles and dangerous delaminations. By enhancing visible and infrared features using the same paradigm (i.e., employing the same or similar network structure and enhancement logic) and multi-directional contextual awareness, the system achieves a more robust representation of the anisotropic characteristics of the chip surface and internal structure. This means that the enhanced visual feature encoding map of the chip sample's surface state more clearly highlights and encodes information such as directional cracks and abnormal packaging material textures. Furthermore, the enhanced visual feature encoding map of the chip sample's internal structure more accurately captures and characterizes defects distributed along specific directions or regions, such as internal voids, delamination, and solder joint anomalies. These high-quality, semantically rich enhanced feature maps provide more reliable input for subsequent multimodal fusion and ultimately intelligent classification, enabling the classifier to more confidently determine the chip's packaging state (e.g., normal, cracked, voided, delamination, etc.), reducing false positives and missed detections. In high-speed batch inspection scenarios, this enhanced feature expression capability means that the assessment of chip package health is more refined and robust, which directly improves the accuracy and reliability of the entire non-destructive analysis system and is crucial for ensuring the quality of high-reliability chips.

[0041] Specifically, in step S1321, the channel feature vector at the (h, w)th pixel position is extracted from the chip sample surface state visual feature coding map as the initial feature vector to be enhanced for the chip sample surface state visual feature, which is expressed as:

[0042] in, is a visual feature coding diagram of the chip sample surface state, the set of real numbers, They are The height, width and number of channels, for middle The channel feature vector of the channel, that is, the initial feature vector of the chip sample surface state visual feature to be enhanced, for No. Pixel location. It's understandable that chip packaging defects (such as cracks, delamination, and solder joint anomalies) often exhibit significant directional characteristics (e.g., the linear extension direction of cracks or the periodic arrangement of solder joint arrays) and long-range dependencies (e.g., material delamination propagating along a specific path). Standard convolution operations, limited by their fixed-size receptive field and isotropic neighborhood sampling pattern, struggle to effectively capture these directionally sensitive defect patterns. In high-density packaging processes, subtle micron-scale defects may be distributed along specific geometric directions. The homogeneous neighborhood aggregation of traditional convolutions can weaken or even lose these directional characteristics. By extracting a channel feature vector for each pixel from the preprocessed multimodal aligned image, we aim to build a foundation for directionally sensitive feature enhancement. Each channel feature vector is essentially a compressed representation of the multispectral, multiscale information within the local receptive field of the convolutional network at a specific spatial location (h, w). It incorporates semantic information learned by the underlying network about key defect criteria, such as surface texture gradients and infrared thermal radiation distribution. By performing directionally sensitive, context-aware enhancement on these high-dimensional feature vectors, we can overcome the limitations of traditional convolution and explicitly model the anisotropic propagation of defects in space, providing more discriminative feature representations for subsequent multimodal fusion and intelligent classification. This operation, essentially through directional enhancement of feature space, fills the gap in the representational capabilities of traditional convolution in directional defect detection.

[0043] In an embodiment of the present application, the step S1322 includes: S1322-1, in the chip sample surface state visual feature coding map, taking the chip sample surface state visual feature initial feature vector to be enhanced as the center along the first direction to perform feature sampling to obtain a set of chip sample surface state visual first direction context feature vectors, and the chip sample surface state visual feature initial feature vector to be enhanced is located at the center position of the set of chip sample surface state visual first direction context feature vectors; S1322-2, in the chip sample surface state visual feature coding map, taking the chip sample surface state visual feature initial feature vector to be enhanced as the center along the second direction to perform feature sampling. Sampling is performed to obtain a set of chip sample surface state visual second direction context feature vectors, wherein the initial feature vector of the chip sample surface state visual feature to be enhanced is located at the center of the set of chip sample surface state visual second direction context feature vectors; S1322-3, in the chip sample surface state visual feature coding diagram, feature sampling is performed along the third direction with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual third direction context feature vectors, wherein the initial feature vector of the chip sample surface state visual feature to be enhanced is located at the center of the set of chip sample surface state visual third direction context feature vectors.

[0044] Specifically, the step S1322 is expressed as follows using the feature sampling formula:

[0045]

[0046] in, and Directions The step vector, for each direction , For the direction Upsampling The channel feature vectors of adjacent positions (symmetrical sampling, the center is ), if the coordinates are out of bounds, use zero to fill in, is the sampling neighborhood value, is the sampling boundary, These are the sets of visual context feature vectors for the chip sample's surface state in the first direction, the second direction, and the third direction, respectively. It should be understood that to address the directional sensitivity limitations of standard convolution and effectively characterize the asymmetric features of spatial context, a multi-directional context feature acquisition mechanism is employed. Specifically, using the multi-dimensional feature vectors of the target pixel in the preprocessed initial feature map as a reference, feature trajectories are tracked along a preset set of directions (including typical geometric directions such as orthogonal axes and oblique directions). This not only aggregates local neighborhood information but also reveals direction-dependent structural evolution patterns through the construction of directional feature sequences. The generated directional feature sequences contain spatial topological features along specific directions, dynamic feature evolution patterns, and long-range correlation information, such as the directionality of lattice defect extension or the periodicity of metallization textures. This directional-constrained feature sampling strategy essentially achieves precise decoupling of the anisotropic characteristics of the material microstructure through directional trajectory modeling.

[0047] Specifically, in step S1323, the set of the chip sample surface state visual first direction context feature vector, the set of the chip sample surface state visual second direction context feature vector, and the set of the chip sample surface state visual third direction context feature vector are respectively input into the direction context perception device based on the converter structure to obtain the chip sample surface state visual first direction context perception implicit coding vector of the object to be enhanced, the chip sample surface state visual second direction context perception implicit coding vector of the object to be enhanced, and the chip sample surface state visual third direction context perception implicit coding vector of the object to be enhanced, which are expressed as the direction context perception formula:

[0048]

[0049] in, is a directional context sensor based on a converter structure, They are respectively the sequence of chip sample surface state visual first direction context perception implicit coding vectors, the sequence of chip sample surface state visual second direction context perception implicit coding vectors and the sequence of chip sample surface state visual third direction context perception implicit coding vectors, For extraction The center position in context-aware latent encoding vectors), These are the first-direction context-aware implicit encoding vector of the chip sample surface state, the second-direction context-aware implicit encoding vector of the chip sample surface state, and the third-direction context-aware implicit encoding vector of the chip sample surface state. It should be understood that to address the directional sensitivity limitations of standard convolution and effectively represent the asymmetric features of spatial context, a multi-directional context feature acquisition mechanism is employed. Specifically, using the multi-dimensional feature vector of the target pixel in the preprocessed initial feature map as a reference, feature trajectories are tracked along a preset set of directions (including typical geometric directions such as orthogonal axes and oblique directions). This not only aggregates local neighborhood information but also reveals direction-dependent structural evolution patterns through the construction of directional feature sequences. The generated directional feature sequences contain spatial topological features along specific directions, dynamic feature evolution patterns, and long-range correlation information, such as the directionality of lattice defect extension or the periodicity of metallization textures. This directional-constrained feature sampling strategy essentially achieves precise decoupling of the anisotropic characteristics of the material microstructure through directional trajectory modeling.

[0050] In an embodiment of the present application, the step S1324 includes: constructing a direction-sensitive attention module based on the first direction, the second direction and the third direction, and inputting the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0051] In an embodiment of the present application, the step S1324 includes: S1324-1, respectively performing anisotropy correction on the first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced to obtain the corrected first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the corrected second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the corrected third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced; S1324-2, constructing a direction-sensitive attention module based on the first direction, the second direction, and the third direction, and inputting the corrected first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the corrected second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the corrected third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

[0052] Specifically, step S1324 is expressed as the direction-sensitive attention encoding formula:

[0053]

[0054] in, for The transposed vector of and Directions and direction The embedding code vector of for The transposed vector of for and The correlation factor between For the natural constant The exponential function value with base , For direction Relative direction The attention score, and All represent the chip sample surface state visual object to be enhanced context-aware implicit coding vector in the first direction, the chip sample surface state visual object to be enhanced context-aware implicit coding vector in the second direction, and the chip sample surface state visual object to be enhanced context-aware implicit coding vector in the third direction. Not equal to , The enhanced chip sample surface state visual feature coding vector is the first vector of the enhanced chip sample surface state visual feature coding map. Channel feature vectors at pixel locations. As can be understood, the network intelligently integrates multi-directional context-aware latent encoding vectors by introducing a direction-sensitive attention mechanism. The core of this mechanism lies in breaking away from the static weighting model of traditional feature fusion and dynamically learning the contribution weights of different directional contexts in specific defect scenarios. Specifically, context-aware latent encoding vectors in three orthogonal directions (such as horizontal, vertical, and diagonal) are first input into a shared Transformer encoder, where a multi-head self-attention mechanism is used to capture the global dependencies of feature sequences in each direction. Building on this foundation, the direction-sensitive attention module further incorporates a directional identity embedding, enabling attention weight calculation based not only on the content similarity of the encoding vectors but also explicitly incorporating directional attributes. For example, when processing periodic defects along wafer scribe lines, the module dynamically enhances the weight of horizontal context while suppressing interfering vertical features. For randomly distributed microcracks, attention weights are evenly distributed across multiple directions. Through this dynamic directional weighting mechanism, the system can adaptively screen out the contextual information most relevant to the current defect type and spatial location. The final generated enhanced chip sample surface state visual feature encoding vector not only retains the texture specificity of local details, but also strengthens the cross-scale directional defect representation capability, providing subsequent classifiers with more discriminative input representation.

[0055] Given that the directional context sensor of the Transformer architecture encodes each directional feature independently, although it strengthens the local feature distribution, the attention mechanism exacerbates the directional sensitivity difference, resulting in feature coupling distortion due to the anisotropic effect when multi-directional features are fused, thereby weakening the discrimination accuracy of the coding vector. To this end, the present application introduces a high-order anisotropic polarization correction mechanism before feature aggregation, performs nonlinear reconstruction on the weighted multi-directional feature vector, compensates for the anisotropic disturbance through the direction-sensitive polarization response, eliminates cross-directional feature coupling interference, and finally generates an enhanced coding vector with high-fidelity directional features. This process effectively suppresses the feature distortion caused by differences in spatial field polarizability by dynamically correcting the directional weight distribution, significantly improving the robustness of defect representation.

[0056] Therefore, before summing, the weighted context-aware implicit encoding vector for each direction after weighting is , perform high-order anisotropic polarization representation:

[0057]

[0058] in, For the The weighted context-aware implicit encoding vector of the direction, For the The weighted context-aware implicit encoding vector of the direction, for Context-aware high-order anisotropic polarization representation vector.

[0059] Therefore, the high-order field polarizability analysis in the predetermined direction is performed in the form of a bilinear response to the other two directions except the predetermined direction.

[0060] Then, under the covariant partial derivative representation of heterogeneous analysis, we can respectively right Find the partial derivatives:

[0061]

[0062] in, is the modified context-aware high-order anisotropic polarization representation vector.

[0063] That is, since the context-aware implicit coding vectors in the other two directions are It is essentially symmetrical, so its partial derivatives are the same, which means that the calculation of partial derivatives makes an anisotropic correction to the overall field polarization response.

[0064] In this way, and Perform point multiplication weighted correction to obtain Then, calculate:

[0065]

[0066] In this way, by introducing a high-order anisotropic polarization correction mechanism, the weighted multi-directional context-aware implicit coding vector is nonlinearly reconstructed. The direction-sensitive polarization response is used to compensate for anisotropic field distortion, effectively suppressing cross-directional feature coupling interference, and ultimately generating an enhanced coding vector with high-fidelity directional features. This process significantly improves the robustness of defect representation by dynamically correcting the directional weight distribution. The fused enhanced chip sample surface state visual feature coding vector retains the specificity of local detail texture while strengthening the ability to represent cross-scale directional defects.

[0067] In the above-mentioned non-destructive analysis method of chip packaging status based on microscopic inspection, the step S140 determines the packaging status detection result based on the chip sample status multimodal fusion representation information between the enhanced chip sample surface status visual features and the enhanced chip sample internal structure visual features, and the packaging status detection result includes a predefined packaging status category label and its confidence, including: S141, fusing the chip sample surface status visual feature coding map and the chip sample internal structure visual feature coding map to obtain a chip sample status multimodal fusion representation map as the chip sample status multimodal fusion representation information; S142, inputting the chip sample status multimodal fusion representation map into a trained classifier to obtain a packaging status detection result, and the packaging status detection result includes a predefined packaging status category label and its confidence.

[0068] Specifically, step S141 fuses the chip sample surface state visual feature encoding map and the chip sample internal structure visual feature encoding map to obtain a multimodal fusion representation map of the chip sample state as the multimodal fusion representation information of the chip sample state. It should be understood that the integrity and reliability of chip packaging need to be evaluated from multiple dimensions. Visible light microscopy can primarily capture surface details, texture, color, and other information about chip packaging, and has strong detection capabilities for surface cracks, scratches, and packaging material appearance abnormalities. Infrared imaging, especially volume data constructed in combination with sequence imaging, can penetrate packaging materials and reveal internal structural defects such as voids, delamination, and internal connection abnormalities. The information obtained by these two modalities is naturally complementary, and no single modality can provide a comprehensive understanding of the overall chip packaging condition. Therefore, in order to overcome the limitations of a single modality and fully utilize the advantages of multi-source information, it is necessary to deeply integrate the high-level features extracted from different modalities. In the technical solution of this application, the chip sample surface state visual feature encoding map and the chip sample internal structure visual feature encoding map are further fused to obtain a multimodal fusion representation map of the chip sample state. This fusion is not simply a data overlay; rather, it is performed within the high-semantic feature space learned by deep learning models, aiming to capture the inherent connections and complex interactions between surface and internal structural features. For example, certain internal defects may cause subtle surface deformation or texture changes, while surface defects may indicate deeper structural issues. Through effective feature fusion, the system can comprehensively consider this cross-modal correlation information to generate a more discriminative representation than any single modality, providing the richest and most effective information foundation for subsequent defect classification. Specifically, given that visible light microscopy primarily captures surface details of chip packages, such as cracks, scratches, or material appearance anomalies, while infrared imaging can penetrate packaging materials to reveal internal structural defects such as voids, delamination, and solder defects, the information provided by these two modalities is naturally complementary. Therefore, the fusion process requires a mechanism that not only effectively combines these data from different dimensions but also mines and exploits the potential connections between them. Deep learning-based methods, particularly those with powerful feature representation capabilities, such as convolutional neural networks (CNNs) or their variants, are suitable for fusing the two encoding maps. By using a pre-trained deep learning model, the two modal input data can be encoded to obtain a high-dimensional abstract feature representation. To further enhance the fusion effect, an attention mechanism may be introduced. This dynamically adjusts the importance of different features, ensuring that the final fusion result focuses more on the most discriminative information.For example, when building a multimodal fusion model, a cross-modal attention module can be used to calculate the correlation score of each local region in each modality relative to the corresponding regions in the other modalities, thereby guiding feature selection and combination. This ensures that, even against complex backgrounds, key feature points that reflect both surface characteristics and internal structure can be accurately identified, providing strong support for subsequent classification tasks. Furthermore, in addition to direct feature-level fusion, higher-level semantic fusion strategies can also be considered. This involves not only considering pixel-level or feature vector-level data fusion, but also exploring how to combine the high-level semantic information contained in different modalities. To this end, techniques from natural language processing, such as the Transformer architecture, can be leveraged. Using sequence modeling, feature sequences from two encoded images are used as input and processed through multiple self-attention mechanisms. This approach has the advantage of automatically learning the optimal combination of features from different modalities, eliminating the need for manual definition of specific fusion rules, thereby improving the model's flexibility and generalization capabilities. Another important consideration in the fusion process is how to address the inconsistent feature distribution between modalities caused by differences in imaging principles. Because visible light images and infrared images reflect different physical properties of an object's surface and interior, respectively, they may exhibit significant statistical differences. To address this, the two encodings can be normalized before fusion to achieve similar scales and distributions, facilitating subsequent operations. Advanced regularization techniques, such as adversarial training, can also be employed to teach the model to minimize cross-modal differences while preserving their uniqueness, thereby promoting more effective fusion.

[0069] Specifically, in step S142, the multimodal fusion representation diagram of the chip sample state is input into a trained classifier to obtain a packaging state detection result, and the packaging state detection result includes a predefined packaging state category label and its confidence. It should be understood that the multimodal fusion representation diagram of the chip sample state is input into a trained classifier to obtain a packaging state detection result, and the packaging state detection result includes a predefined packaging state category label and its confidence. In this way, the trained classifier can be used to automatically and accurately judge the packaging state of the chip sample. Specifically, it is to assign the high-dimensional multimodal fusion representation diagram of the chip sample state to the most likely packaging state category according to the pattern learned by the model, and at the same time give the confidence of the judgment.

[0070] In summary, a non-destructive analysis method for chip packaging conditions based on microscopic inspection according to an embodiment of the present application is illustrated. The method first obtains the original visible light image and the original infrared image sequence and performs image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Next, the key visual features representing the surface state and internal structure of the chip are extracted, and refined and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically integrated, and a trained intelligent classifier is used to automatically identify and judge the packaging status of the chip, and give specific categories (such as normal, voids, cracks, etc.) and their reliability, thereby achieving the evaluation of its packaging quality without destroying the chip.

[0071] Figure 5 FIG. 1 is a system block diagram of a non-destructive analysis system for chip packaging conditions based on microscopic inspection according to an embodiment of the present application. Figure 5 As shown, according to an embodiment of the present application, a non-destructive analysis system 100 for chip packaging conditions based on microscopic inspection includes: a target chip sample data acquisition module 110, which is used to acquire an original visible light image and an original infrared image sequence of a target chip sample; a target chip sample data processing module 120, which is used to perform image preprocessing on the original visible light image and the original infrared image sequence respectively, and then spatially align them to obtain a registered chip sample visible light image and a registered chip sample infrared volume data; a target chip sample data same-paradigm feature enhancement module 130, which is used to perform feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features; a packaging state detection result determination module 140, which is used to determine a packaging state detection result based on the chip sample state multimodal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, wherein the packaging state detection result includes a predefined packaging state category label and its confidence level.

[0072] In the above-mentioned microscopic inspection-based non-destructive analysis system for chip packaging conditions, the target chip sample data same-paradigm feature enhancement module is used to: input the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch to obtain a chip sample surface state visual feature coding map and a chip sample internal structure visual feature coding map; perform same-paradigm feature enhancement on the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map respectively to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature.

[0073] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned chip packaging condition non-destructive analysis system based on microscopic detection have been described in the above reference. Figures 1 to 4 The nondestructive analysis method of chip packaging condition based on microscopic inspection has been introduced in detail, and therefore, its repeated description will be omitted.

[0074] In summary, a non-destructive analysis system for chip packaging conditions based on microscopic inspection according to an embodiment of the present application is illustrated. It first acquires the original visible light image and the original infrared image sequence and performs image preprocessing and spatial registration to ensure that information from different light sources and different depths can be spatially correlated. Next, the key visual features representing the surface state and internal structure of the chip are extracted, and refined and enhanced to improve the discrimination ability and robustness. Finally, these processed and enhanced surface and internal features are organically integrated, and a trained intelligent classifier is used to automatically identify and judge the packaging status of the chip, and give specific categories (such as normal, voids, cracks, etc.) and their reliability, thereby achieving the evaluation of its packaging quality without destroying the chip.

Claims

1. A non-destructive analysis method for chip packaging conditions based on microscopic inspection, characterized in that: include: Acquire the original visible light image and original infrared image sequence of the target chip sample; After performing image preprocessing on the original visible light image and the original infrared image sequence, spatial registration is performed on them to obtain registered chip sample visible light images and registered chip sample infrared volume data; The visible light image of the registered chip sample and the infrared volume data of the registered chip sample are subjected to feature extraction and same-paradigm feature enhancement through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features, including: Inputting the registered chip sample visible light image and the registered chip sample infrared volume data into a multi-input deep learning model including a first feature extraction branch and a second feature extraction branch respectively to obtain a chip sample surface state visual feature coding map and a chip sample internal structure visual feature coding map; The chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map are respectively subjected to same-paradigm feature enhancement to obtain an enhanced chip sample surface state visual feature coding map as the enhanced chip sample surface state visual feature and an enhanced chip sample internal structure visual feature coding map as the enhanced chip sample internal structure visual feature, including: extracting the channel feature vector of the (h, w)th pixel position from the chip sample surface state visual feature coding map as the initial feature vector of the chip sample surface state visual feature to be enhanced; in the chip sample surface state visual feature coding map, taking the initial feature vector of the chip sample surface state visual feature to be enhanced as the center, performing feature sampling along multiple directions to obtain a set of chip sample surface state visual first direction context feature vectors, a set of chip sample surface state visual second direction context feature vectors and a set of chip sample surface state visual third direction context feature vectors; and combining the set of chip sample surface state visual first direction context feature vectors, the chip sample The set of the second direction context feature vector of the present surface state vision and the set of the third direction context feature vector of the chip sample surface state vision are respectively input into the direction context perception device based on the converter structure to obtain the first direction context perception implicit coding vector of the chip sample surface state vision object to be enhanced, the second direction context perception implicit coding vector of the chip sample surface state vision object to be enhanced and the third direction context perception implicit coding vector of the chip sample surface state vision object to be enhanced; the first direction context perception implicit coding vector of the chip sample surface state vision object to be enhanced, the second direction context perception implicit coding vector of the chip sample surface state vision object to be enhanced and the third direction context perception implicit coding vector of the chip sample surface state vision object to be enhanced are subjected to direction-sensitive attention analysis to obtain an enhanced chip sample surface state vision feature coding vector, wherein the enhanced chip sample surface state vision feature coding vector is the channel feature vector of the (h, w)th pixel position of the enhanced chip sample surface state vision feature coding map; Based on the chip sample state multimodal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, a packaging state detection result is determined, and the packaging state detection result includes a predefined packaging state category label and its confidence.

2. The non-destructive analysis method for chip packaging conditions based on microscopic inspection according to claim 1, characterized in that: In the chip sample surface state visual feature coding map, feature sampling is performed along multiple directions with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual first direction context feature vectors, a set of chip sample surface state visual second direction context feature vectors, and a set of chip sample surface state visual third direction context feature vectors, including: In the chip sample surface state visual feature coding map, feature sampling is performed along a first direction with the chip sample surface state visual feature initial feature vector to be enhanced as the center to obtain a set of chip sample surface state visual first direction context feature vectors, wherein the chip sample surface state visual feature initial feature vector to be enhanced is located at the center position of the set of chip sample surface state visual first direction context feature vectors; In the chip sample surface state visual feature coding map, feature sampling is performed along the second direction with the chip sample surface state visual feature initial feature vector to be enhanced as the center to obtain a set of chip sample surface state visual second direction context feature vectors, wherein the chip sample surface state visual feature initial feature vector to be enhanced is located at the center position of the set of chip sample surface state visual second direction context feature vectors; In the chip sample surface state visual feature coding map, feature sampling is performed along the third direction with the initial feature vector of the chip sample surface state visual feature to be enhanced as the center to obtain a set of chip sample surface state visual third direction context feature vectors, and the initial feature vector of the chip sample surface state visual feature to be enhanced is located at the center position of the set of chip sample surface state visual third direction context feature vectors.

3. The non-destructive analysis method for chip packaging conditions based on microscopic inspection according to claim 2, characterized in that: Direction-sensitive attention analysis is performed on the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced to obtain an enhanced chip sample surface state visual feature coding vector, wherein the enhanced chip sample surface state visual feature coding vector is a channel feature vector at the (h, w)th pixel position of the enhanced chip sample surface state visual feature coding map, including: A direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

4. The non-destructive analysis method for chip packaging conditions based on microscopic inspection according to claim 3, characterized in that: A direction-sensitive attention module is constructed based on the first direction, the second direction, and the third direction, and the first direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced, and the third direction context-aware implicit coding vector of the chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector, including: Performing heterogeneity correction on the first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, the second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and the third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, respectively, to obtain a corrected first direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, a corrected second direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced, and a corrected third direction context-aware implicit coding vector of the surface state of the chip sample visual object to be enhanced; A direction-sensitive attention module is constructed based on the first direction, the second direction and the third direction, and the first direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced, the second direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced and the third direction context-aware implicit coding vector of the corrected chip sample surface state visual object to be enhanced are input into the direction-sensitive attention module to obtain the enhanced chip sample surface state visual feature coding vector.

5. The non-destructive analysis method for chip packaging conditions based on microscopic inspection according to claim 4, characterized in that: Based on the multimodal fusion representation information of the chip sample state between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features, a packaging state detection result is determined, wherein the packaging state detection result includes a predefined packaging state category label and its confidence level, including: fusing the chip sample surface state visual feature coding map and the chip sample internal structure visual feature coding map to obtain a chip sample state multimodal fusion representation map as the chip sample state multimodal fusion representation information; The multimodal fusion representation graph of the chip sample state is input into a trained classifier to obtain a packaging state detection result, wherein the packaging state detection result includes a predefined packaging state category label and its confidence.

6. A microscopic inspection-based nondestructive analysis system for chip packaging conditions, used to execute the microscopic inspection-based nondestructive analysis method for chip packaging conditions according to claim 1, characterized in that: include: A target chip sample data acquisition module is used to acquire the original visible light image and original infrared image sequence of the target chip sample; a target chip sample data processing module, configured to perform image preprocessing on the original visible light image and the original infrared image sequence, and then spatially register them to obtain registered chip sample visible light images and registered chip sample infrared volume data; A target chip sample data same-paradigm feature enhancement module is used to perform feature extraction and same-paradigm feature enhancement on the registered chip sample visible light image and the registered chip sample infrared volume data through a dual-branch multi-input deep learning model to obtain enhanced chip sample surface state visual features and enhanced chip sample internal structure visual features; The packaging state detection result determination module is used to determine the packaging state detection result based on the chip sample state multimodal fusion representation information between the enhanced chip sample surface state visual features and the enhanced chip sample internal structure visual features. The packaging state detection result includes a predefined packaging state category label and its confidence.

Citation Information

Patent Citations

  • Substation equipment defect identification method and device

    CN112734692A

  • Infrared and visible light visual information fusion method based on gradient transformation prior

    CN117173063A