Chip packaging defect detection method applied to edge device based on YOLOv11m
By introducing Starnet backbone network, SimAM attention mechanism and GSConv lightweight convolution technology into the YOLOv11m model, the chip package defect detection model is optimized, solving the problems of low efficiency of traditional detection methods and excessive resource utilization of model deployment resources, and achieving efficient and fast detection results.
Patent Information
- Application Number
- CN202510694812.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Traditional chip detection methods have problems such as low efficiency, high misjudgment rate and high equipment cost. The existing YOLOv11m model has slow detection speed, large model size, and excessive resource use when deployed on edge devices.
Based on the YOLOv11m model, Starnet backbone network, SimAM attention mechanism and GSConv lightweight convolution technology are introduced to optimize the model structure to reduce the computational volume and model size, while improving the detection speed and feasibility of edge device deployment.
While maintaining detection accuracy, the speed of chip package defect detection is significantly improved, the model size and calculation amount are reduced, and the feasibility of deployment on edge devices is improved.
Smart Images

Figure CN120219388A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision target detection, and in particular to a chip packaging defect detection method based on YOLOv11m applied to edge devices. Background Art
[0002] At a time when science and technology are changing with each passing day, the semiconductor industry, as the core and cornerstone of the modern information technology industry, is booming. The rise of emerging technologies such as 5G communications, artificial intelligence, and the Internet of Things has led to an explosive growth in chip demand. Not only has the number continued to rise, but the requirements for performance, size, and integration have also become increasingly stringent. As chip integration continues to increase, the chip packaging link has become increasingly critical. During the packaging process, any slight deviation may cause packaging defects on the surface of the chip, and these defects seriously affect the reliability and performance of the chip. In large-scale chip production, it is necessary to accurately detect various surface packaging defects while ensuring efficient detection speed. This has become a key problem that the industry needs to solve urgently.
[0003] Traditional chip inspection methods have many drawbacks. Manual appearance inspection is highly subjective, difficult to detect tiny defects, and extremely inefficient; although automatic optical inspection has improved inspection efficiency to a certain extent, its ability to identify complex textures and multi-layer structures on the chip surface is poor, and when there are interference factors such as reflections and shadows on the chip surface, it is easy to make misjudgments; although electron beam inspection can provide high-resolution images, the equipment procurement cost is high, and there is a potential risk of electron beam damage to the chip surface during the inspection process. The inspection speed is also relatively slow, and the inspection time per square centimeter is long. In the pursuit of efficient and non-destructive testing, large-scale application is hindered.
[0004] Although YOLOv11m performs well in the field of target detection, it still has obvious shortcomings when applied to chip packaging defect detection. The types of chip packaging defects are complex and small in size. The ordinary YOLOv11m is not capable of detecting small targets, which can easily lead to missed detections. At the same time, the chip production line has extremely high requirements for real-time detection. The ordinary YOLOv11m has a large amount of calculation under high-resolution images, and the inference speed is difficult to meet the standards. In addition, the resources of edge computing devices commonly used in chip workshops are limited, and the huge model parameters of ordinary YOLOv11m are difficult to achieve lightweight deployment. These problems make it difficult for traditional YOLOv11m to balance detection accuracy, real-time performance and deployment feasibility, and it is in urgent need of improvement.
[0005] Similar demands also exist in other fields, and there have already been some effective improvement solutions. For example, the lightweight hierarchical architecture of the Starnet backbone network can reduce the computational load and improve the feature extraction ability in fields such as small target detection in medical images and real-time detection in autonomous driving. The SimAM attention mechanism and the GSConv lightweight convolution technology also demonstrate their value in scenarios such as focusing on key features in remote sensing images and optimizing the computational cost of in-vehicle edge devices. In the field of chip packaging defect detection, the hierarchical architecture of Starnet adapts to the characteristics of small targets and reduces the computational load. The SimAM attention mechanism adapts to the complex background and easily disturbed defects of images, and the GSConv technology is suitable for low-computing-power edge devices. These improved technologies have shown good adaptability and application potential in this field.
[0006] Facing these problems, traditional detection methods and existing object detection models have gradually revealed their limitations, making it difficult to meet the growing industry demands. Moreover, the solutions to similar demands in other fields have not been effectively applied to this field. Therefore, there is an urgent need for an innovative and efficient detection model to fill this gap. Summary of the Invention
[0007] Based on the original YOLOv11 model, the present invention provides a method for detecting chip packaging defects based on YOLOv11m for edge devices. While maintaining the detection accuracy, it aims to improve the detection speed of chip packaging defects, reduce the model size and computational load, and enhance the feasibility of deployment on edge devices.
[0008] Specifically, the present invention is implemented through the following technical solutions. A method for detecting chip packaging defects based on YOLOv11m for edge devices proposed according to the present invention comprises the following steps:
[0009] Step 1: Take an image of the packaged chip, mark the coordinates of the defect positions in the image, and export the coordinate file and the image to form a dataset.
[0010] Step 2: Perform data augmentation on the dataset, including cropping and stitching the image, flipping the image, and adjusting the brightness, contrast, and saturation of the image.
[0011] Step 3: Divide the dataset into a training set, a test set, and a validation set according to the ratio of 8:1:1.
[0012] Step 4: Build a YOLOv11-ALE object detection model, whose network structure includes a Backbone module, a Neck module, and a Head module, which are responsible for feature extraction, feature fusion, and target prediction respectively.
[0013] The Backbone module adopts the Starnet network structure, and then replaces the original C2PSA module with the C2CGA module to adjust the feature association between channels, and then adds the SimAM attention module to strengthen the key feature output;
[0014] The Neck module introduces the lightweight convolution technology GSConv to replace the original Conv module, combines the standard convolution with the deep convolution, outputs through a 1×1 convolution mixed channel, and then replaces the original C3k2 module with the MSR-VoVGSCSP module based on GSConv, and completes feature processing and structural optimization through cross-level connection and fusion strategies;
[0015] The Head module adopts the LWNBDet detection head, which uses a bidirectional feature propagation architecture to connect and fuse multi-scale features across scales, determine the importance of features in a weighted manner, and achieve efficient target detection;
[0016] Step 5, import the data set into the network for training to obtain an improved target detection model;
[0017] Step 6: Convert the obtained model format into an edge device compatible format, transplant it to the edge device, and detect chip packaging defects.
[0018] Furthermore, the data enhancement of the data set includes: cropping the defective portion of the image, and then evenly arranging and splicing 4-8 cropped images into one image; flipping the image horizontally or vertically; increasing the brightness, contrast, and saturation of the image by 25% and decreasing it by 25% respectively; and combining the above methods on the same image to make the defect features more prominent, which helps the model better capture the subtle details of the defect edges.
[0019] Furthermore, the Starnet network structure adopts a four-stage hierarchical architecture. First, the input image is downsampled through a convolutional layer to change the size and dimension of the input feature map; batch normalization is used to replace the original normalization and is placed at the back to facilitate the fusion operation during inference; GELU is replaced with ReLU6 to simplify the calculation; then it enters the Star Blocks to further extract features. Second, it is downsampled again through a convolutional layer, the size of the feature map is further reduced, and the number of channels increases. Then, more abstract features are extracted by the Star Blocks. Then, the convolutional downsampling process is repeated, the resolution of the feature map continues to decrease, and the number of channels continues to double. The Star Blocks make the feature expression more abundant. Finally, it is downsampled by convolution again, and then features are extracted by the Star Blocks, and the result is output through global average pooling and a fully connected layer. At the end of each module, depthwise separable convolutions are added. The depthwise convolution extracts local features, and the pointwise convolution integrates channel information; the star operation block has a dual-branch structure. After linear transformation and element-wise multiplication for fusion and residual connection, the output is enhanced for feature expression; the network channel expansion factor is fixed at 4, and the width doubles in each stage. The number of blocks and the embedded channels are adjusted to adapt to different scenarios.
[0020] Furthermore, in the benchmark model network structure of the Backbone module, the C2PSA module is replaced with the C2CGA module. After the C2CGA is embedded in the benchmark model, the input feature map is sliced into multiple groups by channel. The input number of channels is 1024, divided into 16 groups, with 64 channels in each group. Each feature group calculates the attention independently; during cascaded fusion, the previous output is concatenated with the current group for processing, enhancing cross-scale interaction, optimizing the network structure, and adapting to chip package defect detection.
[0021] Furthermore, the SimAM attention mechanism is introduced into the Backbone module. First, it receives the feature map with dimensions N×C×H×W output from the previous layer; then, an energy function is constructed based on the "spatial inhibition" theory in neuroscience, and the importance values of each neuron are obtained through solution. A 3D attention weight covering both the channel and spatial dimensions is generated; then, after these weights are processed by the sigmoid function, they are multiplied element-wise with the original feature map to highlight key features and suppress redundancy; finally, the processed feature map is output.
[0022] Furthermore, the lightweight convolution technology GSConv divides the input feature map into two groups. One group goes through a standard convolution, and the other group goes through a depthwise separable convolution. Then, the features are fused through channel rearrangement, reducing the computational amount while retaining the feature expression ability.
[0023] Furthermore, the MSR-VoVGSCSP module is constructed based on GSConv. The input feature map is divided into two branches through cross-stage connection. GSConv operations are performed on each branch, and multi-scale 3×3 and 5×5 convolutional kernels are introduced in parallel to process the branch features and capture information at different scales. The output of each branch is element-wise added to the shallow features passed through the residual connection for fusion, and then the branch features are fused through concatenation to optimize the feature pyramid structure and enhance the feature representation ability.
[0024] Furthermore, the LWNB Detector head constructs a multi-level feature interaction system, and extracts feature maps with resolutions of 1 / 16, 1 / 32, 1 / 64, and 1 / 128 from different convolutional stages of the backbone network. By building a bidirectional parallel conduction link, element-wise addition and channel concatenation operations are performed on the upsampled and downsampled features; nodes that only transmit as a single data stream are removed, and bypass connections are established between features at the same level to enhance information reuse. For each input feature branch, a trainable weight vector is deployed for channel-wise weighting, and the feature fusion coefficient is calculated based on the normalized fusion formula. Finally, the fused feature matrix is input into the bounding box-class prediction network composed of shared-parameter convolutional layers to achieve defect location regression and class determination. Figure 1 Furthermore, the format of the obtained model is converted into an edge device-compatible format. After being transplanted to the edge device, the following steps are used to implement encapsulated chip defect detection: establish a feature library containing encapsulated defects, scratch defects, pin defects, and stain defects; set a confidence threshold to judge the validity of the result according to the confidence of the model recognition result; implement defect size measurement through the pixel-physical size calibration algorithm; set an early warning mechanism that generates different levels of quality alerts according to production requirements, defect types, sizes, and confidences.
[0025] The YOLOv11-ALE chip encapsulation defect detection model of the present invention optimizes the YOLOv11m model, introduces the StarNet backbone network, and the lightweight hierarchical architecture reduces the computational load and adapts to the extraction of small target defect features; adopts the CSCGA module grouping mechanism to reduce the computational complexity and enhance cross-group feature interaction at the same time; adds the SimAM attention mechanism to focus on the key feature points of the defect; improves the original convolution operation, and GSConv combines standard convolution and depth convolution to reduce the computational cost and accelerate the inference of the Neck layer. The MSR-VoVGSCSP module improved based on GSConv improves the cross-stage connection efficiency and adapts to defect features of different sizes; with the help of the bidirectional feature propagation architecture of the LWNB Detector head, cross-scale connection fuses multi-scale features and weights to determine the importance of features, reduces invalid operations, and while maintaining the detection accuracy, greatly improves the detection speed of chip encapsulation defects, reduces the model size and computational load, and improves the feasibility of deployment on edge devices.
[0026] The YOLOv11-ALE chip encapsulation defect detection model of the present invention optimizes the YOLOv11m model, introduces the StarNet backbone network, and the lightweight hierarchical architecture reduces the computational load and adapts to the extraction of small target defect features; adopts the CSCGA module grouping mechanism to reduce the computational complexity and enhance cross-group feature interaction at the same time; adds the SimAM attention mechanism to focus on the key feature points of the defect; improves the original convolution operation, and GSConv combines standard convolution and depth convolution to reduce the computational cost and accelerate the inference of the Neck layer. The MSR-VoVGSCSP module improved based on GSConv improves the cross-stage connection efficiency and adapts to defect features of different sizes; with the help of the bidirectional feature propagation architecture of the LWNB Detector head, cross-scale connection fuses multi-scale features and weights to determine the importance of features, reduces invalid operations, and while maintaining the detection accuracy, greatly improves the detection speed of chip encapsulation defects, reduces the model size and computational load, and improves the feasibility of deployment on edge devices. Brief Description of the Drawings
[0027] Figure 1 is a flowchart of a lightweight chip packaging defect detection model based on YOLOv11m and its construction method of the present invention.
[0028] Figure 2 is a schematic diagram of the network structure of the improved YOLOv11-ALE object detection model of the present invention.
[0029] Figure 3 is a schematic diagram of the structure of the SimAM attention mechanism module.
[0030] Figure 4 is a schematic diagram of the structure of the GSConv module.
[0031] Figure 5 is a schematic diagram of the MSR-VoVGSCSP structure.
[0032] Figure 6 is a schematic diagram of the structure of the LWNB Detector head.
[0033] Figure 7 is a comparison chart of the training indicators of YOLOv11 and YOLOv11-ALE changing with the number of training rounds, where Figure 7 a is a comparison chart of the precision rates of YOLOv11 and YOLOv11-ALE; Figure 7 b is a comparison chart of the recall rates of YOLOv11 and YOLOv11-ALE; Figure 7 c is a comparison chart of the mean average precision with an intersection over union threshold of 0.5 for YOLOv11 and YOLOv11-ALE; Figure 7 d is a comparison chart of the mean average precision with an intersection over union threshold between 0.5 and 0.95 for YOLOv11 and YOLOv11-ALE.
[0034] Figure 8 is a comparison chart of the loss curves of YOLOv11 and YOLOv11-ALE, where Figure 8 a is a comparison chart of the bounding box losses of the training sets of YOLOv11 and YOLOv11-ALE; Figure 8 b is a comparison chart of the stepwise focal losses of the training sets of YOLOv11 and YOLOv11-ALE; Figure 8 c is a comparison chart of the classification losses of the training sets of YOLOv11 and YOLOv11-ALE; Figure 8 d is a comparison chart of the bounding box losses of the validation sets of YOLOv11 and YOLOv11-ALE; Figure 8 e is a comparison chart of the stepwise focal losses of the validation sets of YOLOv11 and YOLOv11-ALE; Figure 8Figure f shows the comparison of classification losses between YOLOv11 and YOLOv11-ALE on the validation set.
[0035] Figure 9 These are the detection effect diagrams of the YOLOv11-ALE chip packaging defect detection network. Detailed implementation manners
[0036] To make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0037] Figure 1 This is the flowchart of the method of the present invention. A lightweight chip packaging defect detection model based on YOLOv11m and its construction method provided by the present invention specifically include the following steps:
[0038] Step 1: Take images of the packaged chips, mark the coordinates of the defect positions in the images, and export the coordinate files and images to form a dataset.
[0039] Step 2: Perform data augmentation on the dataset, including cropping and stitching the images, flipping the images, and adjusting the brightness, contrast and saturation of the images.
[0040] In step 2, the data augmentation of the dataset includes: cropping out the part of the image containing the defect, and then evenly arranging and stitching 4-8 cropped images into one image; flipping the image horizontally or vertically respectively; increasing and decreasing the brightness, contrast and saturation of the image by 25% respectively; and combining the above methods for the same image, so as to make the defect features more prominent and help the model better capture the subtle details of the defect edges.
[0041] Step 3: Write a Python script file to divide the pictures in the public dataset into a training set, a test set and a validation set according to the ratio of 8:1:1. Subsequently, write a program to convert the labels of the original dataset into the format applicable to the YOLO model, and create a summary file data.yaml storing the relative paths and categories of each dataset.
[0042] Step 4: Build a YOLOv11-ALE object detection model, the network structure of which includes a Backbone module, a Neck module and a Head module, which are responsible for feature extraction, feature fusion and target prediction respectively.
[0043] The StarNet network structure is added to the Backbone module. It adopts a four-stage hierarchical architecture. First, the input image is downsampled through a convolutional layer to change the size and dimension of the input feature map; batch normalization is used to replace the original normalization and is placed after it to facilitate the fusion operation during inference; GELU is replaced with ReLU6 to simplify the calculation; then it enters the Star Blocks to further extract features. Secondly, it is downsampled again through a convolutional layer, the size of the feature map is further reduced and the number of channels increases, and then more abstract features are extracted by the StarBlocks. Then, the process of convolutional downsampling is repeated, the resolution of the feature map continues to decrease and the number of channels continues to double, and the feature expression becomes richer through the Star Blocks. Finally, it is downsampled by convolution again, and then features are extracted by the Star Blocks, and the result is output through global average pooling and a fully connected layer. Depthwise separable convolutions are added at the end of each module, with depthwise convolutions extracting local features and pointwise convolutions integrating channel information; the star operation block has a dual-branch structure, which fuses through linear transformation, element-wise multiplication and residual connection to output, enhancing the feature expression; the network channel expansion factor is fixed at 4, the width doubles in each stage, and the number of blocks and the embedded channels are adjusted to adapt to different scenarios.
[0044] Its mathematical expression is as follows:
[0045]
[0046]
[0047]
[0048] Among them, i and j are used to index the channels, and α is the coefficient of each term:
[0049]
[0050] After rewriting the star operation, it is expanded into a combination of different terms, and each term (related to x) shows a non-linear association, indicating independent and implicit dimensions. Finally, by stacking multiple layers, the implicit dimensions are increased infinitely in a recursive manner.
[0051] The formula for batch normalization is:
[0052]
[0053] Among them, is the original feature value of the th channel in the input feature map.
[0054] Secondly, replace the C2PSA module in the benchmark model network structure of the Backbone module with the C2CGA module. After the C2CGA is embedded in the benchmark model, the input feature map is sliced into multiple groups by channels. The number of input channels is 1024, which is divided into 16 groups, with 64 channels in each group. Each feature group calculates the attention independently. During cascading fusion, the previous output is concatenated with the current group to enhance cross-scale interaction, optimize the network structure, and adapt to chip package defect detection.
[0055] The formula for grouped attention is as follows:
[0056]
[0057] Among them, is the learnable parameter matrix, represents the number of channels of the input feature map, represents the number of groups into which the input feature map is divided in the channel dimension.
[0058] The formula for the attention score is as follows:
[0059]
[0060] Among them, is the scaling factor, which is used to prevent gradient vanishing.
[0061] The formula for cross-scale interaction is as follows:
[0062]
[0063] Among them, Projection is the dimension alignment operation, and Fuse is the fusion function.
[0064] Then, introduce the SimAM attention mechanism into the Backbone module. First, receive the feature map with a dimension of output from the previous layer. Then, construct an energy function based on the "spatial inhibition" theory of neuroscience to solve the importance values of each neuron, and generate 3D attention weights covering both channel and spatial dimensions. Subsequently, after processing these weights through the sigmoid function, multiply them element-wise with the original feature map to highlight key features and suppress redundancy. Finally, output the processed feature map.
[0065] The energy function is as follows:
[0066]
[0067] Among them, is the target neuron, is other neurons within the same channel, is the number of neurons within the channel, and are the weights and biases of the linear transformation, is the regularization coefficient.
[0068] The minimum energy at each position can be obtained by the following formula:
[0069]
[0070] where, represents the mean of all neurons (including the target neuron t) in the same channel of the feature map, represents the variance of all neurons (including the target neuron t) in the same channel of the feature map.
[0071] The sigmoid function is as follows:
[0072]
[0073] where, integrates the , represents element-wise multiplication.
[0074] In the Neck module, the Conv module in the baseline model is replaced with the GSConv module, as Figure 4 shown. GSConv divides the input feature map into two groups. One group undergoes standard convolution, and the other group undergoes depthwise separable convolution. Then, the features are fused through channel rearrangement, reducing the computational complexity while retaining the feature expression ability.
[0075] In the Neck module, the C3k2 module is replaced with the MSR-VoVGSCSP module, as Figure 5 shown. It is constructed based on GSConv. The input feature map is divided into two branches through cross-stage connection. GSConv operations are performed on each branch, and multi-scale 3×3 and 5×5 convolutional kernels are introduced in parallel to process the branch features and capture information at different scales. The output of each branch is element-wise added to the shallow features passed through the residual connection for fusion, and then the branch features are fused through concatenation to optimize the feature pyramid structure and enhance the feature representation ability.
[0076] The multi-scale feature fusion formula is as follows:
[0077]
[0078]
[0079]
[0080] where, is the output feature map of the 3×3 convolution branch, is the output feature map of the 5×5 dilated convolution branch, BN represents batch normalization, and ReLU represents the activation function. represents a 3×3 dilated convolution with a dilation rate of 2. is the output feature map after multi-scale feature fusion.
[0081] The LWNBDet detection head is adopted in the Head module. As Figure 6 shown, it constructs a multi-level feature interaction system, extracts feature maps with resolutions of 1 / 16, 1 / 32, 1 / 64, and 1 / 128 from different convolution stages of the backbone network. By building a two-way parallel conduction link, element-wise addition and channel concatenation operations are performed on the upsampled and downsampled features; nodes that only transmit as a single data stream are removed, and bypass connections are established between features at the same level to enhance information reuse. For each input feature branch, a trainable weight vector is deployed for channel dimension weighting, and the feature fusion coefficient is calculated based on the normalized fusion formula. Finally, the fused feature matrix is input into the bounding box-class prediction network composed of shared parameter convolution layers to achieve defect location regression and class determination. Figure 1 / 16, 1 / 32, 1 / 64, 1 / 128 of the original. Through building a two-way parallel conduction link, element-wise addition and channel concatenation operations are performed on the upsampled and downsampled features; nodes that only transmit as a single data stream are removed, and bypass connections are established between features at the same level to enhance information reuse. For each input feature branch, a trainable weight vector is deployed for channel dimension weighting, and the feature fusion coefficient is calculated based on the normalized fusion formula. Finally, the fused feature matrix is input into the bounding box-class prediction network composed of shared parameter convolution layers to achieve defect location regression and class determination.
[0082] The formula for fast normalized fusion is:
[0083]
[0084] is the normalized weight coefficient of the k-th input feature, ranging from [0,1], indicating the contribution degree of this feature during the fusion process; is the original weight parameter that can be learned and is dynamically adjusted through network training; m is the total number of feature branches participating in the fusion.
[0085] Step 5: Import the dataset into the network for training to obtain the improved YOLOv11-ALE object detection model, and evaluate the training results.
[0086] The operations in Step 5 are as follows: Statistically calculate the computational complexity, number of parameters, model size, and the speed of processing images of the model, as well as Precision and mAP50 on the test set, and compare them with the original YOLOv11 object detection method. The results are shown in Table 1 below.
[0087] Table 1
[0088] It can be seen from the results in Table 1 that compared with the original YOLOv11 model, although the accuracy of the model of the present invention is slightly reduced, the GFLOPs are reduced by 74.2%, the number of parameters is reduced by 56.0%, the model size is reduced by 55.6%, and at the same time, the FPS is increased by 88.2%, meeting the requirements of detection accuracy and the feasibility of deployment on edge devices.
[0089] Step 6: Convert the format of the obtained model into an edge device-compatible format, transplant it to the edge device, and then implement encapsulated chip defect detection through the following steps: establish a feature library containing encapsulated defects, scratch defects, pin defects, and stain defects; set a confidence threshold to judge the validity of the result based on the confidence of the model recognition result; implement defect size measurement through a pixel-physical size calibration algorithm; set up an early warning mechanism to generate different levels of quality alerts according to production requirements, defect types, sizes, and confidence levels.
[0090] The model training environment of the present invention is as follows: the CPU uses Intel(R) Core(TM) i9-13900HX @ 2.2 GHz, the GPU uses NVIDIA GeForce RTX4080 Laptop, the running memory is 32GB, the operating system is Windows 11, the CUDA version is 12.6, the deep learning framework is PyTorch, based on the YOLO11m model, and the embedded development version is NVIDIA JetsonOrin NX 16GB.
Claims
1. A chip package defect detection method based on YOLOv11m applied to edge devices, characterized in that, The steps include: Step 1: Take an image of the packaged chip, mark the coordinates of the defect positions in the image, and export the coordinate file and the image to form a data set; Step 2: Data enhancement of the dataset, including cropping and splicing the images, flipping the images, and adjusting the brightness, contrast, and saturation of the images; Step 3: Divide the data set into training set, test set, and validation set in a ratio of 8:1:1; Step 4: Build the YOLOv11-ALE target detection model. Its network structure includes Backbone module, Neck module and Head module, which are responsible for feature extraction, feature fusion and target prediction respectively. The Backbone module adopts the Starnet network structure, and then replaces the original C2PSA module with the C2CGA module to adjust the feature association between channels, and then adds the SimAM attention module to strengthen the key feature output; The Neck module introduces the lightweight convolution technology GSConv to replace the original Conv module, combines the standard convolution with the deep convolution, outputs through a 1×1 convolution mixed channel, and then replaces the original C3k2 module with the MSR-VoVGSCSP module based on GSConv, and completes feature processing and structural optimization through cross-level connection and fusion strategies; The Head module adopts the LWNBDet detection head, and with the help of a bidirectional feature propagation architecture, it connects and fuses multi-scale features across scales and determines the importance of features in a weighted manner; Step 5, import the data set into the network for training to obtain an improved target detection model; Step 6: Convert the obtained model format into an edge device compatible format, transplant it to the edge device, and detect chip packaging defects.
2. The chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, wherein The data enhancement of the data set includes: cropping the defective portion of the image, and then evenly arranging and splicing 4-8 cropped images into one image; flipping the image horizontally or vertically; and increasing and decreasing the brightness, contrast, and saturation of the image by 25% and 25%, respectively.
3. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, The Starnet network structure adopts a four-stage hierarchical architecture. First, the input image is processed by convolution layer, batch normalization and ReLU6 activation function to extract features preliminarily, and then enters Star Blocks to further extract features; secondly, it is downsampled again through the convolution layer, the feature map size is further reduced, the number of channels is increased, and then Star Blocks extract more abstract features; then, the convolution downsampling process is repeated, the feature map resolution continues to decrease, the number of channels continues to double, and Star Blocks makes the feature expression richer; finally, convolution downsampling is performed again, and then features are extracted by Star Blocks, and the results are output through global average pooling and fully connected layers.
4. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, Replace the C2PSA module in the baseline model network structure of the Backbone module with the C2CGA module. After the C2CGA is embedded in the baseline model, the input feature map is sliced into multiple groups by channels. The number of input channels is 1024, which is divided into 16 groups, and the number of channels in each group is 64. Each feature group calculates the attention independently. During cascade fusion, the previous output is concatenated with the current group to enhance cross-scale interaction and optimize the network structure.
5. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, Introduce the SimAM attention mechanism in the Backbone module. First, receive the feature map output from the previous layer. Then, construct an energy function based on the "spatial inhibition" theory in neuroscience, solve to obtain the importance values of each neuron, and generate 3D attention weights covering both channel and spatial dimensions. Subsequently, after processing these weights through the sigmoid function, multiply them element-wise with the original feature map. Finally, output the processed feature map.
6. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, The lightweight convolution technology GSConv divides the input feature map into two groups. One group undergoes standard convolution, and the other group undergoes depthwise separable convolution. Then, the features are fused through channel rearrangement.
7. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, The MSR-VoVGSCSP module is constructed based on GSConv. The input feature map is divided into two branches through cross-level connection. GSConv operations are performed on each branch, and 3×3 and 5×5 convolutional kernels of multiple scales are introduced in parallel to process the branch features and capture information at different scales. The output of each branch is element-wise added and fused with the shallow features passed through residual connection, and then the branch features are fused through concatenation.
8. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, The LWNBDet detection head is based on a bidirectional feature propagation architecture, obtaining a differential scale feature set from different layers of the backbone network. By constructing a cross-level feature circulation path, redundant unidirectional connection nodes are eliminated, direct connection channels for same-level features are added, and an adaptive weight allocation mechanism is used to parameterize each input feature. After being optimized by the normalization fusion strategy, the integrated features are output to the target localization and classification prediction module.
9. A chip package defect detection method based on YOLOv11m applied to edge devices according to claim 1, characterized in that, Convert the format of the obtained model into a format compatible with edge devices. After transplanting it to the edge device, the encapsulation chip defect detection is realized through the following steps: establish a feature library including encapsulation defects, scratch defects, pin defects, and stain defects; set the confidence threshold according to production needs, and judge the validity of the result based on the confidence of the model recognition result; realize defect size measurement through the pixel-physical size calibration algorithm; set an early warning mechanism that generates different levels of quality alerts according to production requirements, defect types, sizes, and confidences.
Citation Information
Patent Citations
On-load tap-changer fault diagnosis method based on lightweight YOLO11
CN119556128A
Quartz ring defect detection method based on YOLOv11
CN119831941A
Lightweight mobile phone screen defect detection method based on improved YOLOv10
CN119904449A
Photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention
CN119992213A
Building wall defect detection method based on improved YOLOv10 network
CN120014427A
Cited By
Food preservation box detection method and system based on neural network
CN120877056A
Lightweight YOLO network and method for solar panel damage detection
CN121353189A
Lightweight YOLO network and method for solar panel damage detection
CN121353189B
Chip advanced packaging defect detection method and system based on mmmember detection framework
CN121391822A
Screening method and system for key regulatory pathway subsets of large medical model
CN121790010A