Semiconductor detection method and device based on multi-scale lightweight network
By replacing the C2f module with MSEC module in the YOLOv8 model and adding a DAEM module, combined with the EDSDC detection head, the problem of insufficient suppression ability of the YOLOv8 model to complex backgrounds and difficulty in detecting small devices is solved, and the model is lightweight and accurate improvement is achieved.
Patent Information
- Application Number
- CN202510212436.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-18
AI Technical Summary
The existing YOLOv8 model has poor ability to suppress complex backgrounds when detecting semiconductor devices, and is difficult to detect small devices, and the model is not lightweight enough.
Replace some C2f modules in the BackBone network or Neck network of the YOLOv8 model with MSEC module, add DAEM module for feature enhancement, and use EDSDC detection head to build the YOLOv8n-MDE model for multi-scale feature extraction and lightweighting.
It improves the accuracy of the model's detection of small devices and its detection capabilities in complex contexts, reduces the number of model parameters, and has faster detection speed, higher accuracy and better robustness.
Smart Images

Figure CN120339670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semiconductor device detection, and particularly to a semiconductor detection method and device based on a multi-scale lightweight network. Background Art
[0002] With the rapid development of technology, the demand in the semiconductor industry continues to grow, and chip manufacturing technology is constantly advancing towards high density and multi-level directions. In this context, the internal components of a chip exhibit the characteristics of occupying a small pixel area, having a large aspect ratio, and being numerous. Traditional semiconductor device detection mainly relies on manual visual inspection, which is inefficient, subjective, and prone to errors. Therefore, there is an urgent need to develop a more efficient and accurate automatic detection method.
[0003] The application of deep learning technology has brought new possibilities to semiconductor device detection. As a representative lightweight object detection algorithm, the YOLO algorithm is suitable for tasks of automatic detection and classification in SEM images based on ML. The anchor-free single-stage architecture of this model makes it more suitable for training in SEM image datasets of different layers of different semiconductors. Among them, YOLOv8, through its efficient network structure and anchor-free design, combined with the Mosaic data augmentation technology, has excellent performance in various complex scenarios. The YOLOv8 model mainly optimizes by using the C2f module to replace the C3 module of YOLOv5 in both the backbone part and the neck, enabling the model to have stronger feature extraction capabilities and fuse different levels of features from the backbone, ensuring that the model can make full use of multi-scale information and significantly improving the inference speed. At the same time, the decoupled head structure uses two independent branches for object classification and prediction box regression, and different loss functions are used for these two types of tasks.
[0004] However, in the face of problems such as complex background interference in semiconductor images, difficult detection of small devices, and poor machine hardware in actual scenarios, the performance of the existing YOLOv8 model cannot meet the requirements of industrial production. Summary of the Invention
[0005] The technical problems to be solved by the present invention are as follows: how to solve the problems that the existing YOLOv8 model has poor ability to suppress complex backgrounds and difficult detection of small semiconductor devices when detecting semiconductor devices, and how to achieve model lightweighting.
[0006] The present invention solves the above technical problems through the following technical solutions: A semiconductor detection method based on a multi-scale lightweight network, the method comprising:
[0007] Obtain a semiconductor image, preprocess the semiconductor image to obtain a preprocessed semiconductor image;
[0008] Replace some of the C2f modules in the BackBone network or Neck network of the YOLOv8 model with MSEC modules for multi-scale feature extraction of the input features. Input the output features of the 5th module in the BackBone network and the 16th module in the Neck network, the output features of the 7th module in the BackBone network and the 19th module in the Neck network, and the output features of the 10th module in the BackBone network and the 22nd module in the Neck network into a DAEM module for feature enhancement respectively, and input the feature maps output by each DAEM module into the EDSDC detection head. Construct and train the YOLOv8n-MDE model to obtain the trained YOLOv8n-MDE model;
[0009] Input the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain the detection result.
[0010] Beneficial effects: The present invention makes improvements based on the existing YOLOv8 model. Replace some of the C2f modules in the BackBone network or Neck network with MSEC modules for multi-scale feature extraction of the input features, and achieve efficient channel utilization, improve the accuracy of the model for detecting small devices, and reduce the overall number of model parameters to achieve model lightweighting; By adding a DAEM module in the Neck network to construct a feature enhancement network, improve the detection ability of the model for small devices in complex backgrounds; Compared with the existing YOLOv8n benchmark model, the YOLOv8n-MDE semiconductor detection model constructed by the present invention has a faster detection speed and higher accuracy.
[0011] Preferably, replace the 7th and 9th C2f modules in the BackBone network of the YOLOv8 model with MSEC modules, or replace the 13th, 19th, and 22nd C2f modules in the Neck network of the YOLOv8 model with MSEC modules, or replace both the 7th and 9th C2f modules in the BackBone network and the 13th, 19th, and 22nd C2f modules in the Neck network with MSEC modules.
[0012] Beneficial effects: The present invention provides multiple implementable module replacement methods and compares the model performance of each replacement method. It is found through comparison that only replacing the 7th and 9th C2f modules in the BackBone network with MSEC modules can achieve a significant improvement in accuracy while reducing complexity, which is the best balance between performance and efficiency.
[0013] Preferably, the input features of the MSEC module are input into the Spilt layer for equal division to obtain the first feature and the second feature. The first feature is input into the first Channel rearrange layer for rearrangement to obtain the third feature and the fourth feature. The third feature is input into the 3×3Conv layer to obtain the fifth feature. The fourth feature is input into the 5×5Conv layer to obtain the sixth feature. The fifth feature and the sixth feature are input into the second Channel rearrange layer to obtain the seventh feature. The second feature and the seventh feature are input into the concatenation layer to obtain the eighth feature. The eighth feature is input into the 1×1Conv layer for channel fusion to obtain the output features of the MSEC module.
[0014] Beneficial effects: The MSEC module of the present invention uses the grouping idea to equally divide the input layer in channels, effectively utilizes the parameters of the network and maintains the richness of information. The minimum value of the number of channels is set to be not less than 64. One part of the equal division is first rearranged into a tensor containing multiple feature groups, and then grouped in parallel. Convolution operations are respectively performed using 3x3 and 5x5 convolution kernels to capture feature information at different scales, enabling the model to effectively capture details in the image. The multi-scale feature maps after extraction are partially re-stacked and merged into the original dimension, and finally concatenated with the part retaining the original channel information, and channel fusion is performed through the 1x1 convolution layer to form an overall feature map. This method can effectively reduce the mixed redundant calculation between channels, help integrate multi-scale feature information, and at the same time effectively reduce the number of parameters.
[0015] Preferably, the input features of the DAEM module are input into the Channel Recalibrate layer for channel weight recalibration to obtain the self-calibrated feature map. The self-calibrated feature map is respectively input into the Spilt layer and the convolution layer. The feature map after dimensionality reduction by the convolution layer is respectively input into the Local Attention layer and the Global Attention layer. The spatial features are obtained by capturing the image features from the spatial dimension through the Local Attention layer, and the depth features are obtained by capturing the image features from the depth dimension through the Global Attention layer. The spatial features and the depth features are input into the Sigmoid activation function to generate local and global weight factors, and feature fusion is performed on the output features of the Spilt layer according to the weight factors to obtain the output features of the DAEM module.
[0016] Beneficial effects: The DAEM module of the present invention includes a Channel Weight Recalibration sub-module (ChannelRecalibrate). By adaptively adjusting the weights of each channel, the model can pay more attention to the features relevant to the current task. The Spatial Convolution (Local Attention) and Depth Convolution (Global Attention) sub-module groups are introduced to capture image features from the spatial dimension and depth dimension respectively. Through the combination of local attention and global attention, this module can capture both fine-grained features and global context information at the same time. The DAEM module can enhance the model's recognition ability for different detailed features in complex scenes and its anti-interference ability for complex backgrounds, helping the model to perform more precisely and stably when dealing with multi-scale targets and subtle features.
[0017] Preferably, the Channel Recalibrate layer includes an AdaptiveAvgPool2d layer, a first Conv layer, and a second Conv layer. The input features of the DAEM module are dimensionally expanded to obtain a feature map with doubled number of channels. The feature map with doubled number of channels is input into the AdaptiveAvgPool2d layer, where global average pooling is performed in the spatial dimension to extract the global context information of each channel. The output features of the AdaptiveAvgPool2d layer are input into the first Conv layer for dimensionality reduction, and the dimensionally reduced features are input into the second Conv layer for feature mapping to obtain dimensionally expanded features. The dimensionally expanded features generate channel-level dynamic weights through the Sigmoid activation function. Based on the channel-level dynamic weights, the feature map with doubled number of channels is self-calibrated, and the self-calibrated feature map is added to the feature map with doubled number of channels channel by channel to obtain the self-calibrated feature map.
[0018] Preferably, the Local Attention layer includes a first 1×1Conv layer and a second 1×1Conv layer. The feature map dimensionally reduced by the convolutional layer is the input feature of the Local Attention layer. The input features of the Local Attention layer are sequentially input into the first 1×1Conv layer and the second 1×1Conv layer for convolutional dimensionality reduction to obtain spatial features.
[0019] Preferably, the Global Attention layer includes an AdaptiveAvgPool2d layer, a first 1×1Conv layer, and a second 1×1Conv layer. The feature map dimensionally reduced by the convolutional layer is the input feature of the Global Attention layer. The input features of the GlobalAttention layer are sequentially input into the AdaptiveAvgPool2d layer, the first 1×1Conv layer, and the second 1×1Conv layer to obtain depth features.
[0020] Preferably, in the EDSDC detection head, the second - layer Conv module in the detection head of the YOLOv8 model is replaced with an EDSDC module. The output features of the first - layer Conv module are the input features of the EDSDC module. The input features of the EDSDC module are input into the DCovN module to extract and fuse fine - grained local spatial features to obtain the output features of the DCovN module. The output features of the DCovN module are input into the Avg_Pool layer to reduce the output of different - depth convolutions to a spatial dimension of 1×1. The output of the Avg_Pool layer is sequentially input into the first FC layer, ReLU activation function, second FC layer, and Sigmoid activation function to obtain the importance of each channel. The input features of the EDSDC module are fused with the importance of each channel to obtain the output features of the EDSDC module.
[0021] Beneficial effects: When the existing YOLOv8 model processes objects that are relatively small or closely arranged relative to the input image, there may still be problems of missed detection or misclassification, and two sets of 3x3 convolution operations with heavy parameter usage in the detection head make the overall number of parameters of the model large. The present invention uses the lightweight detection head Detect_EDSDC to replace the detection head in the existing YOLOv8 model. By introducing the attention mechanism and depth - separable convolution, not only can the model more accurately capture target features and accurately detect targets of different scales, improving the model's recognition of small targets and complex backgrounds, but also effectively reduce the number of parameters of the model.
[0022] Preferably, the DCovN module includes n DCovNBlock layers and a set of channel feature remapping and enhancement modules. The DCovNBlock layer includes a Conv2d3×3 layer, activation function GELU, and BN2d layer. The channel feature remapping and enhancement module includes a Conv2d1×1 layer, activation function GELU, and BN2d layer. The input features of the EDSDC module are input into n DCovNBlock layers. In each DCovNBlock layer, they sequentially pass through the Conv2d3×3 layer, activation function GELU, and BN2d layer to obtain the output features of the BN2d layer. The output features of the BN2d layer and the features after fusing with the input features of the EDSDC module are input into the next layer. The output of the last DCovNBlock layer sequentially passes through the Conv2d1×1 layer, activation function GELU, and BN2d layer to obtain the output of the DCovN module.
[0023] Beneficial effects: The EDSDC module of the present invention includes a depthwise separable convolution DCovN module and a channel attention mechanism. In the DCovN module, n residual blocks DCovNBlock and a group of channel feature remapping and enhancement modules are introduced. In the DCovNBlock residual block, 3x3 depthwise convolution is used to capture local features within each channel, and the GELU smooth non-linear activation function is used to enhance the non-linear expression ability of the features. Finally, batch normalization (BatchNorm2d) is used to standardize the features, reducing the network's offset phenomenon and improving the training speed and model performance. In the channel feature remapping and enhancement module, pointwise (1x1) convolution is used to adjust the relationship between channels, and the GELU activation function is also used to enhance the features and BatchNorm2d to ensure the stability of the distribution.
[0024] The present invention also provides a semiconductor detection device based on a multi-scale lightweight network, which includes:
[0025] An image acquisition module, used to acquire a semiconductor image, preprocess the semiconductor image, and obtain a preprocessed semiconductor image;
[0026] A model construction module, used to replace some C2f modules in the BackBone network or Neck network of the YOLOv8 model with MSEC modules, input the output features of the 5th module in the BackBone network and the 16th module in the Neck network, the output features of the 7th module in the BackBone network and the 19th module in the Neck network, and the output features of the 10th module in the BackBone network and the 22nd module in the Neck network into the DAEM module for feature enhancement respectively, and input the feature map output by the DAEM module into the EDSDC detection head, construct and train the YOLOv8n-MDE model, and obtain the trained YOLOv8n-MDE model;
[0027] A detection module, used to input the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain a detection result.
[0028] The present invention also provides an electronic device, which includes:
[0029] One or more processors;
[0030] A storage device, used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the described method.
[0031] The advantages provided by the present invention are:
[0032] The detection speed of the YOLOv8n-MDE semiconductor detection model constructed by the present invention has increased by 16%, and mAP50 and mAP50:95 have increased by 4.5% and 4% respectively. The high-precision and lightweight YOLOv8n-MDE model also has better robustness, providing a valuable reference for the development of semiconductor device detection methods. Description of the Drawings
[0033] Figure 1 It is a flowchart of the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0034] Figure 2 It is an architecture diagram of the YOLOv8n-MDE model in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0035] Figure 3 It is a structural diagram of the MSEC module in the YOLOv8n-MDE model in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0036] Figure 4 It is a structural diagram of the DAEM module in the YOLOv8n-MDE model in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0037] Figure 5 It is a structural diagram of the Channel Recalibrate layer in the DAEM module in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0038] Figure 6 It is a structural diagram of the Local Attention layer in the DAEM module in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0039] Figure 7 It is a structural diagram of the Global Attention layer in the DAEM module in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0040] Figure 8 It is a structural diagram of the EDSDC detection head in the YOLOv8n-MDE model in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0041] Figure 9 It is a structural diagram of the EDSDC module in the semiconductor detection method based on a multi-scale lightweight network provided by an embodiment of the present invention;
[0042] Figure 10Schematic diagram of the semiconductor detection device based on the multi-scale lightweight network provided by the embodiments of the present invention;
[0043] Figure 11 Schematic diagram of the classification of semiconductor images collected in the experiment of the present invention;
[0044] Figure 12 Schematic diagram of the comparison of the feature maps processed by the model without adding the DAEM module and adding the DAEM module in the experiment of the present invention;
[0045] Figure 13 Schematic diagram of the comparison of the heat maps between the baseline model (without using the EDSDC detection head) and the YOLOv8n-MDE model (using the EDSDC detection head) in the experiment of the present invention;
[0046] Figure 14 Performance comparison curve of the YOLOv8n-MDE model of the present invention and the existing YOLOv8n model;
[0047] Figures 15(a) and (b) are schematic diagrams of the detection results in the comparative experiment of the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in conjunction with specific embodiments and with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0049] As Figure 1 shown, this embodiment provides a semiconductor detection method based on a multi-scale lightweight network. The method includes the following steps:
[0050] Step 1: Obtain a semiconductor image, preprocess the semiconductor image, and obtain the preprocessed semiconductor image.
[0051] Step 2: Replace some C2f modules in the BackBone network or Neck network of the YOLOv8 model with a multi-scale feature extraction module (Multi-Scale and Efficient Channel, MSEC). The MSEC module is used to perform multi-scale feature extraction on the input features. Input the output features of the 5th module in the BackBone network and the 16th module in the Neck network, the output features of the 7th module in the BackBone network and the 19th module in the Neck network, and the output features of the 10th module in the BackBone network and the 22nd module in the Neck network into a module that fuses spatial and channel attention mechanisms (Dual Attention Enhancement Module, DAEM) for feature enhancement, and input the feature maps output by each DAEM module into an EDSDC (Efficient Depthwise Separable Convolutions and DynamicChannel Attention) detection head. Construct a YOLOv8n-MDE (MSEC+DAEM+EDSDC, MDE) model and train it to obtain a trained YOLOv8n-MDE model.
[0052] Step 3: Input the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain the detection result.
[0053] The present invention makes improvements based on the existing YOLOv8 model. It replaces some C2f modules in the BackBone network or Neck network with MSEC modules, which are used to perform multi-scale feature extraction on the input features and achieve efficient channel utilization, improving the accuracy of the model for detecting small devices and reducing the overall number of model parameters. By adding a DAEM module in the Neck network to construct a feature enhancement network, the detection ability of the model for small devices in complex backgrounds is improved. The EDSDC detection head is a lightweight module that combines depthwise separable convolution and channel attention mechanism, which can help reduce the number of model parameters while improving the model's ability and detection accuracy for detecting devices of different sizes. Compared with the existing YOLOv8n benchmark model, the detection speed of the YOLOv8n-MDE semiconductor detection model constructed by the present invention is increased by 16%, and mAP50 and mAP50:95 are increased by 4.5% and 4% respectively. The high-precision and lightweight YOLOv8n-MDE model also has better robustness, providing valuable reference for the development of semiconductor device detection methods.
[0054] Example 1
[0055] See Figure 2, this embodiment provides a semiconductor detection method based on a multi-scale lightweight network, and the method includes the following steps:
[0056] Step 1: Scan and collect high-resolution pictures of the semiconductor through an SEM machine to obtain a semiconductor image, and preprocess the semiconductor image, including annotating and classifying the semiconductor image to obtain the preprocessed semiconductor image, which is used as a dedicated data set for semiconductor devices.
[0057] Step 2: Replace the seventh and ninth C2f modules in the BackBone network of the YOLOv8 model with MSEC modules respectively. Input the output features of the fifth C2f module in the BackBone network and the sixteenth C2f module in the Neck network, the output features of the seventh MSEC module in the BackBone network and the nineteenth C2f module in the Neck network, and the output features of the tenth SPPF module in the BackBone network and the twenty-second C2f module in the Neck network into a DAEM module for feature enhancement respectively, and input the feature maps output by each DAEM module into the EDSDC detection head. Construct and train the YOLOv8n-MDE model to obtain the trained YOLOv8n-MDE model.
[0058] The YOLOv8n-MDE model constructed in the present invention is an improvement based on the existing YOLOv8n model. The reason for choosing the YOLOv8n model is that based on the obtained semiconductor device data set, through a large number of comparative experiments, the performance of mainstream object detection models in the semiconductor device positioning problem is tested, and the YOLOv8n model is preferably selected as the basic model for semiconductor component positioning.
[0059] Figure 2The architecture diagram of the YOLOv8n-MDE model constructed for this invention. The YOLOv8n-MDE model includes a BackBone network (for feature extraction), a Neck network (for feature enhancement and fusion), and a Head network (for object detection and classification). Among them, the BackBone network includes: the first layer is a Conv module for initial feature extraction, with a 2-fold downsampling; the second layer is a Conv module with a 2-fold downsampling to obtain the P2 / 4 feature map; the third layer is a C2f module for feature extraction; the fourth layer is a Conv module with a 2-fold downsampling to obtain the P3 / 8 feature map; the fifth layer is a C2f module for feature extraction; the sixth layer is a Conv module with a 2-fold downsampling to obtain the P4 / 16 feature map; the seventh layer is an MSEC module for feature extraction; the eighth layer is a Conv module with a 2-fold downsampling to obtain the P5 / 32 feature map; the ninth layer is an MSEC module for feature extraction; the tenth layer is an SPPF module for multi-scale feature fusion to enhance the feature richness. The Conv module encapsulates a convolutional layer, Batch Normalization, and the activation function SiLU. The SPPF improves the feature fusion efficiency through multi-scale pooling.
[0060] The Neck network includes: the eleventh layer is an Upsample module with a 2-fold upsampling; the twelfth layer is a Concat module for concatenating with the seventh layer feature map; the thirteenth layer is a C2f module for feature fusion; the fourteenth layer is an Upsample module with a 2-fold upsampling; the fifteenth layer is a Concat module for concatenating with the fifth layer feature map; the sixteenth layer is a C2f module for feature fusion; the seventeenth layer is a Conv module for dimensionality reduction; the eighteenth layer is a Concat module for concatenating with the thirteenth layer feature map; the nineteenth layer is a C2f module for feature fusion; the twentieth layer is a Conv module for dimensionality reduction; the twenty-first layer is a Concat module for concatenating with the tenth layer feature map; the twenty-second layer is a C2f module for feature fusion.
[0061] The output feature map T3 of the SPPF module is divided into three paths. The first path enters the 11th Upsample module for upsampling to obtain the upsampled features. The second path enters the 21st Concat module, and the third path enters the 25th DAEM module (the third DAEM module). The upsampled features output by the 11th Upsample module are fused with the output features of the 7th MSEC module to obtain the feature map C1. The feature map C1 enters the 13th C2f module for feature extraction and is then divided into two paths. One path enters the 14th Upsample module for upsampling, and the other path enters the 18th Concat module. The upsampled features output by the 14th Upsample module are fused with the output features of the 5th C2f module to obtain the feature map C2. The feature map C2 enters the 16th C2f module for feature extraction and is then divided into two paths. One path enters the 17th Conv module, and the other path enters the first DAEM module. The output features of the 5th C2f module in the BackBone network and the 16th C2f module in the Neck network are input into the first DAEM module for feature enhancement, and the output feature map R1 is obtained.
[0062] The output features of the 17th Conv module are fused with the output features of the 13th C2f module to obtain the feature map C3. The feature map C3 is input into the 19th C2f module for feature extraction and is then divided into two paths. One path enters the 20th Conv module, and the other path enters the second DAEM module. The output features of the 7th MSEC module in the BackBone network and the 19th C2f module in the Neck network are input into the second DAEM module for feature enhancement, and the output feature map R2 is obtained.
[0063] The output features of the 20th Conv module are fused with the second path output features of the 10th SPPF module to output the feature map C4. The feature map C4 is input into the 22nd C2f module, and the output features of the 22nd C2f module are input into the third DAEM module. The third path output of the 10th SPPF module in the BackBone network and the output features of the 22nd C2f module in the Neck network are input into the third DAEM module for feature enhancement, and the output feature map R3 is obtained.
[0064] See Figure 3, is the structural diagram of the MSEC module. The input features of the MSEC module are input into the Spilt layer for division to obtain the first feature and the second feature. The first feature is input into the first Channel rearrange layer for rearrangement to obtain the third feature and the fourth feature. The third feature is input into the 3×3 Conv layer to obtain the fifth feature. The fourth feature is input into the 5×5 Conv layer to obtain the sixth feature. The fifth feature and the sixth feature are input into the second Channel rearrange layer to obtain the seventh feature. The second feature and the seventh feature are input into the splicing layer to obtain the eighth feature. The eighth feature is input into the 1×1 Conv layer for channel fusion to obtain the output features of the MSEC module.
[0065] In the existing YOLOv8 model, the C2f module extracts and aggregates feature information through multiple convolutional layers and residual connections. Although this design can deeply extract target features and retain important feature information during transmission, enhancing the detection ability of the model, compared with traditional convolutional modules, the slightly more complex structure of the C2f module may have its advantages offset by the computational cost in an environment with limited resources. Moreover, in the detection work of semiconductor devices, under the limited detection hardware of the SEM machine, it is necessary to simplify the feature extraction network structure to improve the detection performance of the model.
[0066] In this embodiment, the C2f modules in the seventh and ninth layers of the BackBone network of the existing YOLOv8 model are replaced with MSEC modules. Abandoning the processing of redundant channel information by the C2f module in the traditional feature extraction method, the input layer is equally divided into channels using the grouping idea, effectively utilizing the parameters of the network and maintaining the richness of information. The minimum number of channels is set to be not less than 64. One part of the equal division is first rearranged into a tensor containing multiple feature groups, and then grouped in parallel. Convolution operations are performed using 3x3 and 5x5 convolution kernels respectively to capture feature information at different scales, enabling the model to effectively capture details in the image; the multi-scale feature maps after extraction are re-stacked and merged into the original dimension, and finally spliced with the part retaining the original channel information, and channel fusion is performed through the 1x1 convolutional layer to form an overall feature map. This method can effectively reduce the mixed redundant calculation between channels, help integrate multi-scale feature information, and at the same time effectively reduce the number of parameters. The present invention uses the MSEC module to replace the C2f feature extraction module. Through multi-scale feature extraction and efficient channel utilization, it can effectively reduce the parameters while still extracting rich feature information, taking into account both the feature extraction efficiency and the computational overhead, achieving efficient feature extraction without sacrificing accuracy, and is especially suitable for real-time device detection tasks.
[0067] See Figure 4 and Figure 5, the input features of the DAEM module are input into the Channel Recalibrate layer for channel weight recalibration. The Channel Recalibrate layer includes an AdaptiveAvgPool2d layer, a first Conv layer, and a second Conv layer. The input features of the DAEM module are dimensionally expanded to obtain a feature map with doubled number of channels. The feature map with doubled number of channels is input into the AdaptiveAvgPool2d layer, where global average pooling is performed in the spatial dimension to extract the global context information of each channel. The output features of the AdaptiveAvgPool2d layer are input into the first Conv layer for dimensionality reduction, and the dimensionally reduced features are input into the second Conv layer for feature mapping to obtain dimensionally expanded features. The dimensionally expanded features generate channel-level dynamic weights through the Sigmoid activation function. Based on the channel-level dynamic weights, the feature map with doubled number of channels is self-corrected. The self-corrected feature map is added to the feature map with doubled number of channels channel by channel to obtain a self-calibrated feature map. The self-calibrated feature map is respectively input into the Spilt layer and the convolutional layer. The feature map after dimensionality reduction by the convolutional layer is respectively input into the Local Attention layer and the Global Attention layer. The Local Attention layer captures image features from the spatial dimension to obtain spatial features, and the Global Attention layer captures image features from the depth dimension to obtain depth features. The spatial features and depth features are input into the Sigmoid activation function to generate local and global weight factors, and the output features of the Spilt layer are feature fused according to the weight factors to obtain the output features of the DAEM module.
[0068] See Figure 6 and Figure 7 , the Local Attention layer includes a first 1×1 Conv layer and a second 1×1 Conv layer. The feature map after dimensionality reduction by the convolutional layer is the input feature of the Local Attention layer. The input feature of the Local Attention layer is sequentially input into the first 1×1 Conv layer and the second 1×1 Conv layer for convolutional dimensionality reduction to obtain spatial features. The Global Attention layer includes an AdaptiveAvgPool2d layer, a first 1×1 Conv layer, and a second 1×1 Conv layer. The feature map after dimensionality reduction by the convolutional layer is the input feature of the Global Attention layer. The input feature of the Global Attention layer is sequentially input into the AdaptiveAvgPool2d layer, the first 1×1 Conv layer, and the second 1×1 Conv layer to obtain depth features.
[0069] In semiconductor SEM images, components all have characteristics such as small pixel area occupation, large aspect ratio, and large quantity, which affect the feature extraction and discrimination of different components by the model. At the same time, the complex SEM machine viewing environment will also lead to different imaging qualities of products, resulting in difficulty for the model to extract various components from the background under the influence of complex backgrounds. To address this problem, the present invention adds a DAEM module to the Neck network to enhance the model's extraction of detailed features, reduce the impact of complex backgrounds on the model's detection effect, and improve the model's perception ability of different feature information at multiple scales and its anti-interference ability against complex backgrounds.
[0070] The core module of the DAEM module includes the use of a Channel Recalibrate sub-module, which enables the model to pay more attention to features relevant to the current task by adaptively adjusting the weights of each channel. The Channel Recalibrate sub-module expands the dimension of the input feature map to twice the number of channels. By doubling the channels, subsequent layers can process richer feature representations. The feature map after doubling the channels undergoes global average pooling in the spatial dimension to extract the global context information of each channel, and then the number of channels is reduced through a pointwise (1x1) convolution dimensionality reduction operation to reduce the computational complexity and extract more compact feature representations. The dimensionality-reduced features are expanded in dimension through feature mapping and a channel-level dynamic weight is generated through the Sigmoid activation function to enhance the effectiveness and adaptability of channel expression. After self-correcting the features using the generated weighted weights and adding them to the original input features channel by channel, the original information of the input features is retained to avoid feature loss caused by calibration weight errors. Finally, the self-calibrated feature map is reduced in channels to the original channel dimension, and a group of Local Attention and Global Attention sub-modules are introduced to capture image features from the spatial dimension and the depth dimension respectively. The generated spatial and depth features generate local and global weight factors w through the Sigmoid activation function, and the feature map after feature fusion according to the weight factor w is output to the next layer. Through the combination of local attention and global attention, this module can capture both fine-grained features and global context information simultaneously. The DAEM module can enhance the model's recognition ability of different detailed features in complex scenarios and its anti-interference ability against complex backgrounds, helping the model to perform more accurately and stably when processing multi-scale targets and subtle features.
[0071] See Figure 8 and Figure 9, the EDSDC detection head replaces the second-layer Conv module in the detection head of the existing YOLOv8 model with an EDSDC module. The output features of the first-layer Conv module are the input features of the EDSDC module. The input features of the EDSDC module are input into the DCovN module, which is used to extract and fuse fine-grained local spatial features. The DCovN module includes n DCovNBlock layers and a group of channel feature remapping and enhancement modules. The DCovNBlock layer includes a Conv2d 3×3 layer, an activation function GELU, and a BN2d layer. The channel feature remapping and enhancement module includes a Conv2d 1×1 layer, an activation function GELU, and a BN2d layer. The input features of the EDSDC module are input into n DCovNBlock layers. In each DCovNBlock layer, the output features of the BN2d layer are obtained by sequentially passing through the Conv2d 3×3 layer, the activation function GELU, and the BN2d layer. The features obtained by fusing the output features of the BN2d layer and the input features of the EDSDC module are input into the next layer. The output of the last DCovNBlock layer is sequentially passed through a Conv2d 1×1 layer, an activation function GELU, and a BN2d layer to obtain the output of the DCovN module. The output features of the DCovN module are input into the Avg_Pool layer to reduce the output of different-depth convolutions to a spatial dimension of 1×1. The output of the Avg_Pool layer is sequentially input into the first FC layer, the ReLU activation function, the second FC layer, and the Sigmoid activation function to obtain the importance of each channel. The input features of the EDSDC module are fused with the importance of each channel to obtain the output features of the EDSDC module.
[0072] When the existing YOLOv8 model processes objects that are relatively small or closely arranged relative to the input image, there may still be problems of missed detection or misclassification. Moreover, the two groups of 3x3 convolution operations with heavy parameter usage in the detection head of the YOLOv8 model make the overall number of parameters of the model relatively large. The present invention uses a lightweight detection head Detect_EDSDC to replace the detection head in the existing YOLOv8 model. By introducing an attention mechanism and depthwise separable convolution, not only can the model capture target features more accurately and detect targets of different scales accurately, improving the model's recognition of small targets and complex backgrounds, but also effectively reduce the number of parameters of the model.
[0073] The core part of the efficient dynamic detection head is the EDSDC module, which replaces the second group of 3x3 convolution operations in the original detection head. The EDSDC module includes a depthwise separable convolution DCovN module and a channel attention mechanism. The DCovN module introduces n residual blocks DCovNBlock and a group of channel feature remapping and enhancement modules. In the DCovNBlock residual block, 3x3 depthwise convolution is used to capture local features within each channel, and the smooth non-linear activation function GELU is used to enhance the non-linear expression ability of the features. Finally, batch normalization (BatchNorm2d) is used to standardize the features, reduce the offset phenomenon of the network, and improve the training speed and model performance. In the channel feature remapping and enhancement module, pointwise (1x1) convolution is used to adjust the relationship between channels, and the GELU activation function is also used to enhance the features and BatchNorm2d to ensure the stability of the distribution.
[0074] The depthwise separable convolution, that is, the convolution module separated by channels, can learn the importance of different channels, but ignores the information relationship between channels. To make up for this loss, the detection head introduces a global channel attention mechanism. The output of different depthwise convolutions is reduced to a 1×1 spatial dimension through Avg_Pool to extract the global features of each channel. Then, a two-layer fully connected network is used to reduce and increase the dimension, capture the interaction relationship between channels, and use Sigmoid to normalize the weights, mapping the importance of each channel to the range [0, 1]. By combining the channel attention mechanism, the EDSDC detection head enables the model to capture the complementarity and correlation between channels. By focusing on the global information and the relationship between channels, it provides the detection head with the ability to dynamically adjust the feature weights.
[0075] In the EDSDC detection head, the DCovN module is responsible for extracting and fusing fine-grained local spatial features, while the channel attention mechanism further strengthens feature selection through the dynamic weight assignment of global features. The combined design balances the computational efficiency and detection performance of the model and is suitable for the detection tasks of multiple components in the machine platform environment.
[0076] Step 3: Input the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain the detection results of semiconductor component localization.
[0077] Example 2
[0078] The difference between this example and Example 1 is that in this example, some C2f modules in the Neck network of the YOLOv8 model are replaced by MSEC modules. Specifically, the 13th, 19th, and 22nd C2f modules are replaced by MSEC modules, and other settings are the same. For the details not disclosed in the example, please refer to Example 1 of the present invention and will not be elaborated here.
[0079] Example 3
[0080] The difference between this example and Example 1 is that in this example, some C2f modules in the BackBone network and some C2f modules in the Neck network of the YOLOv8 model are replaced with MSEC modules. Specifically, the seventh and ninth C2f modules in the BackBone network and the thirteenth, nineteenth, and twenty-second C2f modules in the Neck network are replaced with MSEC modules. Other settings are the same. For details not disclosed in this example, please refer to Example 1 of the present invention and will not be elaborated here.
[0081] Example 4
[0082] This is an apparatus example that can be used to execute the method example of the present invention. For details not disclosed in the apparatus example of the present invention, please refer to the method example of the present invention and will not be elaborated here.
[0083] A semiconductor detection device based on a multi-scale lightweight network, comprising:
[0084] An image acquisition module, configured to acquire a semiconductor image, preprocess the semiconductor image, and obtain a preprocessed semiconductor image;
[0085] A model construction module, configured to replace some C2f modules in the BackBone network or Neck network of the YOLOv8 model with MSEC modules, input the output features of the fifth module in the BackBone network and the sixteenth module in the Neck network, the output features of the seventh module in the BackBone network and the nineteenth module in the Neck network, and the output features of the tenth module in the BackBone network and the twenty-second module in the Neck network into the DAEM module for feature enhancement respectively, and input the feature map output by the DAEM module into the EDSDC detection head, construct the YOLOv8n-MDE model and train it to obtain a trained YOLOv8n-MDE model;
[0086] A detection module, configured to input the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain a detection result.
[0087] Experimental analysis
[0088] The size of the experimental input image is 640*640, the maximum number of training epochs is 5000, the batchsize is 16, the SGD optimizer is used, the initial learning rate is 0.01, the momentum is 0.937, and the early stopping mechanism patience is set to 100. The same experimental configuration is used for training, testing, and validation in the experiment.
[0089] Due to the lack of a dedicated labeled dataset, a self-made dataset was used in this experiment. The acquisition device was a scanning electron microscope of the model Zeiss GeminiSEM 360, which was used to acquire images of the internal components of publicly available 15-nm chips. A total of 4,232 high-definition images of semiconductor internal components were collected, among which 2,783 labeled images were used as the training set, 759 as the test set, and 690 as the validation set. The collected images were divided into three major categories according to the device type: gate, dummy, and resistor. The entire dataset was further divided into 13 categories, as follows Figure 11 Some components are shown. Among them, the gates are subdivided by the number into: singleGate, singleGateDsd, singleGateTsd, doubleGate, doubleGateDsd, doubleGateTsd, multiGate, multiGateDsd, multiGateTsd, and 15 is publicGate. The dummies are divided by the position into: circularDummy and insideDummy.
[0090] Considering the requirements for the accuracy and effectiveness of semiconductor device detection, the present invention selects Precision, Recall, mean Average Precision (mAP), the number of parameters (Params), and the amount of computation (GFLOPs) as the main evaluation indicators.
[0091] To verify the impact of the MSEC module on the backbone network, in this experiment, the C2f module of the model was gradually replaced with the MSEC module, that is, the C2f module was gradually replaced at the seventh layer (P6 layer) of the BackBone network, the ninth layer (P8 layer), the thirteenth layer (P12 layer), the nineteenth layer (P18 layer), and the twenty-second layer (P21 layer) of the Neck network, and the impact on the model after replacement was observed. The changes in the accuracy, the number of parameters, and the amount of computation of the model after replacing the C2f module were listed in Table 1, showing the impact of different module positions on the model performance.
[0092] Table 1 Results of the comparative experiment on replacing the C2f module in multiple layers
[0093] P6 P8 P12 P18 P21 mAP50 mAP50:95 Params / M FLOPs / G 0.842 0.701 3.01 8.1 √ √ 0.878 0.734 2.86 7.9 √ √ √ 0.879 0.718 2.86 7.9 √ √ √ √ √ 0.866 0.717 2.72 7.6
[0094] It can be seen that by only replacing the P6 and P8 layers in the backbone network, the mAP50 and mAP50:95 are respectively improved by 4.39% and 2.42% compared with the baseline model, and the number of parameters and the computational amount are respectively reduced by 4.98% and 2.47% compared with the baseline model. The low-level modules (P6, P8) of the backbone network are responsible for extracting low-level features. Replacing them with the MSEC module can significantly enhance the feature expression ability while reducing the number of parameters and the computational amount. When only replacing the P12, P18, and P21 layers in the Neck network, the effect of mAP50 is close to that of only replacing P6 and P8, but mAP@50:95 is slightly lower, and the number of parameters and the computational amount are the same as those of replacing P6 and P8. It can be understood that the Neck module aggregates multi-level features. Replacing it with the MSEC module has limited performance improvement under higher IoU thresholds, indicating that its impact on global feature expression is not as significant as that of the backbone network. When replacing all modules comprehensively, although the overall number of parameters and the computational amount of the model decrease, the overall improvement effect is limited. Although comprehensive replacement reduces the complexity, the performance degradation may be due to insufficient coordination of feature extraction and aggregation caused by excessive replacement. In summary, the model's choice to replace P6 and P8 can achieve a significant improvement in accuracy while reducing complexity, which is the best balance between performance and efficiency.
[0095] To verify the effectiveness of the DAEM module, the features fused from the 15th, 18th, and 21st branches are fed into the detection head in the Neck of YOLOv8. In the present invention, the DAEM module is used to enhance the features of the feature map. The 4th and 15th, 6th and 18th, and 9th and 21st are respectively fused through the DAEM module for feature enhancement to form three new branches and are respectively sent to the detection head, forming a brand-new feature fusion network. Three feature fusion networks, namely HSFPN, FDPN, and GDFPN, are compared. Table 2 shows the experimental results of various networks applied to the YOLOv8 model.
[0096] Table 2 Experimental results of feature enhancement comparison
[0097]
[0098] As can be seen from Table 2, compared with the baseline model YOLOv8, the introduction of GDFPN leads to an overall decline in the model performance and is not suitable for the detection task of this scenario; the model accuracy of HSFPN has not been improved, but the number of parameters has decreased significantly, and the computational amount has slightly decreased; FDPN has improved the model accuracy to a certain extent, but the computational amount has increased significantly. Although the DAME module introduced in the present invention sacrifices some aspects in terms of lightweight, with the number of parameters and the computational amount increasing by 0.44M and 0.5G respectively, it takes into account both local details and global context information and significantly improves the detection performance. In particular, mAP50 and mAP50:95 are respectively improved by 3.2% and 1.4%.
[0099] To further verify the effect of feature enhancement through the DAEM module, the feature maps of the 15th, 18th, and 21st layers without enhancement and the three branch feature maps after enhancement were selected for comparison. Feature map comparison Figure 12 As shown. In the feature map extracted from the 15th branch corresponding to the small target extraction branch, the features of some small targets are not obvious in the baseline model. However, through the DAEM module that combines local attention and global attention, it can better capture fine-grained features and global context information, enhancing the model's ability to recognize the detailed features of small targets and its anti-interference ability against complex backgrounds, helping the model perform more accurately and stably when dealing with multi-scale targets and subtle features. At the same time, the enhancement effect of extracting features in the 18th and 21st branches of larger targets is not obvious.
[0100] It can be seen that the DAEM module can effectively improve the model's feature extraction ability for small targets and is suitable for the detection and recognition of components in complex semiconductor images. In summary, the innovative fusion strategy of the present invention is more suitable for device detection in semiconductor scenarios.
[0101] To verify the effectiveness of the EDSDC detection head, the EDSDCHead of the present invention was compared with mainstream detection head modules. EfficientHead, DyHead, and NMSfreeHead were respectively selected and applied to the detection head of the baseline model, and all were trained and tested in a unified experimental environment. The experimental data are shown in Table 3 below.
[0102] Table 3 Detection head comparison experiment
[0103]
[0104] From the data in Table 3, it can be seen that all indicators of EfficientHead have not improved; DyHead has a significant improvement in accuracy at the expense of a certain amount of parameters and computational complexity, but it is not conducive to the deployment of the model of the present invention on the machine; NMSfreeHead has no obvious overall improvement without increasing the configuration requirements. The EDSDCHead designed by the present invention has a better performance. On the basis of effectively reducing the number of parameters by 10% and reducing the computational complexity by 12%, the accuracy has also increased slightly.
[0105] To demonstrate the detection ability of the EDSDC detection head, the HiResCAM heatmap was selected for visualizing the results. The visualization results are as Figure 13 shown. Comparing the effects of the heatmaps, the highlighted areas in the heatmap of the baseline model are not obvious in dealing with some blurred areas, and the targets in the blurred areas are not successfully noticed. The model after introducing the channel and spatial attention EDSDC detection head can improve the model's recognition of small targets and complex backgrounds.
[0106] Comparative experiment
[0107] Taking Precision, Recall, mAP50, and mAP50:95 as the vertical coordinates and epoch as the horizontal coordinate, the performance curves are plotted. During the training phase, the proposed model YOLOv8n-MDE and the baseline model YOLOv8n are trained separately using the self-built semiconductor component dataset. The performance curves on the self-built semiconductor dataset are as Figure 14 shown.
[0108] The YOLOv8n-MDE with the orange line converges at 1072 epochs, and the baseline model YOLOv8n with the blue line converges at 942 epochs. Although YOLOv8n-MDE converges slower than the baseline model, it can be seen from the Precision and Recall graphs that the proportion of correctly predicted positive samples and the recall rate of true positive samples of the improved model are better than those of the baseline model. At the same time, it can be seen from the mAP graph that the mAP50 and mAP50:95 of the improved model are increased to 88.7% and 74.1% respectively, which are 4.5% and 4% higher than those of the baseline model. This indicates that through the improvement, the model is optimized to varying degrees in terms of precision and recall, significantly improving the object detection performance of the model.
[0109] Based on the comprehensive index data, the effectiveness of the improvement adopted in the present invention in the semiconductor component detection scenario is verified, and the improved model has applicability and practicality in the real scenario.
[0110] To further verify the superiority of the YOLOv8n-MDE algorithm proposed in the present invention, it will be compared with different lightweight detection models. The experiment selects the mainstream YOLO series lightweight models for comparison. To ensure the reliability of the experimental results, all experiments are trained and tested in a unified experimental environment.
[0111] It can be seen from the data in Table 4 that YOLOv8n-MDE shows obvious advantages in comprehensive performance. Its mAP50 reaches 0.887, which is 4%-5% higher than other models. At the same time, in terms of the mAP50:95 index, YOLOv8n-MDE also leads significantly, reaching 0.741. Although the number of parameters and the amount of computation increase slightly, they still remain within a reasonable range (3.11M parameters, 7.2 GFLOPs). In addition, its frame rate FPS is 322.6, indicating that YOLOv8n-MDE can significantly improve the detection accuracy without significantly sacrificing speed, which is particularly crucial for the task of detecting semiconductor components with high requirements for real-time performance and accuracy. The YOLOv8n-MDE proposed in the present invention performs significantly better than other detection algorithms in the task of detecting semiconductor components, and can meet the requirements of industrial production in terms of both accuracy and model size.
[0112] Table 4 Comparative Experiments of YOLO Series Lightweight Models
[0113]
[0114] To more intuitively display the detection effect of the improved model, this paper conducts verification experiments in blurred backgrounds, normal images, and complex environments with high contrast. The detection results of the YOLO series, YOLOv8n, and YOLOv8n-MDE in different images are shown in Figures 15(a) and 15(b). All visualizations are the detection results, which are the blurred background image, the normal background image, and the high-contrast image in sequence. Since the overall detection accuracy is not very high, but when detecting devices in normal images, the YOLO series all show good performance; when detecting components in blurred backgrounds, there are varying degrees of missed detections in the YOLO series, and YOLOv5n, YOLOXn, and YOLOv11n have the most missed detections; when detecting high-contrast images, YOLOXn, YOLOv11n, and YOLOv3-tiny have the most missed detections and false detections; but YOLOv8n-MDE can recall some components in complex background areas. From the overall detection effect, it can be found that it is significantly better than other models in dealing with components in complex backgrounds and detecting small devices.
[0115] Aiming at the requirements of accuracy and speed in the industrial production environment for component detection in the machine platform environment, and considering that the components to be detected have characteristics such as small pixel area, large aspect ratio, and large quantity, and the background environment where the components are located is complex, a YOLOv8n-MDE semiconductor detection model is proposed. First, by introducing the MSEC module, multi-scale feature extraction and efficient utilization of channels are realized, which reduces the model parameters while maintaining the richness of the extracted feature information; second, the DAEM module that combines spatial and channel attention mechanisms is introduced to enhance the feature expression ability, recall small targets and targets with large aspect ratios in complex backgrounds, and reduce the false detection and missed detection rates; finally, the EDSDC detection head that combines depthwise separable convolution and channel attention mechanism is used to lightweight the detection head while improving the detection accuracy. The YOLOv8n-MDE model can effectively improve the detection effect of various-shaped components in complex backgrounds, improve the detection accuracy while saving computational costs and accelerating the detection speed.
[0116] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A semiconductor detection method based on a multi-scale lightweight network, characterized in that: The method includes: Obtaining a semiconductor image, preprocessing the semiconductor image to obtain a preprocessed semiconductor image; Replacing some C2f modules in the BackBone network or Neck network of the YOLOv8 model with MSEC modules for multi-scale feature extraction of input features, inputting the output features of the 5th layer module in the BackBone network and the 16th layer module in the Neck network, the output features of the 7th layer module in the BackBone network and the 19th layer module in the Neck network, and the output features of the 10th layer module in the BackBone network and the 22nd layer module in the Neck network into a DAEM module for feature enhancement respectively, and inputting the feature maps output by each DAEM module into an EDSDC detection head, constructing and training a YOLOv8n-MDE model to obtain a trained YOLOv8n-MDE model; Inputting the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain a detection result.
2. The semiconductor detection method based on the multi-scale lightweight network according to claim 1, characterized in that: Replacing the C2f modules of the 7th and 9th layers in the BackBone network of the YOLOv8 model with MSEC modules, or replacing the C2f modules of the 13th, 19th, and 22nd layers in the Neck network of the YOLOv8 model with MSEC modules, or replacing the C2f modules of the 7th and 9th layers in the BackBone network and the C2f modules of the 13th, 19th, and 22nd layers in the Neck network with MSEC modules.
3. The semiconductor detection method based on the multi-scale lightweight network according to claim 1 or 2, characterized in that: The input feature of the MSEC module is input into a Spilt layer for division to obtain a first feature and a second feature. The first feature is input into a first Channel rearrange layer for rearrangement to obtain a third feature and a fourth feature. The third feature is input into a 3×3Conv layer to obtain a fifth feature. The fourth feature is input into a 5×5Conv layer to obtain a sixth feature. The fifth feature and the sixth feature are input into a second Channel rearrange layer to obtain a seventh feature. The second feature and the seventh feature are input into a concatenation layer to obtain an eighth feature. The eighth feature is input into a 1×1Conv layer for channel fusion to obtain the output feature of the MSEC module.
4. The semiconductor detection method based on a multi-scale lightweight network according to claim 1, wherein: The input feature of the DAEM module is input into a Channel Recalibrate layer for channel weight recalibration to obtain a self-calibrated feature map. The self-calibrated feature map is respectively input into a Spilt layer and a convolutional layer. The feature map after dimensionality reduction by the convolutional layer is respectively input into a Local Attention layer and a Global Attention layer. The Local Attention layer captures image features from the spatial dimension to obtain spatial features. The Global Attention layer captures image features from the depth dimension to obtain depth features. The spatial features and the depth features are input into a Sigmoid activation function to generate local and global weight factors, and the output features of the Spilt layer are feature-fused according to the weight factors to obtain the output feature of the DAEM module.
5. The semiconductor detection method based on the multi-scale lightweight network according to claim 4, wherein: The Channel Recalibrate layer includes an AdaptiveAvgPool2d layer, a first Conv layer, and a second Conv layer, which expand the dimensions of the input features of the DAEM module to obtain a feature map with doubled number of channels. The feature map with doubled number of channels is input into the AdaptiveAvgPool2d layer, where global average pooling is performed in the spatial dimension to extract the global context information of each channel. The output features of the AdaptiveAvgPool2d layer are input into the first Conv layer for dimensionality reduction, and the reduced-dimensional features are input into the second Conv layer for feature mapping to obtain dimensionally expanded features. The dimensionally expanded features generate channel-level dynamic weights through the Sigmoid activation function, and the feature map with doubled number of channels is self-calibrated based on the channel-level dynamic weights. The self-calibrated feature map is added to the feature map with doubled number of channels channel by channel to obtain the self-calibrated feature map.
6. The semiconductor detection method based on the multi-scale lightweight network according to claim 4, characterized in that: The Local Attention layer includes a first 1×1Conv layer and a second 1×1Conv layer. The feature map after dimensionality reduction by the convolutional layer is the input feature of the Local Attention layer. The input feature of the Local Attention layer is sequentially input into the first 1×1Conv layer and the second 1×1Conv layer for convolutional dimensionality reduction to obtain spatial features.
7. The semiconductor detection method based on a multi-scale lightweight network according to claim 4, wherein: The Global Attention layer includes an AdaptiveAvgPool2d layer, a first 1×1Conv layer, and a second 1×1Conv layer. The feature map after dimensionality reduction by the convolutional layer is the input feature of the Global Attention layer. The input feature of the Global Attention layer is sequentially input into the AdaptiveAvgPool2d layer, the first 1×1Conv layer, and the second 1×1Conv layer to obtain depth features.
8. The semiconductor detection method based on a multi-scale lightweight network according to claim 1, characterized in that: The EDSDC detection head replaces the 2nd layer Conv module in the detection head of the YOLOv8 model with the EDSDC module. The output features of the 1st layer Conv module are the input features of the EDSDC module. The input features of the EDSDC module are input into the DCovN module to extract and fuse fine-grained local spatial features to obtain the output features of the DCovN module. The output features of the DCovN module are input into the Avg_Pool layer to reduce the outputs of different-depth convolutions to a spatial dimension of 1×1. The outputs of the Avg_Pool layer are sequentially input into the first FC layer, the ReLU activation function, the second FC layer, and the Sigmoid activation function to obtain the importance of each channel. The input features of the EDSDC module are fused with the importance of each channel to obtain the output features of the EDSDC module.
9. The semiconductor detection method based on a multi-scale lightweight network according to claim 8, wherein: The DCovN module includes n DCovNBlock layers and a group of channel feature remapping and enhancement modules. The DCovNBlock layer includes a Conv2d 3×3 layer, an activation function GELU, and a BN2d layer. The channel feature remapping and enhancement module includes a Conv2d 1×1 layer, an activation function GELU, and a BN2d layer. The input features of the EDSDC module are input into n DCovNBlock layers. In each DCovNBlock layer, the output features of the BN2d layer are obtained by sequentially passing through the Conv2d 3×3 layer, the activation function GELU, and the BN2d layer. The features obtained by fusing the output features of the BN2d layer and the input features of the EDSDC module are input into the next layer. The output of the last DCovNBlock layer is sequentially passed through the Conv2d 1×1 layer, the activation function GELU, and the BN2d layer to obtain the output of the DCovN module.
10. A semiconductor detection device based on a multi-scale lightweight network, characterized in that: The device includes: an image acquisition module, configured to acquire a semiconductor image, preprocess the semiconductor image, and obtain a preprocessed semiconductor image; a model construction module, configured to replace some C2f modules in the BackBone network or Neck network of the YOLOv8 model with MSEC modules, input the output features of the 5th module in the BackBone network and the 16th module in the Neck network, the output features of the 7th module in the BackBone network and the 19th module in the Neck network, and the output features of the 10th module in the BackBone network and the 22nd module in the Neck network into the DAEM module for feature enhancement respectively, and input the feature map output by the DAEM module into the EDSDC detection head, construct and train the YOLOv8n-MDE model to obtain the trained YOLOv8n-MDE model; a detection module, configured to input the preprocessed semiconductor image into the trained YOLOv8n-MDE model to obtain a detection result.