High-precision lightweight focusing window intelligent selection method based on improved YOLOv8

Through the improved YOLOv8 network, deformable convolution, attention mechanism and lightweight module are introduced, which solves the problem of focusing window selection on wafer marking images by traditional methods, and realizes efficient and accurate intelligent selection of focusing windows.

CN120388164APending Publication Date: 2025-07-29NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510475003.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The traditional focus window selection method is difficult to adapt to the complex and changeable wafer mark image characteristics, resulting in focus imaging deviations and affecting the accuracy of the manufacturing process.

Method used

The improved YOLOv8 network is adopted, and by introducing a deformable convolution module, a hybrid domain attention mechanism module and a Shape-IoU loss function, combined with the Ghost module, it is lightweight to improve feature extraction and prediction accuracy and reduce calculation complexity.

Benefits of technology

It realizes intelligent selection of high-precision and lightweight focusing windows, improves detection efficiency and accuracy, adapts to the diversity of wafer mark shapes, and meets real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388164A_ABST
    Figure CN120388164A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision lightweight focusing window intelligent selection method based on improved YOLOv8, and the method comprises the steps: carrying out the optimization of a mark focusing window into a target detection task, and achieving the precise focusing imaging of a mark in a wafer image. By introducing a deformable convolution DCN module, the learning ability of the network for complex mark structure features is enhanced, and the diversity of wafer mark shapes can be more flexibly adapted; in combination with an attention mechanism CBAM module, the attention of the network to the key mark area is enhanced, and the ability of understanding complex images and predicting features is effectively improved; a Shape-IoU loss function is adopted to optimize regression of a target bounding box, so that the prediction precision is improved, and the target positioning capability of the model is enhanced; and the network is improved by using the lightweight Ghost module, so that the calculation complexity and the parameter quantity are remarkably reduced, and the detection efficiency is improved. According to the invention, intelligent selection of the wafer marking image focusing window in the automatic focusing process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of semiconductor optical precision metrology, and particularly relates to a high-precision lightweight focusing window intelligent selection method based on improved YOLOv8. Background Art

[0002] With the rapid progress of modern technology, the semiconductor industry occupies an extremely important position in the global economy and technology fields. Nowadays, the advanced semiconductor chip manufacturing processes are continuously advancing, the feature sizes are continuously shrinking, and the structural complexity is continuously increasing, which also puts forward higher requirements for the accuracy and stability of manufacturing processes. And the precise wafer marking autofocus technology has become a key core element to ensure the smooth implementation of manufacturing processes. Wafer marking images usually contain complex circuit backgrounds and various marking types. In this case, accurate and rapid window selection is an important basis for achieving precise and efficient focused imaging. However, most traditional focusing window selection methods rely on manual experience or simple threshold settings and are difficult to adapt to the complex and changeable wafer marking image features nowadays. Facing wafers of different batches and specifications, these methods often greatly reduce the accuracy of window selection, resulting in deviation of focused imaging and affecting the precision of subsequent manufacturing processes. Therefore, the present invention optimizes the selection of the focusing window of the marking image into an object detection process to meet the intelligent selection needs of the focusing window of a single wafer marking image.

[0003] At present, typical object detection algorithms can be roughly divided into three categories: traditional algorithms, two-stage and one-stage algorithms. Traditional object detection algorithms have high detection flexibility, strong object localization ability, and simple principles with good interpretability. However, they have high computational complexity, are sensitive to noise and occlusion, and have subjectivity and blindness in window or region selection. The two-stage object detection algorithm is proposed to make up for the deficiencies of traditional object detection algorithms in terms of robustness and multi-class detection. This type of algorithm combines deep learning technology and greatly improves the detection accuracy by means of a phased strategy. However, this method is limited by a relatively complex network structure, and the overall detection speed is relatively slow, making it difficult to meet scenarios with extremely high real-time requirements. To overcome the shortcomings of the two-stage object detection algorithm R-CNN, some optimization methods have also been proposed in subsequent research: Fast R-CNN improves the computational efficiency by sharing convolutional features; Faster R-CNN introduces a region proposal network to achieve end-to-end training, further improving the detection speed and accuracy; Mask R-CNN adds an instance segmentation function on the basis of Faster R-CNN, enabling the simultaneous detection and segmentation of objects. Compared with two-stage algorithms, one-stage object detection algorithms directly regard the object detection task as an end-to-end regression problem and can achieve a one-time prediction process from the input image to the output object category and position coordinates. In the face of complex image scenes with diverse object sizes, this type of algorithm can still maintain high detection accuracy and stability, effectively improving the overall performance and applicability of object detection. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a high-precision lightweight focusing window intelligent selection method based on improved YOLOv8, which enhances the network's learning ability for complex marker structure features, can more flexibly adapt to the diversity of wafer marker shapes; effectively improves the ability to understand complex images and predict features; not only improves the prediction accuracy, but also enhances the model's object localization ability; significantly reduces the computational complexity and the number of parameters, and improves the detection efficiency.

[0005] To achieve the above object, the present invention provides a high-precision lightweight focusing window intelligent selection method based on improved YOLOv8, including:

[0006] Obtain a number of single-layer and double-layer wafer marker images, preprocess the single-layer and double-layer wafer marker images to obtain preprocessed images;

[0007] Construct an initial improved YOLOv8 model, and train the initial improved YOLOv8 model according to the preprocessed images to obtain an optimized improved YOLOv8 model;

[0008] Obtain the target single-layer and double-layer wafer marking images, input them into the optimized improved YOLOv8 model, and complete the intelligent selection of the high-precision lightweight focusing window for the target single-layer and double-layer wafer marking images.

[0009] Optionally, constructing the initial improved YOLOv8 model includes:

[0010] Introduce a deformable convolution module into the backbone network and feature fusion layer of the YOLOv8 network, and extract the features of the wafer marking by dynamically adjusting the offset of the convolution kernel sampling points;

[0011] Add a hybrid-domain attention mechanism module between the backbone network and the feature fusion layer to perform channel and spatial attention weighting on the input feature map;

[0012] Use the Shape-IoU loss function to replace the traditional loss function and calculate the shape and scale matching between the predicted box and the ground truth box;

[0013] Introduce the Ghost module into the convolutional layer, generate the original feature layer and supplementary feature layer through linear transformation, and stack them to form the output feature map.

[0014] Optionally, the introduction positions of the deformable convolution module include the last C2f module of the backbone network and the last layer of the feature fusion layer, and it is only replaced with the deformable convolution module in the convolutional layer at the introduction position.

[0015] Optionally, the implementation process of the hybrid-domain attention mechanism module includes:

[0016] Perform channel attention weighting on the input feature map to generate a channel weight matrix;

[0017] Perform spatial attention weighting on the input feature map to generate a spatial weight matrix;

[0018] Multiply the channel weight matrix and the spatial weight matrix element by element to obtain the final feature weight;

[0019] Multiply the final feature weight and the original feature map element by element.

[0020] Optionally, the calculation of the Shape-IoU loss function includes the matching of two parts: area ratio and shape ratio:

[0021]

[0022]

[0023]

[0024]

[0025]

[0026] Among them, scale is the scaling factor, and w t and h t are the length and width of the ground truth box respectively, and w p and h p are the length and width of the predicted box respectively, x t and y t are the center coordinates of the ground truth box, and x p and y p are the center coordinates of the predicted box, I w and I h are the weight coefficients in the horizontal and vertical directions of the ground truth box respectively, Ω shape is the shape weight, L IoU is the traditional loss function, and L Shape-IoU is the Shape-IoU loss function.

[0027] Optionally, the construction of the Ghost module includes:

[0028] Generating an original feature layer through a conventional convolution operation;

[0029] Performing a depthwise separable convolution operation on the original feature layer to generate a supplementary feature layer;

[0030] Superimposing the original feature layer and the supplementary feature layer along the channel dimension.

[0031] Optionally, the channel attention weighting is achieved through the joint output of global average pooling and max pooling, and the spatial attention weighting is achieved through the concatenation of the pooled feature maps and a convolution operation.

[0032] Optionally, the depthwise separable convolution operation includes pointwise convolution and depthwise convolution, where the kernel size of the pointwise convolution is 3×3 and the kernel size of the depthwise convolution is 1×1.

[0033] Technical effects of the present invention:

[0034] (1) In the present invention, a DCN module is added to optimize the key structures (backbone network and feature fusion layer) of the network, improving the quality of feature extraction of the YOLOv8 network for labeled images, enhancing its robustness to shape offset, rotation, and deformation of the labels, and ensuring the detection effect in scenarios with blurred label edges or double-layer stacking.

[0035] (2) The CBAM module added by the present invention between the backbone network and the feature fusion layer, on the one hand, ensures that the backbone network focuses on the region of interest in the image and effectively extracts target features; on the other hand, it further strengthens the attention to key feature points before feature fusion, thereby improving the accuracy of feature prediction.

[0036] (3) The present invention selects the Shape-IoU loss function to specifically optimize the loss function in the YOLOv8 network, focusing on the matching situation between the predicted bounding box and the ground truth bounding box in terms of shape and scale. Based on IoU, it directly optimizes the aspect ratio of the bounding box, making the shapes of the predicted bounding box and the ground truth bounding box more similar, and effectively improving the fitting degree of the predicted bounding box to the target shape.

[0037] (4) The present invention introduces a Ghost module in the convolutional layer. This module reduces the number of conventional convolutions through simple linear transformations, thereby achieving lightweight improvement of the network, minimizing the computational amount as much as possible, and alleviating the demand for the performance of the computing platform in network detection training. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings constituting a part of this application are used to provide further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0039] Figure 1 is the optimization flowchart of the YOLOv8 network in the embodiment of the present invention;

[0040] Figure 2 is the schematic structural diagram of the deformable convolutional network DCN model in the embodiment of the present invention;

[0041] Figure 3 is the schematic structural diagram of the attention mechanism CBAM model in the embodiment of the present invention;

[0042] Figure 4 is the schematic structural diagram of the lightweight Ghost model in the embodiment of the present invention;

[0043] Figure 5 is the schematic structural diagram of the optimized YOLOv8-Optimized network in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine with the embodiments to detail this application.

[0045] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0046] As Figure 1 shown, in this embodiment, a high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 is provided, including:

[0047] Obtain a number of single-layer and double-layer wafer marking images, preprocess the single-layer and double-layer wafer marking images to obtain a preprocessed image;

[0048] Construct an initial improved YOLOv8 model, and train the initial improved YOLOv8 model according to the preprocessed image to obtain an optimized improved YOLOv8 model;

[0049] Obtain a target single-layer and double-layer wafer marking image, input it into the optimized improved YOLOv8 model, and complete the intelligent selection of the high-precision lightweight focusing window of the target single-layer and double-layer wafer marking image.

[0050] Further, constructing the initial improved YOLOv8 model includes:

[0051] Introduce a deformable convolution module in the backbone network and the feature fusion layer of the YOLOv8 network, and extract the features of the wafer marking by dynamically adjusting the offset of the convolution kernel sampling points;

[0052] Add a hybrid-domain attention mechanism module between the backbone network and the feature fusion layer to perform channel and spatial attention weighting on the input feature map;

[0053] Use the Shape-IoU loss function to replace the traditional loss function to calculate the shape and scale matching between the predicted box and the ground truth box;

[0054] Introduce a Ghost module in the convolutional layer, generate the original feature layer and the supplementary feature layer through linear transformation, and stack them to form an output feature map.

[0055] Further, the introduction positions of the deformable convolution module include the last C2f module of the backbone network and the last layer of the feature fusion layer, and it is only replaced with the deformable convolution module in the convolutional layer at the introduction position.

[0056] Further, the implementation process of the hybrid-domain attention mechanism module includes:

[0057] Perform channel attention weighting on the input feature map to generate a channel weight matrix;

[0058] Spatially attention-weight the input feature map to generate a spatial weight matrix;

[0059] Element-wise multiply the channel weight matrix and the spatial weight matrix to obtain the final feature weight;

[0060] Element-wise multiply the final feature weight and the original feature map.

[0061] Furthermore, the calculation of the Shape-IoU loss function includes the matching of two parts: the area ratio and the shape ratio:

[0062]

[0063]

[0064]

[0065]

[0066]

[0067] where scale is the scaling factor, w t and h t are the length and width of the ground truth box respectively, w p and h p are the length and width of the predicted box respectively, x t and y t are the center coordinates of the ground truth box, x p and y p are the center coordinates of the predicted box, I w and I h are the weight coefficients in the horizontal and vertical directions of the ground truth box respectively, Ω shape is the shape weight, L IoU is the traditional loss function, and L Shape-IoU is the Shape-IoU loss function.

[0068] Furthermore, the construction of the Ghost module includes:

[0069] Generate the original feature layer through conventional convolution operations;

[0070] Perform depthwise separable convolution operations on the original feature layer to generate a supplementary feature layer;

[0071] Stack the original feature layer and the supplementary feature layer along the channel dimension.

[0072] Furthermore, the channel attention weighting is achieved through the combined output of global average pooling and max pooling, and the spatial attention weighting is achieved through the concatenation of the pooled feature maps and convolution operations.

[0073] Furthermore, the depthwise separable convolution operation includes channel-wise convolution and pointwise convolution. The kernel size of the channel-wise convolution is 3×3, and the kernel size of the pointwise convolution is 1×1.

[0074] The present invention introduces the deformable convolutional network DCN into the YOLOv8 network model, as Figure 2 shown. By introducing learnable offsets on the basis of traditional convolution, the sampling points of the convolution kernel become dynamically adjustable, enabling it to more flexibly adapt to diverse changes in the shape, size, and position of the marks. In DCN, by learning the offsets of each sampling point, the shape and size of the convolution kernel are adaptively optimized and adjusted, thereby accurately obtaining the most suitable convolution kernel shape for that area. For the input feature map, each sampling point therein undergoes convolution processing through an adaptive convolution layer to obtain the corresponding offset, resulting in a more accurate output feature map. In this way, the wafer mark feature extraction can be accurately and efficiently achieved, laying a solid foundation for the subsequent analysis and processing of the network.

[0075] The present invention adds a typical hybrid-domain CBAM attention mechanism module to the YOLOv8 network model, as Figure 3 shown. This module simultaneously adopts average pooling and max pooling strategies, enabling the network to effectively promote the fusion of features at different levels and different channels, helping the network better capture multi-scale information, and enhancing the ability to understand complex mark images. On the premise of ensuring that the network performance is not affected, this method reduces the computational complexity to a certain extent, enabling the model to maintain a high operating efficiency while maintaining high performance, thus meeting the dual requirements of efficiency and accuracy in the wafer mark autofocus scenario.

[0076] The present invention optimizes the loss function Shape-IoU in the YOLOv8 network model. The core calculation part of this function usually includes the matching of two parts: the area ratio and the shape ratio:

[0077]

[0078]

[0079]

[0080]

[0081]

[0082] Where scale is the scaling factor, which is determined by the size of the targets in the training dataset. Factors such as the distribution characteristics, mean, and variance of the target sizes in the training dataset will all affect the value of the scaling factor, and it plays a key role in adjusting the overall scale throughout the calculation process. I w 、I h represent the weight coefficients in the horizontal and vertical directions of the ground truth box respectively. When the horizontal size of the ground truth box is relatively long, I w may have a relatively large value, and when the length of the ground truth box in the vertical direction is more prominent, I h will increase accordingly.

[0083] The present invention adopts the Ghost module lightweight module in the YOLOv8 network model, as Figure 4 shown. This module first processes the input features using a conventional convolution to obtain the necessary feature information contained in the input features, and based on this, an original feature layer is generated. Based on this original feature layer, depthwise separable convolutions are used to perform pointwise convolution operations to generate additional feature layers. This feature layer is a powerful supplement to the original feature layer, enriching the overall feature representation and enabling the network to analyze the information of the input features from more dimensions. Finally, the two feature layers obtained above are superimposed and integrated to finally form a complete output feature map. The Ghost module reduces the computational complexity while ensuring that the network can output high-quality feature maps, providing strong support for subsequent object detection and other tasks, and enabling the YOLOv8 network to achieve a better balance between accuracy and efficiency.

[0084] By improving and optimizing these four aspects of the YOLOv8 network, its performance in the wafer marking detection task has been comprehensively improved. The synergistic effect of these improvements has successfully achieved the balanced optimization of the network's object detection performance and efficiency. The optimized YOLOv8-Optimized network structure is as Figure 5 shown.

[0085] To verify the performance improvement of the improved YOLOv8 network above, the present invention conducted comparative experiments on the deformable convolution DCN, the attention mechanism CBAM, the loss function Shape-IoU, and the lightweight Ghost module. At the same time, ablation experiments were designed to verify the overall performance enhancement effect brought to the model when these four optimization strategies work together.

[0086] (1) Experimental dataset construction

[0087] In the semiconductor field, there is relatively little publicly available information on wafer marking-related research on the Internet, and there is no publicly available dataset that can be directly used for the experimental verification of the present invention. Therefore, the present invention uses the simulation software Lumerical FDTD Solutions to generate wafer marking images and construct a dataset. During the simulation process, to further enrich the dataset and improve its generalization ability, 1000 single- and double-layer wafer marking images with different shapes, sizes, and clarity levels are generated by selecting different numerical apertures, wavelengths, magnification factors, and imaging methods. Subsequently, the generated image set is sorted and labeled to construct the training dataset for the network. Among them, 800 images are randomly selected for network training, and the remaining 200 images are used for the testing session.

[0088] (2) Experimental parameter settings

[0089] The present invention uses the PyTorch deep learning framework with a simple framework, flexible operation, and rich toolchain to conduct experimental verification on the improved and optimized YOLOv8 network. To ensure the accuracy and reliability of the experimental results, during the experiment, all experimental environments are set according to Table 1, and the hyperparameters in the network are set as shown in Table 2.

[0090] Table 1

[0091]

[0092]

[0093] Table 2

[0094]

[0095] (3) Experiments on the deformable convolution DCN module

[0096] For the introduced position of the selected deformable convolution module in the analysis of the present invention (the last layer of the backbone network and the feature fusion layer), five groups of network models are designed: the general YOLOv8 network model, the YOLOv8-DCN-B and YOLOv8-DCN-N network models with replacements in the last layer of the backbone network and the feature fusion layer respectively, the YOLOv8-DCN-BN network model with replacements in both places, and the YOLOv8-DCN-A network model with replacements in all convolutional layers. Experimental verification is carried out on them using the wafer marking dataset, and the obtained performance indicators are shown in Table 3.

[0097] Table 3

[0098]

[0099] As can be seen from Table 3, with the increase in the deformable convolution module, the mAP@.5:.95 detection accuracy gradually improves. The YOLOv8-DCN-BN network model has a 2.8% improvement compared to the general network model, and the computational complexity is reduced by 4.9%. From the performance indicators of the YOLOv8-DCN-A network model, after replacing all the convolutional layers in the network, the number of parameters and the size of the network model increase by 15.2% and 12.7% respectively compared to the general network model, but the mAP@.5:.95 shows a downward trend, decreasing by 3.8% compared to YOLOv8-DCN-BN. It can be seen that the selection of the position of the deformable convolution module introduced in the present invention can not only exert the performance of the convolutional network and improve the detection accuracy, but also effectively avoid the problem that the computational complexity increases sharply due to the addition of too many modules, resulting in a decline in network performance.

[0100] (4) Experiment on the CBAM Module of the Attention Mechanism

[0101] For the CBAM attention mechanism module added in the present invention, two other typical attention mechanism modules (SE and ECA attention mechanisms) were added at the same position respectively, and then the network after addition was verified by comparative experiments using the wafer marking dataset. The obtained performance indicators are shown in Table 4.

[0102] Table 4

[0103]

[0104] As can be seen from Table 4, the increase in the attention mechanism can effectively improve the mAP@.5:.95 detection accuracy of the network, and the number of parameters and the size of the network model after improvement only increase slightly. Among them, the addition of the CBAM attention mechanism has the best effect on improving the network performance. The YOLOv8-CBAM network model has only a 1.2% increase in the number of parameters compared to the general network model, but the mAP@.5:.95 increases by 4.9%.

[0105] (5) Experiment on the Shape-IoU Module of the Loss Function

[0106] Based on the general YOLOv8 network, the loss functions were replaced with GIoU, WIoU, and Shape-IoU loss functions respectively to generate three groups of control network models. The four groups of networks were verified by comparative experiments using the wafer marking dataset. The obtained performance indicators are shown in Table 5.

[0107] Table 5

[0108]

[0109]

[0110] As can be seen from Table 5, after replacing the WIoU loss function with the general network model, the mAP@.5:.95 detection accuracy decreased by 1.6%, while both the GIoU and Shape-IoU loss functions improved, and the Shape-IoU loss function had the best improvement effect, increasing by 3.3%.

[0111] (6) Experiments on lightweight Ghost module

[0112] As can be seen from the above three sets of experimental data, after optimizing and improving the general YOLOv8 network, along with the improvement of performance, the number of parameters, the amount of computation, and the network size all increased. To avoid the increased performance requirements of the network for the computing platform, the present invention uses the Ghost module to perform lightweight processing on the unimproved convolutional module. To verify the lightweight effect of this module, two sets of control networks were designed, and two other typical lightweight modules, MobileNet and ShuffleNet, were used to conduct experimental verification, and the obtained performance indicators are shown in Table 6.

[0113] Table 6

[0114]

[0115] From the above data analysis, it can be obtained that in the lightweight model, compared with the general network model, the YOLOv8-Ghost network model has significantly reduced the number of network parameters by 33.7%, reduced the amount of computation by 30.9%, and also reduced the network model size by 31.7%, showing the best lightweight effect among the three modules. The lightweight of the model will inevitably have a certain impact on the detection accuracy of the network. Compared with the other two lightweight modules, the Ghost module has the least decrease in mAP@.5:.95, only decreasing by 1.6%.

[0116] (7) Ablation experiments

[0117] To verify the overall performance gain of the YOLOv8-Optimized network after introducing deformable convolution, adding attention mechanism, replacing the loss function, and performing lightweight improvement, the present invention optimizes the network model in the above order on the basis of the general YOLOv8 network, and uses the wafer marking dataset for experimental verification. The performance indicators of all optimized networks are shown in Table 7.

[0118] Table 7

[0119]

[0120] As can be seen from Table 7, after the first three optimizations, the detection accuracy of the network has increased by 7.1% compared to the general network. The number of network model parameters gradually increases, and the final number of network parameters obtained increases by 11.4% compared to the general network. At the same time, although the deformable convolution DCN module can effectively reduce the increase in computational complexity, after optimizing the network with the attention mechanism CBAM and the loss function Shape-IoU, the computational complexity of the network still increases by 4.9% compared to the general network. After lightweighting the network model, based on the network model obtained from the first three optimizations, the number of parameters is reduced by 15.8%, the computational complexity is reduced by 31.8%, and the mAP@.5:.95 detection accuracy only drops by 1.3%. The overall performance of the finally obtained YOLOv8-Optimized network is also far superior to the unoptimized general network. The number of parameters is reduced by 6.2%, the computational complexity is reduced by 8.4%, the mAP@.5:.95 detection accuracy is increased by 5.8%, and the size of the network model is reduced by 4.8%.

[0121] Based on the problems of poor adaptability, insufficient accuracy, and low detection efficiency of the traditional focusing windowing algorithm for complex marker structures and backgrounds, and aiming at the application background of the precise bounding box selection requirement for markers in wafer images, this invention optimizes the marker focusing windowing into an object detection task, and proposes a high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8.

[0122] To achieve this goal, the technical solutions adopted by this invention specifically include the following steps:

[0123] To improve the quality of feature extraction of the YOLOv8 network for marker images, enhance its robustness to marker shape offset, rotation, and deformation, and ensure the detection effect in the scenarios of blurred marker edges or double-layer stacked scenes, this invention adds a DCN module to optimize the key structures of the network - the backbone network and the feature fusion layer (Neck).

[0124] (1) This invention chooses to introduce a deformable convolution module at the position of the last C2f module in the backbone network. Introducing the deformable convolution at this position can not only make full use of its deformable convolution performance to accurately capture complex features such as possible shape offset, rotation, and deformation of wafer markers, but also reduce the consumption of overall computing resources to ensure the efficiency and stability of network operation.

[0125] (2) To give full play to the advantages of deformable convolution in improving feature fusion and capturing scale changes, and to control the negative impact brought by the increase in computational complexity, this invention chooses to introduce a deformable convolution module at the last layer of this part.

[0126] To improve the feature extraction performance of the YOLOv8n network, the present invention adds a typical hybrid-domain CBAM attention mechanism module to its network model.

[0127] The added CBAM module is located between the backbone network and the feature fusion layer. By assigning weights to the target features, on the one hand, it ensures that the backbone network focuses on the regions of interest in the image and effectively extracts the target features; on the other hand, it further strengthens the attention to key feature points before feature fusion, thereby improving the accuracy of feature prediction.

[0128] Aiming at the deficiencies of traditional loss functions, the present invention selects the Shape-IoU loss function to optimize the loss function in the network. Compared with the traditional IoU loss function, this function focuses on the matching situation of the shape and scale between the predicted bounding box and the ground truth bounding box, directly optimizes the aspect ratio of the bounding box based on IoU, making the shapes of the predicted bounding box and the ground truth bounding box more similar, effectively improving the fitting degree of the predicted bounding box to the target shape. At the same time, it also reasonably constrains the area difference between the predicted bounding box and the ground truth bounding box, avoiding problems such as the generated predicted bounding box being too large or too small and not conforming to the actual situation.

[0129] To minimize the computational amount as much as possible and alleviate the demand for the performance of the computing platform in network detection training, the present invention introduces a Ghost module in the convolutional layer. The Ghost module reduces the number of conventional convolutions through simple linear transformations, thereby achieving lightweight improvement of the network.

[0130] This module first processes the input features using conventional convolutions to obtain the necessary feature information contained in the input features, and generates an original feature layer based on this. Based on this original feature layer, depthwise separable convolutions are used to perform pointwise convolution operations to generate additional feature layers. This feature layer is a powerful supplement to the original feature layer, enriching the overall feature representation and enabling the network to analyze the information of the input features from more dimensions. Finally, the two feature layers obtained above are superimposed and integrated to finally form a complete output feature map. The Ghost module not only reduces the computational complexity but also ensures that the network can output high-quality feature maps, providing strong support for subsequent object detection and other tasks, and enabling the YOLOv8 network to achieve a better balance between accuracy and efficiency.

[0131] (1) The present invention optimizes by adding a DCN module to the key structures (backbone network and feature fusion layer) of the network, improving the quality of feature extraction of the YOLOv8 network for labeled images, enhancing its robustness to shape offset, rotation, and deformation of the labels, and ensuring the detection effect in scenarios with blurred label edges or double-layer stacking.

[0132] (2) The CBAM module added by the present invention between the backbone network and the feature fusion layer ensures, on the one hand, that the backbone network focuses on the regions of interest in the image and effectively extracts target features; on the other hand, it further strengthens the attention to key feature points before feature fusion, thereby improving the accuracy of feature prediction.

[0133] (3) The present invention selects the Shape-IoU loss function to specifically optimize the loss function in the YOLOv8 network, focusing on the matching of the predicted bounding box and the ground truth bounding box in terms of shape and scale. Based on the IoU, it directly optimizes the aspect ratio of the bounding box, making the shapes of the predicted bounding box and the ground truth bounding box more similar, and effectively improving the fitting degree of the predicted bounding box to the target shape.

[0134] (4) The present invention introduces a Ghost module in the convolutional layer. This module reduces the number of conventional convolutions through simple linear transformations, thereby achieving lightweight improvement of the network, minimizing the computational amount as much as possible, and alleviating the requirements of network detection training for the performance of the computing platform.

[0135] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An intelligent selection method for high-precision lightweight focusing windows based on improved YOLOv8, characterized in that, Including: Obtain a number of single - and double - layer wafer mark images, preprocess the single - and double - layer wafer mark images to obtain a preprocessed image; Construct an initial improved YOLOv8 model, and train the initial improved YOLOv8 model according to the preprocessed image to obtain an optimized improved YOLOv8 model; Obtain a target single - and double - layer wafer mark image, input it into the optimized improved YOLOv8 model to complete the intelligent selection of a high - precision lightweight focusing window for the target single - and double - layer wafer mark image.

2. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 1, wherein Constructing the initial improved YOLOv8 model includes: Introduce a deformable convolution module in the backbone network and the feature fusion layer of the YOLOv8 network, and extract the features of the wafer mark by dynamically adjusting the offset of the convolution kernel sampling points; Add a hybrid - domain attention mechanism module between the backbone network and the feature fusion layer to perform channel and spatial attention weighting on the input feature map; Use the Shape - IoU loss function to replace the traditional loss function to calculate the shape and scale matching between the predicted bounding box and the ground - truth bounding box; Introduce a Ghost module in the convolutional layer, generate an original feature layer and a supplementary feature layer through linear transformation, and stack them to form an output feature map.

3. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 2, characterized in that, The introduction positions of the deformable convolution module include the last C2f module of the backbone network and the last layer of the feature fusion layer, and it is replaced with the deformable convolution module only in the convolutional layers at the introduction positions.

4. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 2, characterized in that, The implementation process of the hybrid - domain attention mechanism module includes: Perform channel attention weighting on the input feature map to generate a channel weight matrix; Perform spatial attention weighting on the input feature map to generate a spatial weight matrix; Multiply the channel weight matrix and the spatial weight matrix element - by - element to obtain a final feature weight; Multiply the final feature weight and the original feature map element - by - element.

5. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 2, characterized in that, The calculation of the Shape - IoU loss function includes the matching of two parts: area ratio and shape ratio: Among them, scale is the scaling factor, w t , h t are the length and width of the ground truth box respectively, w p , h p are the length and width of the predicted box respectively, x t , y t are the center coordinates of the ground truth box, x p , y p are the center coordinates of the predicted box, I w , I h are the weight coefficients in the horizontal and vertical directions of the ground truth box respectively, Ω shape is the shape weight, L IoU is the traditional loss function, L Shape-IoU is the Shape-IoU loss function.

6. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 2, characterized in that, The construction of the Ghost module includes: Generate an original feature layer through a conventional convolution operation; Perform a depth - wise separable convolution operation on the original feature layer to generate a supplementary feature layer; Stack the original feature layer and the supplementary feature layer along the channel dimension.

7. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 4, characterized in that, The channel attention weighting is realized through the joint output of global average pooling and max - pooling, and the spatial attention weighting is realized through the splicing and convolution operation of the pooled feature map.

8. The high-precision lightweight focusing window intelligent selection method based on the improved YOLOv8 according to claim 6, wherein The depth - wise separable convolution operation includes depth - wise convolution and point - wise convolution, where the convolution kernel size of the depth - wise convolution is 3×3, and the convolution kernel size of the point - wise convolution is 1×1.

Citation Information

Cited By

  • Locomotive cab personnel leaving detection method based on YOLOV8s and related equipment

    CN121170766A