A lightweight bolt detection method based on improved YOLOv11

CN122821077APending Publication Date: 2026-09-25WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610709980.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有骨干网络的长距离依赖建模能力有限,全局上下文信息不足,导致螺栓目标易被复杂背景淹没

Benefits of technology

[0040](1)通过在骨干网络前端引入自适应光照尺度协同特征增强模块,对输入特征图进行光照补偿与通道筛选,使后续网络层在光照校准后的特征上执行提取与融合。该模块改善了强光与低照度条件下的早期特征质量,降低了光照突变对螺栓识别稳定性的影响,使后续条带池化与动态上采样模块能够在更稳定的特征分布上工作。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821077A_ABST
    Figure CN122821077A_ABST
Patent Text Reader

Abstract

The application discloses a kind of light weight bolt detection methods based on improved YOLOv11, with YOLOv11 as basic framework to carry out six core improvements: inserting adaptive light scale collaborative feature enhancement module in front-end of backbone network, improve the stability of feature under complex illumination;The C3k2 module of backbone network is replaced with light weight C3k2_Dual module, reduce model parameter quantity and calculation amount;After each C3k2_Dual module, add strip pooling module, enhance global long distance feature perception;The up-sampling layer of neck network is replaced with dynamic up-sampling module, optimize small target feature recovery ability;Add foreground and background decoupling gate fusion module at the neck feature fusion place, inhibit steel structure background interference;After detection head, add detection after verification module based on bolt array geometry prior, eliminate false detection and compensate for missing detection.The application realizes model light weight while ensuring detection accuracy, effectively improves the real-time performance, accuracy and environmental adaptability of bolt detection in substation aerial work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a lightweight bolt inspection method based on an improved YOLOv11. Background Technology

[0002] In high-altitude operations on steel structures in substations, bolts are critical fasteners, and their installation quality directly affects the safety of equipment operation. Because bolts are typically small in size, numerous, and densely distributed, and the working environment includes strong light, low illumination, and interference from complex steel structures, traditional manual inspection methods are inefficient and dangerous, making it difficult to meet the demands of large-scale, high-frequency inspections.

[0003] Deep learning-based target detection methods offer a feasible path for automated bolt inspection. The YOLO series of algorithms, due to their end-to-end detection architecture and superior inference speed, have been widely applied in industrial target inspection tasks. However, when dealing with high-altitude operations on steel structures in substations, existing YOLO-based detection methods still face the following technical challenges:

[0004] First, the lighting conditions at substation sites are complex and variable, with alternating periods of strong light and low illumination. This causes distortion in the early feature extraction stage of the image, making it difficult for subsequent network layers to recover effective information from the degraded features, which directly affects the stability of bolt target recognition.

[0005] Second, background components such as steel structure supports, beams, and insulators are highly similar to the bolt targets in grayscale features, and their spatial distribution exhibits strong structural interference. Existing backbone networks have limited long-range dependency modeling capabilities and insufficient global contextual information, causing bolt targets to be easily obscured by complex backgrounds. Furthermore, during cross-layer feature fusion in the neck network, interference features from the steel structure background diffuse towards the detection head during the upsampling process, further deteriorating the detection environment for small targets.

[0006] Third, bolts are typical small targets, and after multiple downsampling operations in deep networks, spatial detail information is severely lost. Traditional upsampling methods use fixed interpolation kernels, which fail to adaptively adjust the sampling position according to the feature distribution, resulting in insufficient edge localization accuracy for small targets and making it difficult to meet the accuracy requirements of installation quality inspection.

[0007] Fourth, the existing YOLO network model has a large number of parameters and computational load. When directly deployed on edge computing devices or mobile platforms, it suffers from high inference latency and large memory consumption, making it difficult to meet the real-time detection requirements in high-altitude operation scenarios.

[0008] Fifth, bolts on steel structures are usually distributed in collinear or equally spaced arrays according to design specifications. The post-processing of existing detection methods only relies on non-maximum suppression and confidence threshold screening, without utilizing the spatial geometric laws of bolt distribution, which leads to isolated false detections or missed detections in complex backgrounds.

[0009] Therefore, how to reduce model complexity while ensuring detection accuracy, and effectively cope with the challenges of harsh lighting, complex backgrounds and small target positioning, has become an urgent technical problem to be solved in the inspection of bolts in substation steel structures. Summary of the Invention

[0010] The present invention provides a lightweight bolt inspection method based on the improved YOLOv11 to solve the problems existing in the prior art.

[0011] The technical solutions adopted in this invention are as follows:

[0012] A lightweight bolt inspection method based on an improved YOLOv11 includes the following steps:

[0013] Acquire and preprocess image datasets of substation steel structure bolts under different lighting conditions;

[0014] An improved detection model based on YOLOv11 was established and trained using a preprocessed dataset. The improved detection model is based on the YOLOv11 network model, which includes a backbone network, a neck network, and a detection head. The improvements are as follows:

[0015] (1) An adaptive illumination scale collaborative feature enhancement module is inserted after the first convolutional layer of the backbone network and before the first C3k2 module;

[0016] (2) Replace all C3k2 modules in the backbone network with C3k2_Dual modules, wherein the C3k2_Dual modules are constructed by replacing standard convolutions with lightweight dual-branch convolutions;

[0017] (3) A strip pooling module is inserted after the last output convolutional layer of each C3k2_Dual module in the backbone network. The strip pooling module located at the end of the backbone network is connected to the spatial pyramid pooling module.

[0018] (4) Replace all upsampling layers in the neck network with dynamic upsampling modules;

[0019] (5) At each lateral connection point in the neck network, before the features output by the dynamic upsampling module are fused with the corresponding layer features of the backbone network, a foreground-background decoupling gating fusion module is inserted.

[0020] (6) After the candidate box is output by the detection head and before non-maximum suppression, a post-detection verification module based on bolt array geometric prior is added;

[0021] The images of the substation steel structure bolts to be detected are input into the improved detection model after training. The improved detection model is processed sequentially through the backbone network, neck network, detection head, and post-detection verification module to output the position coordinates, category, and confidence score of the bolt candidate boxes.

[0022] Furthermore, the adaptive illumination scale collaborative feature enhancement module includes:

[0023] The illumination factor estimation branch is used to perform global average pooling on the input feature map and generate illumination modulation vector and bias vector through two 1×1 convolution layers, respectively.

[0024] Scale response gating is used to generate channel weights through a channel attention mechanism;

[0025] The collaborative modulation unit is used to perform residual modulation output on the input feature map according to the illumination modulation vector, bias vector and channel weight.

[0026] Furthermore, the lightweight dual-branch convolution divides the input channels into two groups in a 1:1 ratio. The first group is processed in parallel by a 3×3 depthwise convolution and a 1×1 pointwise convolution, and then the two groups are added together. The second group is processed only by a 1×1 pointwise convolution. Finally, the two groups of outputs are spliced ​​together in the channel dimension.

[0027] Furthermore, the strip pooling module processes the input tensor through parallel horizontal strip pooling and vertical strip pooling to obtain horizontal and vertical features respectively. The horizontal and vertical features are then fused and weights are generated by 1×1 convolution and the Sigmoid function, and multiplied with the input tensor for output.

[0028] Furthermore, the dynamic upsampling module generates an offset through a linear layer, superimposes the offset onto the initial sampling grid to generate a dynamic sampling point set, and uses a grid sampling function to resample the input feature map based on the dynamic sampling point set to achieve dynamic upsampling.

[0029] Furthermore, the foreground / background decoupling gating fusion module includes:

[0030] Background suppression mask generation unit, used to generate background suppression mask;

[0031] A local feature purification unit is used to perform background suppression and purification on the corresponding layer features of the backbone network based on the background suppression mask.

[0032] An adaptive gating fusion unit is used to perform weighted fusion of the features output by the dynamic upsampling module and the corresponding layer features of the purified backbone network using the learned channel gating weights.

[0033] Furthermore, the background suppression mask generation unit generates a spatial attention mask based on the global context features output by the strip pooling module, using a 1×1 convolution and a sigmoid function.

[0034] The local feature purification unit multiplies the corresponding hierarchical features of the backbone network with the inverse mask obtained by subtracting the spatial attention mask from 1 to obtain the purified features.

[0035] The adaptive gating fusion unit generates channel gating weights through a two-layer fully connected network, and performs weighted fusion of the features output by the dynamic upsampling module and the purified features.

[0036] Furthermore, the detection and verification module based on the prior geometry of the bolt array uses a random sampling consensus algorithm to fit the center point of the candidate box output by the detection head to obtain the bolt distribution geometric model. It calculates the vertical distance from the center point of each candidate box to the bolt distribution geometric model and determines the candidate box with the vertical distance exceeding the adaptive threshold as a false detection and removes it.

[0037] Furthermore, the adaptive threshold is determined based on the bolt diameter and the image scale.

[0038] Furthermore, the detection and verification module based on the prior geometry of the bolt array performs a spacing consistency verification on the gaps between adjacent inspected bolts on the bolt distribution geometry model. If the spacing conforms to the bolt installation specifications and there are candidate boxes with confidence levels lower than a preset threshold in the corresponding area, the confidence level of the candidate box is increased.

[0039] The present invention has the following beneficial effects:

[0040] (1) By introducing an adaptive illumination scale collaborative feature enhancement module at the front end of the backbone network, illumination compensation and channel filtering are performed on the input feature map, so that subsequent network layers can perform extraction and fusion on the features after illumination calibration. This module improves the early feature quality under strong light and low illumination conditions, reduces the impact of sudden illumination changes on the stability of bolt recognition, and enables the subsequent strip pooling and dynamic upsampling modules to work on a more stable feature distribution.

[0041] (2) By replacing the C3k2 module in the backbone network with the C3k2_Dual module composed of lightweight dual-branch convolution, and using the parallel grouping processing of depthwise convolution and pointwise convolution to replace the traditional standard convolution, the number of parameters and computation of the convolutional layer are reduced while ensuring the continuity of cross-channel information, making the improved model more suitable for deployment on edge devices with limited computing resources.

[0042] (3) By inserting a strip pooling module after each C3k2_Dual module in the backbone network, strip pooling in the horizontal and vertical directions is used to capture the long-distance spatial dependencies of steel structure components and supplement global context information. This module enhances the network's ability to understand the spatial layout of the steel structure and improves the feature discrimination between bolt targets and background components.

[0043] (4) By replacing the upsampling layer in the neck network with a dynamic upsampling module, the upsampling grid position is dynamically adjusted by adaptively learning the sampling point offset using the input feature distribution, replacing the traditional fixed interpolation method. This module improves the spatial detail recovery capability of small targets in the feature map reconstruction stage, which helps to improve the positioning accuracy of bolt targets.

[0044] (5) By introducing a foreground-background decoupling gating fusion module at the lateral connection of the neck network, a background suppression mask is generated based on global context features to perform background purification on the corresponding hierarchical features from the backbone network, and the fusion ratio of deep features and shallow features is controlled by adaptive gating weights. This module suppresses the propagation of steel structure background noise to the detection head and reduces the impact of background interference on candidate box generation.

[0045] (6) By introducing a post-detection verification module based on the geometric prior of the bolt array after the detection head, the geometric model of the bolt distribution is fitted using a random sampling consensus algorithm. Isolated false detections are eliminated based on the spatial consistency deviation between the center point of the candidate box and the fitted model, and the confidence of low-confidence candidate boxes in empty areas that meet the installation specification spacing is increased. This verification module, as a zero-parameter post-processing step, assists in verifying the detection results using the spatial rules of the bolt array without increasing the model inference burden, which helps to reduce the probability of isolated false detections and improve the situation of missed detections. Attached Figure Description

[0046] Figure 1 This is a flowchart of the present invention.

[0047] Figure 2 This is a structural diagram of the adaptive illumination scale collaborative feature enhancement module.

[0048] Figure 3 This is a diagram of the C3k2_Dual module and its lightweight dual-branch convolutional structure.

[0049] Figure 4 This is a structural diagram of the strip pooling module.

[0050] Figure 5 This is a structural diagram of the dynamic upsampling module.

[0051] Figure 6 Structure diagram of the foreground / background decoupled gating fusion module. Detailed Implementation

[0052] The invention will now be further described with reference to the accompanying drawings.

[0053] This embodiment provides a lightweight bolt inspection method based on the improved YOLOv11, including the following steps:

[0054] S1: Dataset acquisition and preprocessing.

[0055] A dataset of bolt images from high-altitude operations on substation steel structures was acquired. This dataset includes bolt images under strong light, low light, and normal lighting conditions. The dataset was labeled with information including the bolt category, location coordinates, and bounding box dimensions. Preprocessing was performed on the dataset, including image size normalization, pixel value normalization, and data augmentation. The preprocessed dataset was then divided into training and validation sets.

[0056] S2: Establish an improved detection model.

[0057] An improved detection model based on YOLOv11 was established. This improved model is based on the YOLOv11 network model, which includes a backbone network, a neck network, and a detection head. The improvements of the improved detection model are as follows:

[0058] (1) Adaptive illumination scale collaborative feature enhancement module.

[0059] An adaptive illumination scale collaborative feature enhancement module is inserted after the first convolutional layer of the backbone network and before the first C3k2 module.

[0060] The adaptive illumination scale-coordinated feature enhancement module includes an illumination factor estimation branch, a scale response gating unit, and a cooperative modulation unit, wherein:

[0061] The illumination factor estimation branch performs global average pooling on the input feature map and then processes it through two 1×1 convolution layers to generate an illumination modulation vector and a bias vector, respectively. Specifically:

[0062] Let the input feature map be ,in For the number of channels, For height, For width, global average pooling is used to... Compressed into channel statistical vectors Channel statistical vectors The input is a first 1×1 convolution, which is processed by an activation function and then input into a second 1×1 convolution. The output is an illumination modulation vector. and bias vector .

[0063] Scale response gating generates channel weights through a channel attention mechanism, specifically: for channel statistical vectors... The channel weights are obtained by sequentially passing through a fully connected layer, an activation function, and another fully connected layer, and then mapped using the Sigmoid function. Channel weights are used to characterize the importance of each channel in identifying bolt targets.

[0064] The co-modulation unit performs residual modulation on the input feature map based on the illumination modulation vector, bias vector, and channel weights. Specifically, the co-modulation unit performs the following operations:

[0065] ,

[0066] ,

[0067] in, This indicates that the channel dimension is broadcast multiplied. This is the output feature map of the adaptive illumination scale collaborative feature enhancement module.

[0068] (2) C3k2_Dual module and lightweight dual-branch convolution.

[0069] All C3k2 modules in the backbone network are replaced with C3k2_Dual modules, which are constructed by replacing standard convolutions with lightweight bi-branch convolutions.

[0070] The lightweight two-branch convolution divides the input channels into two groups in a 1:1 ratio. Let the number of input feature channels be... The first group includes the previous One channel, the second group includes the following One channel.

[0071] The first set of features is processed in parallel using 3×3 depthwise convolutions and 1×1 pointwise convolutions, then added and fused. Specifically, the first set of features is first processed by 3×3 depthwise convolutions to extract spatial features, and then by 1×1 pointwise convolutions to adjust the channel representation. The outputs of the two branches are then added element-wise at the same spatial resolution. The second set of features is processed only by 1×1 pointwise convolutions to maintain the continuity of cross-channel information and reduce computational complexity. Finally, the fused result of the first set of features is concatenated with the pointwise convolution result of the second set along the channel dimension, resulting in an output channel count of [number missing]. The feature map.

[0072] The C3k2_Dual module uses the aforementioned lightweight dual-branch convolution as the basic convolutional unit. It is reorganized according to the residual connections and bottleneck structure of the original YOLOv11 C3k2 module, while maintaining the original module's output dimension and connection relationships unchanged.

[0073] (3) Strip pooling module.

[0074] A strip pooling module is inserted after the last output convolutional layer of each C3k2_Dual module in the backbone network. The strip pooling module located at the end of the backbone network is connected to the spatial pyramid pooling module.

[0075] The strip pooling module processes the input tensor through parallel horizontal and vertical strip pooling.

[0076] Horizontal strip pooling performs global average pooling on the input tensor along the horizontal direction, compressing the width dimension to obtain horizontal features. .

[0077] Vertical strip pooling performs global average pooling on the input tensor along the vertical direction, compressing the height dimension to obtain vertical features. .

[0078] The horizontal and vertical features are fused; specifically, the horizontal and vertical features are fused. Copy and expand in the width direction to ,Will Copy and expand in the vertical direction to The two expanded features are added element by element to obtain the fused features. .

[0079] Fusion features The number of channels is adjusted using 1×1 convolution, and then a weight map is generated using the Sigmoid function. .

[0080] Weight graph The enhanced feature map is output by multiplying the input tensor element by element.

[0081] (4) Dynamic upsampling module.

[0082] Replace all upsampling layers in the neck network with dynamic upsampling modules.

[0083] The dynamic upsampling module generates offsets through a linear layer, specifically:

[0084] Let the input feature map be The target upsampling factor is ,Will After a 1×1 linear convolutional layer, the output offset field is obtained. The offset field contains the two-dimensional offset of each target sampling point in the horizontal and vertical directions.

[0085] offset The points are superimposed onto the initial sampling grid to generate a dynamic sampling point set. The initial sampling grid... This is a pre-generated, regularized two-dimensional coordinate grid based on the spatial resolution of the input feature map, the target upsampling ratio, and bilinear interpolation rules. It represents the default sampling position of the upsampled target pixel on the original feature map. Dynamic sampling point set. By offset With the initial sampling grid It is obtained by adding each element together.

[0086] Using the grid sampling function based on the dynamic sampling point set The input feature map is resampled. The grid sampling function uses bilinear interpolation to calculate the feature values ​​of the sampling points, consistent with the interpolation method of YOLOv11's native upsampling, and outputs the upsampled feature map. This enables dynamic upsampling.

[0087] (5) Foreground and background decoupling gating fusion module.

[0088] At each lateral connection point in the neck network, before fusing the features output by the dynamic upsampling module with the corresponding layer features of the backbone network, a foreground / background decoupling gating fusion module is inserted.

[0089] The foreground-background decoupling gating fusion module includes a background suppression mask generation unit, a local feature purification unit, and an adaptive gating fusion unit.

[0090] The background suppression mask generation unit generates a spatial attention mask based on the global context features output by the strip pooling module, using a 1×1 convolution and a sigmoid function. Specifically, at each lateral connection, the global context features output by the strip pooling module are used. After dimensionality reduction via 1×1 convolution, a spatial attention mask is obtained by mapping using the Sigmoid function. The range of the spatial attention mask is The higher the value, the higher the probability that the area belongs to a steel structure background.

[0091] The local feature purification unit multiplies the corresponding layer features of the backbone network with the inverse mask obtained by subtracting the spatial attention mask from 1 to obtain the purified features. Specifically, the inverse mask is calculated. The backbone network corresponds to the hierarchical features With inverse mask Element-by-element multiplication yields the purified characteristics. .

[0092] The adaptive gating fusion unit generates channel gating weights through a two-layer fully connected network, and performs weighted fusion of the features output by the dynamic upsampling module and the purified features. Specifically, the features output by the dynamic upsampling module... Characteristics after purification After concatenation along the channel dimension, global average pooling is used to obtain the channel statistical vector. Then, two fully connected layers and a Softmax function are applied to generate two sets of channel gating weights. .

[0093] The output of the adaptive gating fusion unit is: ,in, The fused feature map is then fed into subsequent layers of the neck network for processing.

[0094] (6) Detection and verification module based on bolt array geometric prior.

[0095] After the candidate box is output by the detection head and before non-maximum suppression, a post-detection verification module based on bolt array geometric prior is added.

[0096] S3: Model training.

[0097] The improved detection model is trained using a preprocessed training set, with the training epochs, batch size, and learning rate set. During training, the detection head outputs candidate box category predictions, location regressions, and confidence predictions. The total loss is calculated using classification loss functions, regression loss functions, and confidence loss functions. The network parameters are updated using the backpropagation algorithm until the model converges.

[0098] S4: Bolt detection reasoning.

[0099] The images of the substation steel structure bolts to be detected are input into the improved detection model after training. The model is then processed sequentially through the backbone network, the neck network, the detection head, and the post-detection verification module based on the geometric prior of the bolt array. The model outputs the position coordinates, category, and confidence level of the bolt candidate boxes.

[0100] Specifically, the detection head processes the fused feature map output by the neck network to generate a set of candidate boxes. Each candidate box Including center point coordinates Width, height, and category confidence And category prediction.

[0101] The detection-and-verification module based on bolt array geometric prior processes the candidate box set:

[0102] The random sampling consensus algorithm is used to fit the center points of the candidate bounding boxes output by the detection head to obtain the geometric model of the bolt distribution. From the set of candidate bounding box center points... The smallest subset is randomly selected from the candidate geometric models. A straight line model or a regular grid model is fitted as the candidate geometric model. The vertical distance from all center points to the candidate model is calculated, and the number of interior points is counted. The above process is iteratively executed, and the candidate model with the most interior points is selected as the final bolt distribution geometric model. The number of iterations is adaptively adjusted according to the number of candidate boxes, and the default setting is 100 times.

[0103] Calculate the vertical distance from the center point of each candidate box to the bolt distribution geometry model. The adaptive threshold is determined based on the bolt diameter and image scale. Specifically, let the average pixel diameter of the bolt in the image be... Image scaling factor is Then the adaptive threshold Determine by the following formula:

[0104] ,

[0105] in This is a preset scaling factor, ranging from 1.5 to 3.0. The vertical distance... Exceeding the adaptive threshold The candidate boxes were identified as false positives and removed.

[0106] The detection and verification module based on the geometric prior of the bolt array performs a spacing consistency verification on the gaps between adjacent detected bolts on the bolt distribution geometric model.

[0107] The spacing between the center points of adjacent inspected bolt candidate frames is calculated along the bolt distribution geometric model. If the spacing conforms to the bolt installation specifications and there are candidate frames with a confidence level lower than a preset threshold in the corresponding area, the confidence level of the candidate frame is increased by 0.2, and the increased confidence level does not exceed 0.9. The preset threshold is determined based on the confidence level distribution statistics on the validation set, and its value ranges from 0.3 to 0.5.

[0108] After the candidate box set is processed by the detection and verification module, overlapping boxes are removed by non-maximum suppression, and the final bolt candidate box position coordinates, category and confidence score are output.

[0109] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be considered within the scope of protection of the present invention.

Claims

1. A lightweight bolt inspection method based on an improved YOLOv11, characterized in that: Includes the following steps: Acquire and preprocess image datasets of substation steel structure bolts under different lighting conditions; An improved detection model based on YOLOv11 was established and trained using a preprocessed dataset. The improved detection model is based on the YOLOv11 network model, which includes a backbone network, a neck network, and a detection head. The improvements are as follows: (1) An adaptive illumination scale collaborative feature enhancement module is inserted after the first convolutional layer of the backbone network and before the first C3k2 module; (2) Replace all C3k2 modules in the backbone network with C3k2_Dual modules, wherein the C3k2_Dual modules are constructed by replacing standard convolutions with lightweight dual-branch convolutions; (3) A strip pooling module is inserted after the last output convolutional layer of each C3k2_Dual module in the backbone network. The strip pooling module located at the end of the backbone network is connected to the spatial pyramid pooling module. (4) Replace all upsampling layers in the neck network with dynamic upsampling modules; (5) At each lateral connection point in the neck network, before the features output by the dynamic upsampling module are fused with the corresponding layer features of the backbone network, a foreground-background decoupling gating fusion module is inserted. (6) After the candidate box is output by the detection head and before non-maximum suppression, a post-detection verification module based on bolt array geometric prior is added; The images of the substation steel structure bolts to be detected are input into the improved detection model after training. The improved detection model is processed sequentially through the backbone network, neck network, detection head, and post-detection verification module to output the position coordinates, category, and confidence score of the bolt candidate boxes.

2. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 1, characterized in that: The adaptive illumination scale collaborative feature enhancement module includes: The illumination factor estimation branch is used to perform global average pooling on the input feature map and generate illumination modulation vector and bias vector through two 1×1 convolution layers, respectively. Scale response gating is used to generate channel weights through a channel attention mechanism; The collaborative modulation unit is used to perform residual modulation output on the input feature map according to the illumination modulation vector, bias vector and channel weight.

3. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 1, characterized in that: The lightweight dual-branch convolution divides the input channels into two groups in a 1:1 ratio. The first group is processed in parallel by a 3×3 depthwise convolution and a 1×1 pointwise convolution, and then the two groups are added together. The second group is processed only by a 1×1 pointwise convolution. Finally, the two groups of outputs are concatenated in the channel dimension.

4. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 1, characterized in that: The strip pooling module processes the input tensor through parallel horizontal and vertical strip pooling to obtain horizontal and vertical features respectively. The horizontal and vertical features are then fused and weights are generated by 1×1 convolution and the Sigmoid function, and multiplied with the input tensor for output.

5. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 1, characterized in that: The dynamic upsampling module generates an offset through a linear layer, superimposes the offset onto the initial sampling grid to generate a dynamic sampling point set, and uses the grid sampling function to resample the input feature map based on the dynamic sampling point set to achieve dynamic upsampling.

6. The lightweight bolt inspection method based on improved YOLOv11 as described in claim 1, characterized in that: The foreground / background decoupling gating fusion module includes: Background suppression mask generation unit, used to generate background suppression mask; A local feature purification unit is used to perform background suppression and purification on the corresponding layer features of the backbone network based on the background suppression mask. An adaptive gating fusion unit is used to perform weighted fusion of the features output by the dynamic upsampling module and the corresponding layer features of the purified backbone network using the learned channel gating weights.

7. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 6, characterized in that: The background suppression mask generation unit generates a spatial attention mask based on the global context features output by the strip pooling module, using a 1×1 convolution and a sigmoid function. The local feature purification unit multiplies the corresponding hierarchical features of the backbone network with the inverse mask obtained by subtracting the spatial attention mask from 1 to obtain the purified features. The adaptive gating fusion unit generates channel gating weights through a two-layer fully connected network, and performs weighted fusion of the features output by the dynamic upsampling module and the purified features.

8. The lightweight bolt inspection method based on improved YOLOv11 as described in claim 1, characterized in that: The detection and verification module based on bolt array geometric prior uses a random sampling consensus algorithm to fit the center point of the candidate box output by the detection head to obtain the bolt distribution geometric model. It calculates the vertical distance from the center point of each candidate box to the bolt distribution geometric model and determines the candidate box with a vertical distance exceeding an adaptive threshold as a false detection and removes it.

9. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 8, characterized in that: The adaptive threshold is determined based on the bolt diameter and image scale.

10. The lightweight bolt inspection method based on the improved YOLOv11 as described in claim 8, characterized in that: The detection and verification module based on the geometric prior of the bolt array performs a spacing consistency verification on the gaps between adjacent inspected bolts on the bolt distribution geometric model. If the spacing conforms to the bolt installation specifications and there are candidate boxes with confidence levels lower than a preset threshold in the corresponding area, the confidence level of the candidate box is increased.