Disease intelligent detection method and device for steel structure bolt and medium

By improving the YOLOv8-GDFPN network model, the accuracy and robustness of bolt defect detection in steel structure bridges were enhanced, solving the problems of insufficient extraction of subtle features and external interference in the existing technology, and realizing efficient and accurate defect detection.

CN121280353APending Publication Date: 2026-01-06SHANDONG JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511380884.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing methods for detecting bolts in steel structure bridges suffer from insufficient ability to extract and locate subtle features, making them susceptible to external interference and affecting the accuracy of defect detection models.

Method used

The YOLOv8-GDFPN network model is adopted. The backbone network is improved by introducing the PSConv module, and the neck network is optimized by combining the CSPStage and DySample modules to enhance the spatial feature extraction capability and noise resistance, and a multi-scale feature fusion mechanism is constructed.

Benefits of technology

It significantly improves the accuracy and anti-interference ability of bolt defect detection, maintains high inference efficiency, reduces reliance on manual detection, and increases the degree of automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121280353A_ABST
    Figure CN121280353A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent disease detection method and device for a steel structure bolt and a medium, belongs to the technical field of computer vision, and is used for solving the problems that bolt detection in an existing steel structure bridge has limitations: the bolt fine feature extraction and positioning capability is insufficient, and the detection is easily interfered by external conditions; and the identification accuracy of the disease detection model is influenced. The method comprises the following steps: performing convolution optimization processing of a structured coverage mode under related multi-scale convolution on spatial scale features in a YOLOv8 network model to obtain an improved backbone network; carrying out fusion optimization processing between deep semantic information and shallow spatial information on the neck network, and carrying out sampling optimization processing on detection precision of targets with different scales to obtain an improved neck network; based on the improved backbone network, the improved neck network and the head network, constructing a YOLOv8-GDFPN network model; and obtaining a current bolt disease detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular to an intelligent detection method, device and medium for defects in steel structure bolts. Background Technology

[0002] Steel structures, with their high strength, lightweight, ease of construction, and excellent seismic performance, occupy an important position in modern engineering construction. Bolted connections, as a crucial method of joint connection, are widely used in transportation hub projects such as cross-sea bridges and urban viaducts due to their simplicity of construction, ease of disassembly, and excellent fatigue resistance. However, steel bridges in long-term service are susceptible to bolt corrosion, loosening, and missing bolts due to the combined effects of vehicle loads, structural vibrations, and complex environmental factors (such as salt spray corrosion, freeze-thaw cycles, and acid rain corrosion). These defects not only weaken the bolt preload and connection strength, affecting the overall structural integrity, but can even lead to local component failure and catastrophic accidents such as bridge collapse. Therefore, timely and accurate detection and assessment of the health status of steel structure bolts are crucial for ensuring structural operational safety and extending service life.

[0003] In the field of steel structure bolt inspection, traditional manual inspection methods face problems such as high safety risks from high-altitude operations, low efficiency of visual inspection, and susceptibility to missed detections and misjudgments in complex environments. With the maturity of drone technology, drones equipped with high-resolution imaging devices can flexibly navigate the complex structural spaces of bridges, achieving rapid acquisition of high-definition images of bolts from multiple angles and precise location of their positions and shapes. Meanwhile, deep learning algorithms in the field of artificial intelligence, with their powerful image feature recognition and analysis capabilities, have demonstrated significant application results in industrial quality inspection, fault diagnosis, and other fields. Therefore, combining image data acquired by drones with deep learning algorithms to construct image-based intelligent defect detection technology has become a new research direction in the field of civil engineering inspection.

[0004] However, bolt defects in steel structure bridges are characterized by small size, diverse shapes, and susceptibility to environmental interference. The limitations of existing models in feature extraction and noise resistance restrict their application in practical engineering. In bolt defect detection scenarios, using the YOLOv8 network model still has some limitations: firstly, bolts are mostly small targets, and the original model lacks the ability to extract and locate subtle features; secondly, outdoor shooting environments are easily affected by changes in lighting, background interference, and other noise, requiring improvement in the model's anti-interference performance. Summary of the Invention

[0005] This application provides an intelligent detection method, device, and medium for defects in steel structure bolts, addressing the following technical problems: Existing bolt detection methods for steel structure bridges have limitations, including insufficient ability to extract and locate subtle bolt features, susceptibility to external interference, and reduced accuracy of defect detection models. For example, the insufficient ability to extract and locate subtle bolt features and susceptibility to external interference affect the accuracy of defect detection models.

[0006] The embodiments of this application adopt the following technical solutions:

[0007] On one hand, this application provides an intelligent detection method for bolt defects in steel structures, including: labeling and classifying pre-collected bolt defect images with relevant defect features and performing image surface preprocessing to obtain a bolt defect detection image set; performing multi-scale feature coverage processing on the spatial features in the YOLOv8 network model under multi-scale convolution with multi-dilation rate to obtain an improved backbone network based on the YOLOv8 network model; and performing fusion optimization processing on the neck network in the YOLOv8 network model with relevant deep semantic information and shallow spatial information through a preset CSPStage module and a dynamic upsampling module, and further optimizing the process. The detection accuracy of targets at different scales in the YOLOv8 network model is sampled and optimized to obtain an improved neck network based on the YOLOv8 network model. A YOLOv8-GDFPN network model is constructed based on the improved backbone network, the improved neck network, and the head network. The YOLOv8-GDFPN network model is trained using the bolt defect detection image set to obtain the trained YOLOv8-GDFPN network model. The current bolt defect detection image set is then input into the YOLOv8-GDFPN network model to obtain the current bolt defect detection result.

[0008] This application improves the backbone network by introducing the PSConv module (multi-scale convolution module), enhancing its spatial feature extraction capabilities. It also optimizes the neck network by combining the CSPStage and DySample modules (dynamic upsampling modules), improving the efficiency of multi-scale feature fusion through an improved feature fusion mechanism. This significantly enhances the detection accuracy and noise resistance of the YOLOv8-GDFPN network model. Furthermore, the YOLOv8-GDFPN network model shows a substantial improvement in accuracy for bolt defect detection, maintaining high inference efficiency while ensuring detection accuracy compared to the original YOLOv8 network model. It can better integrate deep semantic information and shallow spatial information, providing richer features and contributing to improved detection accuracy. The optimized YOLOv8-GDFPN network model can process image data faster, resulting in a faster detection response time. Moreover, the YOLOv8-GDFPN network model is more robust and resistant to interference from image noise and illumination changes. Simultaneously, it reduces reliance on manual detection, increases the automation level of the detection process, and lowers labor costs.

[0009] In one feasible implementation, the pre-collected bolt defect images are annotated and classified according to relevant defect features, including: using a drone's shooting system to collect key areas related to bolt connections in each span of the main bridge steel truss and the key structural areas of the approach bridges of a steel structure bridge, obtaining an original bolt area image set; wherein, the key areas include at least: bolt threads and nut mating surfaces; using an image annotation tool, the bolt targets in the original bolt area image set are annotated with relevant defects under rectangular boxes, obtaining an initial bolt defect annotated image set; wherein, the defects include: corrosion features, loosening features, missing features, and no defects features; and the preset typical bolt defect sample dataset is fused with the defects features in the initial bolt defect annotated image set according to relevant feature distributions to achieve a balanced distribution of the defects features, thereby generating the bolt defect annotated image set.

[0010] In one feasible implementation, the pre-acquired bolt defect images are labeled and classified according to relevant defect features and the image surface is preprocessed to obtain a bolt defect detection image set. Specifically, this includes: dividing the bolt defect labeled images in the bolt defect labeled image set into blocks to obtain several image blocks; performing histogram equalization processing on the several image blocks; limiting the range of pixel grayscale values ​​in the several image blocks and optimizing the contrast of local image regions within the range limitation to obtain a bolt defect preprocessed image set; performing geometric transformation processing on the bolt defect preprocessed image set under the offline preprocessing stage to obtain a bolt defect offline preprocessed image set; wherein the offline preprocessing stage includes: random flipping, rotation transformation, and adaptive cropping; performing pixel-level transformation processing on the bolt defect preprocessed image set under the online enhancement stage to obtain a bolt defect online enhanced image set; wherein the online enhancement stage includes: dynamically and randomly adjusting brightness, contrast transformation, and saturation transformation; and generating the bolt defect detection image set based on the bolt defect offline preprocessed image set and the bolt defect online enhanced image set.

[0011] In one feasible implementation, the spatial features in the YOLOv8 network model are subjected to multi-scale feature coverage processing with multi-dilation rates under multi-scale convolution to obtain an improved backbone network based on the YOLOv8 network model. Specifically, this includes: integrating periodic dilation rates into different convolution kernels of a single convolutional filter based on the convolutional module in the YOLOv8 network model, and performing fine-grained to coarse-grained structured coverage processing on the feature scale space through a single convolutional layer to obtain a multi-scale convolution operation; wherein, a single element of the output feature map in the multi-scale convolution operation is: performing convolution operations with several filters of defined sizes with the input features respectively, and obtaining an element of the output feature map with multiple channels based on the dilation rate between the output channels and input channels in the input features; replacing the original convolution operation in the YOLOv8 network model with the multi-scale convolution operation, and performing convolution optimization on the original backbone network in the YOLOv8 network model to obtain the improved backbone network.

[0012] In one feasible implementation, a preset CSPStage module is used to perform fusion optimization processing on the neck network of the YOLOv8 network model, involving the fusion of deep semantic information and shallow spatial information. Specifically, this includes: dividing the input feature map of the YOLOv8 network model into two parallel branches to obtain a first branch network and a second branch network; using the first branch network, performing 1×1 convolution on the input features in the input feature map for channel compression processing to output a first feature; using the second branch network, performing 1×1 convolution on the input features, and then fusing the phased features based on the RepConv module and 3×3 convolution operation to obtain a second feature; using a preset Concat network, merging the first feature and the second feature, and then performing 1×1 convolution operation to output the final fused feature; wherein, the final fused feature represents the fusion feature between deep semantic information and shallow spatial information under the characteristics of subtle defects and overall spatial distribution.

[0013] In one feasible implementation, a dynamic upsampling module is used to optimize the detection accuracy of targets at different scales in the YOLOv8 network model, resulting in an improved neck network based on the YOLOv8 network model. Specifically, this includes: replacing the target detection-related upsampling module in the YOLOv8 network model with the dynamic upsampling module; wherein the upsampling module includes nearest-neighbor interpolation sampling or bilinear interpolation sampling; generating sampling points for the input feature map using a sampling point generator in the dynamic upsampling module, obtaining a sampling set; resampling the input feature map using a grid sampling function in the dynamic upsampling module, based on the position information of the sampling set, to generate a new sampled input feature map; and optimizing the original neck network in the YOLOv8 network model based on the improved new sampled input feature map generated by the dynamic upsampling module and the final fusion feature generated by the CSPStage module, to obtain the improved neck network.

[0014] In one feasible implementation, the sampling point generator in the dynamic upsampling module performs sampling point generation processing on the input feature map related to the upsampling scale factor to obtain a sampling set. Specifically, this includes: performing linear layer processing on the input feature map through the input and output channels of the sampling point generator to generate an offset at a first preset size; wherein, the linear layer processing includes: linear transformation and generating a dynamic range factor through a Sigmoid activation function; after obtaining the offset, the first preset size is resized by pixel rewashing to obtain a second preset size; the second preset size is added to the original sampling network to generate the sampling set.

[0015] In one feasible implementation, the bolt defect detection image set is used to analyze the bolt defects.

[0016] The YOLOv8-GDFPN network model is trained by learning data features to obtain the trained YOLOv8-GDFPN network model. Specifically, this includes: uniformly adjusting the pixel size of the bolt defect detection image set; iteratively training the YOLOv8-GDFPN network model based on the bolt defect detection image set according to the preset stochastic gradient descent optimizer, initial learning rate, and weight decay coefficient, to obtain the trained YOLOv8-GDFPN network model.

[0017] Secondly, embodiments of this application also provide an intelligent detection device for defects in steel structure bolts, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, enabling the at least one processor to execute an intelligent detection method for defects in steel structure bolts as described in any of the above embodiments.

[0018] Thirdly, embodiments of this application also provide a non-volatile computer storage medium, which is a non-volatile computer-readable storage medium storing at least one program. Each program includes instructions, which, when executed by a terminal, cause the terminal to execute the intelligent detection method for defects in steel structure bolts described in any of the above embodiments.

[0019] This application provides an intelligent detection method, equipment, and medium for defects in steel structure bolts. Compared with the prior art, the embodiments of this application have the following beneficial technical effects:

[0020] 1. Unmanned aerial vehicles (UAVs) were used to collect bolt images under various scenarios, lighting conditions, and weather conditions. After image preprocessing, key point annotation, and data augmentation, a bolt defect detection dataset covering four categories of bolt conditions—corrosion, loosening, missing bolts, and standard bolts—was constructed. This dataset features comprehensive defect coverage and strong anti-interference capabilities, meeting the needs of model training and detection.

[0021] 2. By introducing the PSConv module (multi-scale convolution module), the backbone network is improved, enhancing the spatial feature extraction capability; the neck network is optimized by combining the CSPStage and DySample modules (dynamic upsampling modules), and the multi-scale feature fusion efficiency is improved by improving the feature fusion mechanism, significantly enhancing the model's detection accuracy and noise resistance.

[0022] 3. The YOLOv8-GDFPN network model significantly improves the accuracy of bolt defect detection. Compared with the original YOLOv8 network model, it can maintain high inference efficiency while improving detection accuracy. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0024] Figure 1 A flowchart of an intelligent detection method for defects in steel structure bolts provided in this application embodiment;

[0025] Figure 2 This application provides a schematic diagram of the internal process of multi-scale convolution.

[0026] Figure 3 A schematic diagram of the network structure of a CSPStage module provided in an embodiment of this application;

[0027] Figure 4 A schematic diagram of the network structure of a dynamic upsampling module provided in an embodiment of this application;

[0028] Figure 5 This is a schematic diagram of the structure of a sampling point generator provided in an embodiment of this application;

[0029] Figure 6 A schematic diagram of the network structure of a YOLOv8-GDFPN network model provided in this application embodiment;

[0030] Figure 7This is a structural schematic diagram of an intelligent detection device for defects in steel structure bolts provided in an embodiment of this application. Detailed Implementation

[0031] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0032] It should be noted that bolt defects in steel structure bridges are characterized by small size, diverse morphology, and susceptibility to environmental interference. The shortcomings of existing models in feature extraction and noise resistance limit their application in practical engineering. Therefore, this application proposes a YOLOv8-GDFPN model based on multi-scale convolution and dynamic upsampling mechanisms. By introducing multi-dilation rate convolution kernels to enhance spatial feature extraction capabilities, and combining this with a dynamic upsampling strategy to optimize the network feature fusion structure, the model's performance in bolt defect feature extraction under small target detection and complex background interference is strengthened, thus improving model robustness. This application first utilizes UAVs to collect images of the bolt areas of steel structure bridges from multiple angles and in multiple scenes, obtaining bolt defect sample images under different lighting and weather conditions. The LabelImg tool is used to perform fine-grained annotation of typical defects such as bolt corrosion, loosening, and missing bolts in the sample images, constructing a bolt defect detection dataset containing features from multiple scenes. Based on this dataset, the YOLOv8-GDFPN model is trained and tested. Finally, the experimental results show that the improved model significantly improves the accuracy and precision of bolt defect detection, fully verifying its effectiveness and reliability in practical detection applications.

[0033] Among them, the CSPStage module is a deep learning component based on the improved CSP (Cross Stage Partial) architecture, mainly used to optimize feature fusion efficiency.

[0034] This application provides an intelligent detection method for defects in steel structure bolts, such as... Figure 1 As shown, the intelligent detection method for defects in steel structure bolts specifically includes steps S101-S106:

[0035] S101. The pre-collected bolt defect images are labeled and classified according to relevant defect features and the image surface is preprocessed to obtain a bolt defect detection image set.

[0036] Specifically, the first step involves using a drone's imaging system to capture images of key areas related to bolt connections in the main steel truss sections and approach bridges of the steel structure bridge, resulting in an original set of bolt area images. These key areas include at least the bolt threads and the nut mating surfaces.

[0037] In one embodiment, the flight path during bolt image acquisition can be constructed using the DJI Terra platform and intelligent obstacle avoidance system. This involves planning a spiral-shaped, hovering shooting path covering all sections of the main bridge steel truss and key structural areas of the approach bridge, ensuring comprehensive coverage of bolt connections. During flight, the drone integrates multi-source data from visual, ultrasonic, and laser sensors to perceive the surrounding environment in real time and automatically avoid complex obstacles such as bridge railings and braces. A close-up shooting mode is activated for key nodes to acquire high-definition images of important areas such as bolt threads and nut mating surfaces.

[0038] Furthermore, using image annotation tools, the bolt targets in the original bolt area image set are annotated with relevant defects under rectangular boxes to obtain an initial bolt defect annotated image set. These defects include: rust features, loosening features, missing features, and no-defect features.

[0039] In one embodiment, typical defects of steel structure bolts mainly include three categories: corrosion, loosening, and missing bolts. To construct training samples for the model, the open-source image annotation tool LabelImg was selected to annotate bolt targets in the dataset images with rectangular boxes, categorized as rusty, loose, missing, and normal. Considering the imbalanced distribution of these four types of samples in actual bridge images, typical defect samples were supplemented by fusing the COCO dataset during the annotation stage to improve the coverage and diversity of the annotated data.

[0040] Furthermore, the pre-set typical bolt disease sample dataset and the disease features in the initial bolt disease labeled image set are fused under the relevant feature distribution to achieve a balanced distribution of disease features and generate a bolt disease labeled image set.

[0041] Furthermore, the bolt defect labeled images in the bolt defect labeled image set are divided into blocks to obtain several image blocks; and histogram equalization processing is performed on several image blocks.

[0042] Furthermore, the range of pixel grayscale values ​​in several small image blocks is restricted, and the contrast of local image regions within the restricted range is optimized to obtain a preprocessed image set of bolt defects.

[0043] As a feasible implementation method, the bolts of in-service steel bridges are exposed to complex outdoor environments for extended periods, leading to issues such as surface stains, blurred textures, and shadow occlusion in images. Image preprocessing techniques can significantly improve the visual quality of image data through noise suppression, detail enhancement, and illumination correction, laying the foundation for subsequent feature extraction. Contrast Limited Adaptive Histogram Equalization (CLAHE) is an improved image preprocessing algorithm. Compared to the traditional Histogram Equalization (HE) algorithm, this algorithm divides the image into several small blocks, performs histogram equalization on each block, and then limits the grayscale value range of each pixel in the block, optimizing contrast in local image regions. This effectively suppresses noise amplification while enhancing image contrast.

[0044] Furthermore, the bolt defect preprocessed image set is subjected to geometric transformations during the offline preprocessing stage to obtain the bolt defect offline preprocessed image set. The offline preprocessing stage includes: random flipping, rotation transformation, and adaptive cropping.

[0045] Furthermore, the preprocessed bolt defect image set is subjected to pixel-level transformation processing in the online enhancement stage to obtain the online enhanced bolt defect image set. The online enhancement stage includes: dynamically and randomly adjusting brightness, contrast transformation, and saturation transformation.

[0046] As a feasible implementation method, high-quality datasets are fundamental to ensuring the generalization ability of models in the fields of machine learning and computer vision. Data augmentation techniques, by performing a series of transformations on the original samples, can effectively expand the variation space of data samples, thereby enhancing the robustness and generalization ability of models to complex scenes. This application adopts a staged data augmentation strategy to expand the scale of the dataset and balance the labeled categories: in the offline preprocessing stage, geometric transformations such as random flipping, rotation transformation, and adaptive cropping are performed on the dataset, which can achieve a significant increase in the number of samples in a short time; in the online augmentation stage, pixel-level transformations such as brightness, contrast, and saturation are dynamically and randomly adjusted to further enhance the data diversity during training.

[0047] Furthermore, based on the offline preprocessed image set of bolt defects and the online enhanced image set of bolt defects, a bolt defect detection image set is generated.

[0048] In one embodiment, in the target detection task, data augmentation needs to achieve coordinated adjustment of image transformation and bounding box coordinates. When the image undergoes spatial geometric transformations such as rotation or distortion, the pixel coordinates of the bounding box will undergo non-linear mapping. In this case, it is necessary to transform the coordinates of the four vertices of the bounding box using a perspective transformation matrix, and then recalculate the minimum bounding rectangle based on the pixel positions of the transformed vertices to determine the position, size, and morphological parameters of the new bounding box. Finally, after image preprocessing, key point annotation, and staged data augmentation, a bolt defect detection image set is generated.

[0049] S102. Perform multi-scale feature coverage processing on the spatial features in the YOLOv8 network model under multi-scale convolution with multi-dilation rate to obtain an improved backbone network based on the YOLOv8 network model.

[0050] It should be noted that the YOLOv8 model is an object detection model. This model adopts the classic "backbone-neck-head" architecture and supports three major tasks: object detection, instance segmentation, and image classification. Its network structure is as follows: Figure 5 As shown. Compared to earlier versions of the YOLO series models, the main improvements of the YOLOv8 model are as follows: The backbone network adopts an improved CSPDarknet structure, replacing the C3 module in YOLOv5 and other versions with a C2f module, achieving lightweighting by reducing redundant computation and improving gradient propagation efficiency; the neck network optimizes the PAN-FPN structure, replacing redundant convolutional layers with C2f modules to enhance feature fusion efficiency; the head network adopts an anchor-free design and decouples the head structure, combined with the DFL+CIoU loss function, significantly improving detection accuracy and speed.

[0051] However, in bolt defect detection scenarios, this model has some limitations: firstly, bolts are mostly small targets, and the original model is insufficient in extracting and locating subtle features; secondly, outdoor shooting environments are easily affected by changes in lighting, background interference, and other noise, and the model's anti-interference performance needs to be improved, while the backbone network determines the model's detection accuracy. Therefore, this application improves the backbone network of YOLOv8 to meet the needs of bolt defect detection.

[0052] Specifically, firstly, based on the convolutional module in the YOLOv8 network model, the periodic dilation rate is integrated into different convolutional kernels of a single convolutional filter. Then, through a single convolutional layer, the feature scale space is subjected to structured coverage processing from fine-grained to coarse-grained, resulting in a multi-scale convolution operation. In this multi-scale convolution operation, a single element of the output feature map is obtained by performing convolution operations between several filters of defined sizes and the input features, and based on the dilation rate between the output and input channels in the input features, an element of the output feature map with multiple channels is obtained.

[0053] Furthermore, multi-scale convolution operations are replaced with the original convolution operations in the YOLOv8 network model, and the original backbone network in the YOLOv8 network model is optimized by convolution to obtain an improved backbone network.

[0054] As a feasible implementation method, Figure 2 This application provides a schematic diagram of the internal process of multi-scale convolution, as illustrated in the embodiments of this application. Figure 2 As shown in the figure, different kernel colors represent different dilation rates. The backbone network plays the role of the core feature extractor in computer vision tasks, and its feature representation ability directly determines the model's adaptability to complex scenes. While the convolutional modules used in the traditional YOLOv8 model emphasize scale feature extraction, they are limited by fixed receptive fields and local connectivity, resulting in shortcomings in handling small targets and spatial details. Especially when the bolt target has a large scale variation or is occluded, it is prone to decreased positioning accuracy or missed detections. To address these bottlenecks, this application introduces Poly-Scale Convolution (PSConv) to replace the original convolutional operation, thereby enhancing the model's ability to capture spatial detail information. The core principle of this method is to integrate the periodic dilation rate into different convolutional kernels of a single convolutional filter, enabling a single convolutional layer to cover the feature scale space from fine-grained to coarse-grained in a structured manner, improving the model's robustness to scale changes without increasing parameters or computational complexity.

[0055] In one embodiment, such as Figure 2 As shown, for a single convolutional layer, the input features are F∈R Cin×H×W C in Let H be the number of channels, and H and W be the height and width of the image, respectively. A set of channels has a quantity of C. out A filter of size K×K ∈ R Cout ×Cin×K×K Convolution operations are performed with the input feature F respectively to obtain the feature C. out The output feature map H∈R of each channel Cout×H'×W' The PSConv operation can be represented as:

[0056]

[0057] In the formula: H c,x,y D represents a single element in the output feature map H. (c,k) The expansion rate between the corresponding output channel c and the input channel k.

[0058] S103. Using the preset CSPStage module and dynamic upsampling module, the neck network in the YOLOv8 network model is subjected to fusion optimization processing between deep semantic information and shallow spatial information, and the detection accuracy of targets at different scales in the YOLOv8 network model is sampled and optimized to obtain an improved neck network based on the YOLOv8 network model.

[0059] It should be noted that the neck network, as the feature scheduling center for the target detection task, is responsible for multi-scale feature fusion and enhancement. In the bolt defect detection task, optimizing the neck network structure can effectively integrate semantic information and positional details from different levels in the Feature Pyramid Network (FPN), balance the representation capabilities of local bolt defect features and overall structural features, and improve robustness to complex conditions such as occlusion interference and size differences.

[0060] Specifically, the input feature map in the YOLOv8 network model needs to be divided into two parallel branches to obtain the first branch network and the second branch network.

[0061] Furthermore, the input features in the input feature map are subjected to channel compression processing through 1×1 convolutions via the first branch network, and the first feature is output.

[0062] Furthermore, the input features are subjected to a 1×1 convolution operation through the second branch network, and then the stage features under the 3×3 convolution operation are fused based on the RepConv module to obtain the second feature.

[0063] Furthermore, the first and second features are merged using a pre-defined Concat network, and the final fused feature is output through a 1×1 convolution operation. This final fused feature represents the fusion of deep semantic information and shallow spatial information beneath the subtle defects and overall spatial distribution features.

[0064] As a feasible implementation method, Figure 3 A schematic diagram of the network structure of a CSPStage module provided in an embodiment of this application is shown below. Figure 3 As shown, to address the problem of insufficient ability of existing models to fuse subtle defects and overall spatial distribution features, this application introduces the CSPStage module to replace the C2f module, thereby achieving efficient fusion of deep semantic information and shallow spatial information. The structure of the CSPStage module is as follows: Figure 3As shown, the CSPStage module divides the input feature map into two parallel branches. One branch compresses the channel through a 1×1 convolution, while the other branch performs a 1×1 convolution operation followed by a RepConv module and a 3×3 convolution operation for staged feature fusion. Finally, the features from the two branches are merged by Concat and then output as the final fused features through a 1×1 convolution operation.

[0065] Among them, such as Figure 3 As shown, the CSPStage module, through its asymmetric branching design, enables the network to learn features at different levels on different branches. This helps the network capture multi-scale information of the image and improves the model's expressive power. Simultaneously, by removing the channel splitting operation of the C2f structure, it improves feature utilization efficiency while reducing computational resource requirements.

[0066] Furthermore, the dynamic upsampling module replaces the upsampling module for object detection in the YOLOv8 network model. This upsampling module includes either nearest-neighbor interpolation or bilinear interpolation.

[0067] Furthermore, the sampling point generator in the dynamic upsampling module performs sampling point generation processing on the input feature map with respect to the upsampling scaling factor to obtain the sampling set.

[0068] As a feasible implementation, the input feature map is processed using linear layers through the input and output channels of the sampling point generator to generate an offset at a first preset size. This linear layer processing includes linear transformation and generation of a dynamic range factor using a sigmoid activation function. After obtaining the offset, the first preset size is resized by pixel reshuffling to obtain a second preset size. The second preset size is then added to the original sampling network to generate the sampling set.

[0069] Furthermore, the input feature map is resampled using the grid sampling function in the dynamic upsampling module and based on the location information of the sampling set to generate a new sampled input feature map.

[0070] As a feasible implementation method, Figure 4 This is a schematic diagram of the network structure of a dynamic upsampling module provided in an embodiment of this application. The network structure of the DySample module is as follows: Figure 4As shown, in object detection tasks, upsampling enhances the model's detection accuracy for targets at different scales by improving the details of low-resolution feature maps, and is a core step in achieving multi-scale feature fusion. The nearest neighbor interpolation or bilinear interpolation methods used in the YOLOv8 model are both upsampling methods based on fixed rules. These methods require determining the number of pixels in the output feature map based on the number of nearest or adjacent pixels in the input feature map, often leading to blurred details, jagged edges, or loss of key information, making it difficult to capture small target features. To address these issues, this application introduces a dynamic upsampling (DySample) module to replace the original interpolation sampling methods.

[0071] In one embodiment, Figure 5 This is a schematic diagram of a sampling point generator provided in an embodiment of this application, as shown below. Figure 4 as well as Figure 5 As shown, given an upsampling scaling factor s and an input feature map F of size C×H×W, this feature map is first processed by a sampling point generator to generate a sampling set S of size 2×sH×sW. Subsequently, the grid sampling function (Gridsample) resamples the input feature map F using the positions in the sampling set S, generating an upsampled feature map F' of size C×sH×sW, as shown below:

[0072] F' = grid_sample(F,S). In the sampling point generator structure, the input feature map F has C input channels and 2s output channels. 2 Linear layer processing generates a size of 2s 2 The offset O is calculated as ×H×W. To enhance the flexibility of the offset, a dynamic range factor is generated through an additional linear transformation and a sigmoid activation function to adjust the range of values ​​for the offset. The formula for calculating the offset is as follows:

[0073] O = 0.5sigmoid(linear1(x))·linear2(x); After the offset O is generated, its size is reshaped to 2s×H×W through pixel shuffle, and added to the original sampling grid G ​​to obtain the sampling set S, that is, S = O + G.

[0074] Furthermore, based on the improved newly generated sampled input feature map from the dynamic upsampling module, and

[0075] The final fusion features generated by the CSPStage module are used to optimize the original neck network in the YOLOv8 network model, resulting in an improved neck network.

[0076] S104. Based on the improved backbone network, improved neck network, and head network, a system is constructed.

[0077] YOLOv8-GDFPN network model.

[0078] Specifically, Figure 6 A schematic diagram of the network structure of a YOLOv8-GDFPN network model provided in this application embodiment is shown below. Figure 6 As shown, based on the improved backbone network and improved neck network, and combined with the head network, the YOLOv8-GDFPN network model was obtained. The improved backbone network (Backbone) changed from C2f stacking to alternating PSConv and C2f to achieve multi-scale feature interaction. The improved neck network (Neck) sampling method changed from Upsample to DYSample dynamic upsampling, and the CSPStage module was used to extract features to improve information fusion quality. The head network (Head) still consists of three Decets corresponding to feature layers of different scales.

[0079] S105. Using the bolt defect detection image set, the YOLOv8-GDFPN network model is trained by learning the data features to obtain the trained YOLOv8-GDFPN network model.

[0080] Specifically, the pixel size of the bolt defect detection image set is uniformly adjusted. Then, based on the preset stochastic gradient descent optimizer, initial learning rate, and weight decay coefficient, the YOLOv8-GDFPN network model is iteratively trained to obtain the trained YOLOv8-GDFPN network model.

[0081] In one embodiment, the YOLOv8-GDFPN network model can be trained to learn data features based on the official default parameter configuration of YOLOv8. Parameters are kept consistent throughout the training, testing, and validation of all models, and all experiments are performed in the same hardware and software environment to ensure the reliability and reproducibility of the experimental results. The experimental environment configuration is shown in Table 1. During the training phase, a stochastic gradient descent (SGD) optimizer is used, with an initial learning rate of 0.01 and a weight decay coefficient of 0.0005. Input images are uniformly adjusted to a size of 640×640 pixels, the total number of model training iterations is 600, and the batch size is set to 32 based on hardware performance.

[0082] Table 1 Experimental Environment Configuration

[0083]

[0084] In one embodiment, in the field of machine learning, the algorithm's prediction results for samples can be divided into four categories: TP (True Positives), FP (False Positives), FN (False Negatives), and TN (True Negatives). TP refers to the number of samples that are actually positive and correctly predicted as positive by the model; FP is the number of samples that are actually negative but misclassified as positive by the model; FN represents samples that are actually positive but incorrectly predicted as negative by the model; and TN represents samples that are actually negative and correctly predicted as negative by the model. To comprehensively evaluate model performance, this application uses accuracy (Accuracy, Acc), precision (Precision, P), recall (Recall, R), mean average precision (mAP), F1 score, and frame rate (FPS) as evaluation metrics, specifically defined as follows:

[0085] (1) Accuracy (Accuracy, Acc)

[0086] Accuracy reflects the overall correctness of the model's predictions, and refers to the proportion of all correctly predicted samples in the total sample. It is calculated as follows:

[0087] (2) Precision (P)

[0088] Precision measures the proportion of samples that the model predicts to be positive but are actually positive. It reflects the accuracy of the model's predictions of positive classes; that is, how many samples predicted as positive are actually positive. The calculation method is as follows:

[0089]

[0090] (3) Recall (R)

[0091] Recall measures the proportion of actual positive samples that are correctly predicted, and is used to evaluate the model's ability to capture positive samples. In other words, it measures how many actual positive samples are successfully identified by the model. Its calculation method is as follows:

[0092]

[0093] (4) F1 value

[0094] The F1 score comprehensively considers precision and recall, balancing them through a harmonic average to avoid the influence of extreme values ​​in a single metric, thus providing a more complete reflection of model performance. Its calculation method is as follows:

[0095]

[0096] (5) Mean accuracy (mAP)

[0097] mAP is used for tasks such as multi-class object detection. It measures the average accuracy of the model across different classes by calculating the average accuracy for each class and then averaging it across all classes. This provides a comprehensive evaluation of the model's performance on complex tasks. The calculation method is as follows:

[0098]

[0099] (6) Frame Rate (FPS)

[0100] FPS is a metric for measuring a system's real-time rendering capability, reflecting the number of consecutive image frames generated per unit of time. A higher frame rate indicates clearer dynamic details and directly determines visual smoothness. Its calculation method is as follows:

[0101] In the formula: t pre For preprocessing time, t inf Let t be the inference time. post This refers to post-processing time.

[0102] In one embodiment, after constructing the YOLOv8-GDFPN model, model training was performed based on the experimental environment shown in Table 1. During training, both the bounding box loss (train / box_loss) and classification loss (train / cls_loss) showed a stable decreasing trend, indicating that the model could effectively learn data features. Although the bounding box loss (val / box_loss) and classification loss (val / cls_loss) on the validation set fluctuated slightly, they showed an overall decreasing trend, indicating that the model had good generalization ability. In terms of performance metrics, the model precision (metrics / precision(B)) and recall (metrics / recall(B)) gradually increased and tended to stabilize, remaining above 0.8 and 0.7, respectively. The mean accuracy (metrics / mAP50(B)) at an intersection-union ratio of 50% eventually stabilized at around 0.8, and the mAP50-95 (metrics / mAP50-95(B)) metric performed well at a stricter threshold. In summary, the convergence trends of the training and validation metrics are consistent, indicating that the model training process is stable and the detection performance meets expectations. Finally, after the training process is complete, the system will automatically generate a weight file. The weight paths in the test code must be replaced with the newly generated best weight file, and all other relevant parameter settings must be completely consistent with those in the training phase. After completing the above operations, run the test code, and the system will immediately output the test results. Experimental comparison results show that the improved algorithm significantly outperforms the original YOLOv8 model in the bridge bolt defect detection task. In the same detection task, the model's mAP index reaches 83.5%, effectively overcoming the limitations of the original version, significantly improving the overall detection effect, and enabling more efficient and accurate identification of bridge defects.

[0103] S106. Input the current bolt defect detection image set into the YOLOv8-GDFPN network model to obtain the current bolt defect detection results.

[0104] Specifically, the preprocessed image and its corresponding annotation file are used as input data and loaded into...

[0105] In the YOLOv8-GDFPN network model, ensure that the input data format conforms to the model's requirements, including channel order and image size. Then, use the trained YOLOv8-GDFPN network model to perform inference on the input image. The model will automatically identify bolt defects in the image and output the detected defect location, category, and confidence level; ultimately, it will derive the current bolt defect detection result.

[0106] As a feasible implementation method, to systematically evaluate the contribution of the improved backbone and neck network strategies in the YOLOv8-GDFPN model to the target recognition performance, this application conducted ablation comparison experiments on a self-built dataset. Using the original YOLOv8 model as a baseline, and maintaining complete consistency in the detection head, training strategy, dataset, data augmentation strategy, and all parameters, six detection models were constructed and trained sequentially. These six models reflect the impact of each improvement on the bolt defect detection results. The experimental results show that while using only PSConv convolution operations reduces the number of model parameters, it does not significantly improve the model accuracy; in fact, the mAP metric even decreases. While the model optimized only with CSPStage and DySample achieves improvements in mAP and accuracy, it increases the amount of floating-point operations. The YOLOv8-GDFPN model constructed in this application achieved an mAP of 83.5%, representing improvements of 4.9%, 5.5%, 3.8%, 1.8%, and 1.3% compared to the previous five comparative models, respectively; its F1 score reached 74.0%, representing improvements of 11.4%, 6.0%, 6.2%, 4.3%, and 1.1% compared to the previous five models, respectively. Ablation experiments showed that the improvements to the backbone and neck networks in the YOLOv8-GDFPN model contributed to both a lightweight model and improved accuracy.

[0107] In addition, this application also provides an intelligent detection device for defects in steel structure bolts, such as... Figure 7 As shown, the intelligent detection equipment 700 for defects in steel structure bolts specifically includes:

[0108] At least one processor 701; and a memory 702 communicatively connected to the at least one processor 701; wherein the memory 702 stores instructions executable by the at least one processor 701 to enable the at least one processor 701 to execute:

[0109] The pre-collected bolt defect images are labeled and classified according to relevant defect features and the image surface is preprocessed to obtain a bolt defect detection image set;

[0110] The spatial features in the YOLOv8 network model are subjected to multi-scale feature coverage processing under multi-scale convolution with multi-dilation rate, resulting in an improved backbone network based on the YOLOv8 network model.

[0111] By using the pre-defined CSPStage module and dynamic upsampling module, the neck network in the YOLOv8 network model is subjected to fusion optimization processing between deep semantic information and shallow spatial information. The detection accuracy of targets at different scales in the YOLOv8 network model is also optimized by sampling, resulting in an improved neck network based on the YOLOv8 network model.

[0112] Based on the improved backbone network, improved neck network, and head network, a YOLOv8-GDFPN network model was constructed.

[0113] The YOLOv8-GDFPN network model was trained by learning data features using a set of bolt defect detection images, resulting in the trained YOLOv8-GDFPN network model.

[0114] The current bolt defect detection image set is input into the YOLOv8-GDFPN network model to obtain the current bolt defect detection results.

[0115] This application improves the backbone network by introducing the PSConv module (multi-scale convolution module), enhancing spatial feature extraction capabilities; and optimizes the neck network by combining the CSPStage and DySample modules (dynamic upsampling modules). By improving the feature fusion mechanism, the efficiency of multi-scale feature fusion is enhanced, significantly improving the detection accuracy and noise resistance of the YOLOv8-GDFPN network model. Furthermore, the YOLOv8-GDFPN network model significantly improves the accuracy in bolt defect detection, maintaining high inference efficiency while improving detection accuracy compared to the original YOLOv8 network model. It can better integrate deep semantic information and shallow spatial information, providing richer features and contributing to improved detection accuracy. The optimized YOLOv8-GDFPN network model can process image data faster, resulting in a faster detection response time. Moreover, the YOLOv8-GDFPN network model has stronger resistance to interference from image noise, illumination changes, and other factors, exhibiting better robustness and improving detection reliability. Simultaneously, it reduces reliance on manual detection, increases the automation level of the detection process, and lowers labor costs.

[0116] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0117] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0118] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0119] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0122] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0123] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0124] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0125] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0126] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of this specification.

Claims

1. A method for intelligently detecting diseases of a steel structure bolt, characterized by, The method comprises: The pre-acquired bolt disease image is subjected to annotation classification and image surface preprocessing related to disease characteristics to obtain a bolt disease detection image set; The spatial features in the YOLOv8 network model are subjected to multi-expansion rate multi-scale feature coverage processing under multi-scale convolution to obtain an improved backbone network based on the YOLOv8 network model; Through a pre-set CSPStage module and a dynamic upsampling module, the neck network in the YOLOv8 network model is subjected to fusion optimization processing between deep semantic information and shallow spatial information, and the detection accuracy of different scale targets in the YOLOv8 network model is subjected to sampling optimization processing to obtain an improved neck network based on the YOLOv8 network model; Based on the improved backbone network, the improved neck network and the head network, a YOLOv8-GDFPN network model is constructed; The YOLOv8-GDFPN network model is subjected to learning and training of data features through the bolt disease detection image set to obtain a trained YOLOv8-GDFPN network model; The current bolt disease detection image set is input into the YOLOv8-GDFPN network model to obtain a current bolt disease detection result.

2. The method according to claim 1, wherein The pre-acquired bolt disease image is subjected to annotation classification and image surface preprocessing related to disease characteristics to obtain a bolt disease detection image set, specifically comprising: Through a shooting operation system in the unmanned aerial vehicle, key area image acquisition processing of bolt connection parts is performed on each hole segment of the main bridge steel truss and the key structure area of the approach bridge in the steel structure bridge to obtain an original bolt area image set; wherein the key area at least includes bolt threads and nut engagement surfaces; Through an image annotation tool, disease feature annotation processing of bolt targets in the original bolt area image set is performed under a rectangular frame to obtain an initial bolt disease annotation image set; wherein the disease features include rust features, loosening features, missing features and non-disease features; A pre-set bolt typical disease sample data set is subjected to fusion processing with the disease features in the initial bolt disease annotation image set under feature distribution to make the disease features evenly distributed, and a bolt disease annotation image set is generated.

3. The method according to claim 2, wherein The pre-acquired bolt disease image is subjected to annotation classification and image surface preprocessing related to disease characteristics to obtain a bolt disease detection image set, specifically comprising: The bolt disease annotation images in the bolt disease annotation image set are subjected to block processing to obtain a plurality of image blocks; and the plurality of image blocks are subjected to histogram equalization processing; The pixel gray level value range in the plurality of image blocks is subjected to range limitation, and the image local area in the range limitation is subjected to contrast optimization to obtain a bolt disease pretreatment image set; The bolt disease pretreatment image set is subjected to geometric transformation processing under an offline preprocessing stage to obtain a bolt disease offline pretreatment image set; wherein the offline preprocessing stage includes random flipping, rotation transformation and adaptive cropping; perform pixel-level transformation processing on the bolt disease pretreatment image set in an online enhancement stage to obtain a bolt disease online enhancement image set; wherein the online enhancement stage includes dynamic random adjustment of brightness, contrast transformation, and saturation transformation; generate the bolt disease detection image set based on the bolt disease offline pretreatment image set and the bolt disease online enhancement image set.

4. The intelligent disease detection method of a steel structure bolt according to claim 1, characterized in that, perform multi-expansion-rate multi-scale feature coverage processing on the spatial features in the YOLOv8 network model under multi-scale convolution to obtain an improved backbone network based on the YOLOv8 network model, specifically including: based on the convolution module in the YOLOv8 network model, integrate periodic expansion rates into different convolution kernels of a single convolution filter, and through a single convolution layer, perform structured coverage processing on the feature scale space from fine granularity to coarse granularity to obtain a multi-scale convolution operation; wherein a single element of the output feature map in the multi-scale convolution operation is: performing convolution operations on the input features with several filters of defined sizes, and based on the expansion rate between the output channels and the input channels in the input features, obtaining an element of the output feature map output with multiple channels; replace the original convolution operation in the YOLOv8 network model with the multi-scale convolution operation, and perform convolution optimization on the original backbone network in the YOLOv8 network model to obtain the improved backbone network.

5. The intelligent disease detection method of a steel structure bolt according to claim 1, characterized in that, perform fusion optimization processing between deep semantic information and shallow spatial information on the neck network in the YOLOv8 network model through a preset CSPStage module, specifically including: divide the input feature map in the YOLOv8 network model into two parallel branches to obtain a first branch network and a second branch network; through the first branch network, perform channel compression processing on the input features in the input feature map through 1x1 convolution to output a first feature; through the second branch network, after 1x1 convolution operation on the input features, based on stage feature fusion under RepConv module and 3x3 convolution operation, obtain a second feature; through a preset Concat network, merge the first feature and the second feature, and output a final fusion feature through 1x1 convolution operation; wherein the final fusion feature represents the fusion feature between deep semantic information and shallow spatial information under fine defects and overall spatial distribution features.

6. The intelligent disease detection method of a steel structure bolt according to claim 5, characterized in that, perform sampling optimization processing on the detection accuracy of different scale targets in the YOLOv8 network model through a dynamic upsampling module to obtain an improved neck network based on the YOLOv8 network model, specifically including: replace the upsampling module related to the target detection task in the YOLOv8 network model with the dynamic upsampling module; wherein the upsampling module includes nearest neighbor interpolation sampling or bilinear interpolation sampling; through a sampling point generator in the dynamic upsampling module, perform sampling point generation processing on the input feature map with respect to the upsampling ratio factor to obtain a sampling set; Resample the input feature map through a grid sampling function in the dynamic upsampling module and based on position information of the sampling set to generate a new sampling input feature map; Optimize the original neck network in the YOLOv8 network model based on the new sampling input feature map improved by the dynamic upsampling module and the final fusion feature improved by the CSPStage module to obtain the improved neck network.

7. The intelligent disease detection method of a steel structure bolt according to claim 6, characterized in that, Generate a sampling set by a sampling point generator in the dynamic upsampling module through sampling point generation processing of the input feature map related to an upsampling scale factor, specifically including: Perform linear layer processing on the input feature map through an input channel and an output channel in the sampling point generator to generate an offset under a first preset size; wherein the linear layer processing includes linear transformation and generation of a dynamic range factor through a Sigmoid activation function; After obtaining the offset, perform size reshaping on the first preset size through pixel reflushing to obtain a second preset size; Add the second preset size to the original sampling network to obtain the sampling set.

8. The method of claim 1, wherein the method is characterized by: Learn and train data features of the YOLOv8-GDFPN network model through the bolt disease detection image set to obtain a trained YOLOv8-GDFPN network model, specifically including: Uniformly adjust the pixel size of the bolt disease detection image set; According to a preset random gradient descent optimizer, an initial learning rate, and a weight decay coefficient, and based on the bolt disease detection image set, iteratively train the YOLOv8-GDFPN network model to obtain a trained YOLOv8-GDFPN network model.

9. A steel structure bolt disease intelligent detection equipment, characterized in that, The device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to execute a steel structure bolt disease intelligent detection method according to any one of claims 1-8.

10. A non-transitory computer storage medium, comprising, The storage medium is a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores at least one program, each of which includes instructions that, when executed by a terminal, cause the terminal to execute a steel structure bolt disease intelligent detection method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Automatic detection method suitable for railway steel bridge bolt diseases

    CN114387261A

  • YOLOv8-based bolt looseness detection method, system, equipment, medium and program

    CN119579525A