A steel defect detection model training and application method, equipment and medium

By improving the YOLOv8 model, using residual convolution and depthwise separable convolution to optimize feature extraction, combined with dynamic snake convolution and frequency-adaptive dilation factor, the real-time and adaptability issues in high-speed rail track defect detection are solved, and efficient and accurate defect detection is achieved.

CN120279011BActive Publication Date: 2025-09-30EAST CHINA JIAOTONG UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510756595.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-30
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Existing detection methods have problems with real-time performance, sensitivity, and adaptability in high-speed rail track defect detection. Traditional methods have limited ability to characterize complex defects, while deep learning models have high computational complexity and are difficult to meet the needs of intelligent operation and maintenance of high-speed rail networks.

Method used

By improving the YOLOv8 model, using residual convolution and depth-wise separable convolution to replace the convolution module of the feature extraction backbone network, combined with dynamic snake convolution and frequency-adaptive dilation factor, the detection capability of complex defects is enhanced, and a high-resolution prediction layer and frequency-adaptive dilated convolution are constructed to optimize the model's computational efficiency.

Benefits of technology

It significantly improves the detection accuracy and scenario generalization capability of complex track defects, reduces computational complexity, increases detection speed and adaptability, and meets the intelligent operation and maintenance needs of high-speed rail networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279011B_ABST
    Figure CN120279011B_ABST
Patent Text Reader

Abstract

The present application discloses a steel defect detection model training and application method, device and medium, which relate to the field of image recognition technology. The method comprises: obtaining a data set; taking a sample steel surface image in the data set as input and a defect category corresponding to the sample steel surface image as a label, training a steel defect detection model to obtain a trained steel defect detection model; the steel defect detection model is an improved YOLOv8 model; in the improved YOLOv8 model, the second to fourth convolution modules of the feature extraction backbone network in the YOLOv8 model are replaced with residual convolution modules, and the first convolution module and the fifth convolution module of the feature extraction backbone network in the YOLOv8 model are replaced with depthwise separable convolution modules; the residual convolution module is a convolution module that introduces a residual structure, which can improve the speed of steel defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a steel defect detection model training and application method, equipment and medium. Background Art

[0002] With the rapid expansion of my country's high-speed rail network and the continued growth of operating mileage, accurate detection of defects such as microcracks and tubular scratches on track surfaces has become a key component in ensuring operational safety. However, high-speed rail defects often exhibit complex morphology, small size, and strong background interference. Existing detection methods face significant challenges in real-time performance, sensitivity, and adaptability. Traditional machine learning relies on manually designed features (such as background subtraction and Gabor filtering), has limited ability to characterize complex defects, and is susceptible to uneven lighting and noise interference. While deep learning models can automatically extract features, the high computational complexity of mainstream algorithms (such as the improved YOLO series of models) results in slow steel defect detection, making it difficult to meet the urgent needs of intelligent operation and maintenance of the high-speed rail network. Summary of the Invention

[0003] The purpose of this application is to provide a steel defect detection model training and application method, equipment and medium, which can improve the speed of steel defect detection.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides a steel defect detection model training method, comprising:

[0006] Acquire a data set; the data set includes a plurality of sample steel surface images and a defect category corresponding to each of the sample steel surface images;

[0007] The sample steel surface image is used as input and the defect category corresponding to the sample steel surface image is used as a label to train a steel defect detection model to obtain a trained steel defect detection model; the steel defect detection model is an improved YOLOv8 model; the improved YOLOv8 model replaces the second to fourth convolution modules of the feature extraction backbone network in the YOLOv8 model with residual convolution modules, and replaces the first convolution module and the fifth convolution module of the feature extraction backbone network in the YOLOv8 model with depthwise separable convolution modules; the residual convolution module is a convolution module that introduces a residual structure.

[0008] Optionally, the residual convolution module includes a first residual branch, a second residual branch, a third residual branch, a fusion layer and an activation function layer; the first residual branch includes a first convolution layer and a first normalization processing layer connected in sequence; the second residual branch includes a second convolution layer and a third normalization processing layer connected in sequence; the third residual branch includes a third normalization processing layer; the fusion layer is used to perform weighted fusion on the output of the first residual branch, the output of the second residual branch, and the output of the third residual branch to obtain a weighted fusion feature; the weighted fusion feature is the input of the activation function layer.

[0009] Optionally, the depthwise separable convolution module includes a depthwise convolution unit and a pointwise convolution unit; the depthwise convolution unit includes three groups of first depthwise convolution layers, which respectively extract features from the three color channels of the input feature map of the depthwise separable convolution module to obtain a first intermediate output feature map, a second intermediate output feature map and a third intermediate output feature map; the pointwise convolution unit includes several second depthwise convolution layers, each second depthwise convolution layer performs weighted fusion on the first intermediate output feature map, the second intermediate output feature map and the third intermediate output feature map to obtain a cross-channel fused high-order feature map corresponding to each second depthwise convolution layer.

[0010] Optionally, the improved YOLOv8 model further introduces dynamic snake convolution into the Bottleneck structure in the detection C2f module, and the detection C2f module is a C2f module connected to the detection head.

[0011] Optionally, the improved YOLOv8 model also adds a fusion convolution module after the last C2f module of the Neck network part of the YOLOv8 model, and adds a small defect prediction layer to the detection head module; the fusion convolution module includes a Conv module and a C2f module, the Conv module is used to upsample the output of the last C2f module of the Neck network part of the YOLOv8 model to obtain shallow features; the shallow features and the output of the first C2f module of the feature extraction backbone network are spliced ​​along the channel dimension to obtain a fusion feature map; the C2f module in the fusion convolution module is used to perform cross-scale feature interaction on the fusion feature map to obtain an optimized fusion feature map, and the optimized fusion feature map is the input of the small defect prediction layer.

[0012] Optionally, the improved YOLOv8 model further introduces a frequency adaptive expansion factor in the detection head module of YOLOv8, and the frequency adaptive expansion factor is used to adjust the expansion rate of the convolution kernel.

[0013] Optionally, the steel defect detection model training method further includes:

[0014] Evaluate the performance of the trained steel defect detection model.

[0015] In a second aspect, the present application provides a steel defect detection model application method, comprising:

[0016] Acquire target steel surface image;

[0017] The target steel surface image is input into a trained steel defect detection model to obtain a defect detection result corresponding to the target steel surface image; the trained steel defect detection model is a model trained using the above-mentioned steel defect detection model training method.

[0018] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned steel defect detection model training method or steel defect detection model application method.

[0019] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned steel defect detection model training method or steel defect detection model application method.

[0020] According to the specific embodiments provided in this application, this application discloses the following technical effects: This application provides a steel defect detection model training and application method, device and medium, by replacing the second to fourth convolution modules of the feature extraction backbone network in the YOLOv8 model with residual convolution modules, and replacing the first convolution module and the fifth convolution module of the feature extraction backbone network in the YOLOv8 model with depthwise separable convolution modules, an improved YOLOv8 model is obtained, and using the improved YOLOv8 model as a steel defect detection model for steel defect detection can reduce the computational complexity of the model and improve the detection speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is an application environment diagram of a steel defect detection model training method or a steel defect detection model application method in one embodiment of the present application.

[0023] Figure 2A flow chart of a steel defect detection model training method provided in Example 1 of the present application.

[0024] Figure 3 Schematic diagram of 6 types of original defect images in the data set provided in Example 1 of this application.

[0025] Figure 4 This is a schematic diagram of the structure of the residual convolution module provided in Example 1 of the present application.

[0026] Figure 5 A schematic diagram of the process of improving the feature extraction network using dynamic snake convolution provided in Example 1 of the present application.

[0027] Figure 6 A schematic diagram showing a visual comparison of the defect detection results of various categories using the improved YOLOv8 model provided in Example 1 of the present application and the existing YOLOv8 model.

[0028] Figure 7 A flow chart of a steel defect detection model application method provided in Example 2 of the present application.

[0029] Figure 8 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0031] Traditional detection methods struggle to balance the accuracy of tiny defect localization with real-time detection requirements due to issues such as the limited geometric modeling capabilities of fixed-structure convolution, inefficient multi-scale feature fusion, and high computational redundancy. In this context, there is an urgent need to develop an efficient, robust, and fine-grained defect detection solution for high-speed rail scenarios to address these challenges and meet the urgent needs of intelligent operation and maintenance of the high-speed rail network.

[0032] Existing defect detection technologies suffer from three common flaws: First, there is a mismatch between feature extraction and defect morphology. Traditional methods (such as region growing and morphological segmentation) rely on manually pre-set rules, insufficiently modeling the long-range continuity of tubular cracks, and are prone to missegmentation due to background interference. Related technologies use Gabor filtering to enhance directional features, but filters with fixed frequency and direction struggle to adapt to the multi-scale curvature variations of defects. Second, small target detection performance is limited. While the improved YOLO algorithm based on deep learning (dense network and residual attention mechanism) improves semantic representation through feature reuse, it suffers from severe loss of shallow high-resolution information, making it incapable of capturing the details of millimeter-level cracks. Third, there is a conflict between computational efficiency and accuracy. Existing lightweight designs (such as shallow feature fusion and bidirectional feature pyramids) reduce the number of parameters, but the simplified modules weaken the ability to detect complex defects, making it difficult to balance real-time and robustness requirements.

[0033] To address the above issues, this application proposes a multi-dimensional collaborative optimization algorithm for high-speed rail track defect detection, which improves the YOLOv8 model: simulating the continuous deformation characteristics of tubular defects through dynamic snake convolution, breaking through the geometric constraints of fixed convolution kernels, and enhancing the adaptive tracking capability of crack boundaries; constructing 160 160 high-resolution prediction layers and frequency-adaptive dilated convolutions collaboratively capture the fine-grained texture and multi-scale contextual information of small targets, addressing the problems of shallow feature loss and spectral aliasing. Structural reparameterization and depthwise separable convolutions are used to reconstruct the backbone network, achieving decoupled optimization between feature expression capabilities during the training phase and computational efficiency during the inference phase. Integrating prior knowledge of tubular structures in medical imaging, channel dimensionality reduction and curvature constraint mechanisms are designed to reduce complexity while improving the model's adaptability to typical high-speed rail defects. Through the deep integration of geometric modeling, spectral optimization, and lightweight computing, this solution significantly improves the detection accuracy and scenario generalization capabilities of complex track defects, providing reliable technical support for the safe operation and maintenance of high-speed rail.

[0034] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0035] Example 1.

[0036] The steel defect detection model training method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send a data set to the server 104. After receiving the data set, the server 104 uses the sample steel surface image in the data set as input and the defect category corresponding to the sample steel surface image as a label to train the steel defect detection model to obtain a trained steel defect detection model. The server 104 can feed back the obtained trained steel defect detection model to the terminal 102. In addition, in some embodiments, the steel defect detection model training method can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly perform model training on the data set, or the server 104 can obtain the data set from the data storage system and perform model training on the data set.

[0037] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0038] In an exemplary embodiment, Figure 2 As shown, a steel defect detection model training method is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, including the following steps 201 to 202.

[0039] Step 201 : Acquire a data set; the data set includes a number of sample steel surface images and a defect category corresponding to each of the sample steel surface images.

[0040] Step 202: Using the sample steel surface image as input and the defect category corresponding to the sample steel surface image as a label, a steel defect detection model is trained to obtain a trained steel defect detection model; the steel defect detection model is an improved YOLOv8 model; in the improved YOLOv8 model, the second to fourth convolution modules of the feature extraction backbone network in the YOLOv8 model are replaced with residual convolution modules, and the first convolution module and the fifth convolution module of the feature extraction backbone network in the YOLOv8 model are replaced with depthwise separable convolution modules; the residual convolution module is a convolution module that introduces a residual structure.

[0041] By implementing the above steps 201 to 202, an improved YOLOv8 model is obtained by replacing the second to fourth convolution modules of the feature extraction backbone network in the YOLOv8 model with residual convolution modules, and replacing the first and fifth convolution modules of the feature extraction backbone network in the YOLOv8 model with depthwise separable convolution modules. Using the improved YOLOv8 model as a steel defect detection model for steel defect detection can reduce the computational complexity of the model and improve the detection speed.

[0042] First, we build the dataset for the model. Specifically, we download the steel defect detection dataset NEU-DET from the Internet. This dataset contains 1,800 grayscale images. The original resolution of each image is 200×200 pixels. It contains six defect categories: cracks (Cr), inclusions (In), pits (Ps), patches (Pa), rolled-in scale (Rs), and scratches (Sc). Each defect category has 300 images. The original defect images are as follows: Figure 3 As shown, the original defect image is the sample steel surface image from step 201. First, the dataset's original XML annotation information is converted to txt format. The dataset is then randomly divided into training, test, and validation sets in a 6:4 ratio. The training set contains 1054 images, while the test and validation sets each contain 746 images. This completes dataset construction, preparing for the model training phase in step 306.

[0043] The training, evaluation and validation of the steel defect detection model can be divided into the following steps 301 to 308.

[0044] Step 301: Build the YOLOv8n basic network. The YOLOv8n basic network is the existing YOLOv8 model. The residual structure is introduced to optimize the backbone network topology. Specifically: First, download the YOLOv8n model and complete the deployment of the basic model. Then replace the ordinary convolution modules of the 2nd to 4th convolution modules of the backbone (feature extraction backbone network) with convolution modules that introduce the residual structure. The specific implementation steps are as follows: Figure 4 shown.

[0045] The residual convolution module includes a first residual branch, a second residual branch, a third residual branch, a fusion layer and an activation function layer; the first residual branch includes a first convolution layer and a first normalization processing layer connected in sequence; the second residual branch includes a second convolution layer and a third normalization processing layer connected in sequence; the third residual branch includes a third normalization processing layer; the fusion layer is used to perform weighted fusion on the output of the first residual branch, the output of the second residual branch, and the output of the third residual branch to obtain a weighted fusion feature; the weighted fusion feature is the input of the activation function layer.

[0046] like Figure 4 The figure shows the residual convolution module designed by this application after introducing the residual structure, in which the 3 3 convolution paths are the original convolution module structure. This application adds a 1 on the basis of the original module structure. 1 convolution branch and the original image splicing branch, then the second residual branch is 1 1 convolution branch, the third residual branch is the original image splicing branch. Specifically, the residual convolution module first copies the input feature map f into 3 copies, copy 1, copy 1 and copy 3, and copy 1 is fed into 3 3 The first residual branch of the convolution and obtain the first residual output: f1 = Conv3 3-BN; copy 2 is fed into 1 1 The second residual branch of the convolution and obtain the second residual output: f2 = Conv1 1-BN; copy 3 is retained as the third residual branch of the identity mapping, and the third residual output is obtained: f3 = Identity (f3). Among them, Conv3 3 is the 3 of the middle path (first residual branch) 3 convolutional layers, mainly responsible for extracting local feature information of the input feature map; Conv1 1 is the second residual branch, responsible for controlling the output channels and reducing computational complexity. Identity is the third residual branch, which does no processing and is primarily responsible for preserving the complete information of the original input feature map f to prevent degradation in deep networks. The BN layer is a normalization layer. After the first step, the three branches are batch normalized to obtain intermediate results, namely the first residual output f1, the second residual output f2, and the third residual output f3.

[0047] The first residual output f1, the second residual output f2, and the third residual output f3 of the three branches are weighted fused to obtain the weighted fusion feature, whose formula is: , where W1, W2, W3 are the weights of each residual branch, and The weight of the residual branch By the formula After calculation, the value is obtained according to the proportion. Residual branch The scaling factor learned in the BN layer, Indicates the The standard deviation of the residual branch calculated in the BN layer; =1, 2, 3.

[0048] Standard deviation The calculation process is as follows: For each input feature map (a numerical matrix) fi, (i=1, 2, 3), the BN layer first calculates its mean and variance , the standard deviation is the square root of the variance, ,in is a very small constant used to stabilize the value. The weighted fusion feature F1 is finally activated and a nonlinear factor is introduced to obtain the final output F of the residual convolution module.

[0049] During the model inference phase, all parameters on the second and third residual branches are loaded into the first residual branch. The merged single branch has a higher response speed than the multi-branch structure, and the model detection speed will be greatly improved.

[0050] Step 302: Dynamic serpentine convolution improves the feature extraction network. Specifically, an in-depth analysis is conducted on the difficulty of identifying tubular structure features in high-speed rail track defects. By comparing the morphological characteristics of vascular structures in the field of medical imaging, it is found that tubular defects such as plaques and scratches generated during the steel rolling process have significant similarities in topological structure: they all show long-range continuity, obvious curvature changes, and blurred boundaries. Based on this discovery, the present application introduces dynamic serpentine convolution into the Bottleneck structure of the C2f module of the YOLOv8 model. Specifically, the improved YOLOv8 model also introduces dynamic serpentine convolution into the Bottleneck structure in the detection C2f module. The detection C2f module is a C2f module connected to the detection head.

[0051] Specific steps are as follows Figure 5 As shown, the input image is set to , H, W, and C represent the height, width, and number of channels of the input image, respectively. First, the input image channel is reduced in dimensionality: a 1×1 convolution is used to compress the input channel number to 64 dimensions to reduce the amount of computation.

[0052] (1);

[0053] Where: is the intermediate output result, Represents the convolution kernel used for channel number compression, Represents the parameters of the convolution kernel itself, BN represents the batch normalization operation immediately after each convolution operation, X represents the input image, and finally the input image with C channels is reduced to 64 channels.

[0054] Next, the intermediate result Xmid is input into the 64 groups of 3 The convolution layer consists of 3 convolution kernels. The ordinary convolution kernels in the convolution layer are replaced by dynamic snake convolution to better fit the defects for feature learning.

[0055] For intermediate output results The first step is to generate 9 dynamic sampling point coordinates for each position (p, q) of the feature map. The center point (p, q) is the initial coordinate. , generate the coordinates of the other 8 dynamic sampling points according to the following recursive formula:

[0056] (2);

[0057] Where: and Respectively The horizontal and vertical coordinates of the dynamic sampling points; and Respectively The horizontal and vertical coordinates of the dynamic sampling points; , which means the The learnable parameter of the offset, that is, each dynamic sampling point is allowed to be 1 around the previous dynamic sampling point Swing within the range of 1; Represents the curvature adjustment coefficient, which is used to control the curvature constraint strength. is the curvature constraint function, the curvature constraint function It is used to control the curvature of the convolution kernel sampling points so that it can adaptively fit the tubular structure.

[0058] The mathematical expression of the curvature constraint function is:

[0059] (3);

[0060] Where: Indicates that the input feature map I is at the coordinate point ( , ), which is used to quantify the local curvature; is the Frobenius norm of the Hessian matrix, reflecting the curvature intensity of the point; is an adjustable parameter used to control the sensitivity of the curvature constraint. When it is larger, When it belongs to the first setting value range, Stronger constraints are imposed on high curvature areas (such as sharp turns), forcing the sampling points to fit closely to the structure boundaries. Smaller When the sampling point falls within the second set value range, it is allowed to deviate freely to adapt to the tubular structure in the flat area. The lower limit of the first set value range is greater than or equal to the upper limit of the second set value range. The output is constrained to (0, 1], and the final output result is the constraint weight, which participates in the offset calculation of the coordinates of the above 8 dynamic sampling points.

[0061] The second step is to calculate the coordinates of each dynamic sampling point ( , ) Obtain the actual sampling value through bilinear interpolation. The formula is as follows:

[0062] (4);

[0063] Where, , is the dynamic coordinate, i.e. the four nearest integer grid points around is the adjacent integer coordinate point; the overall process of the formula is based on the floating point coordinate ( , ) Find four adjacent integer points , calculate the bilinear weight of each neighboring point based on the distance , and finally the eigenvalues ​​of the four neighboring points Perform weighted summation to obtain the coordinate point ( , ) interpolation result , It also serves as an intermediate result and participates in the final calculation.

[0064] The third step is to aggregate the dynamic sampling value with the weight matrix, and the process is expressed as:

[0065] (5);

[0066] in, For the The weight matrix of the sampling points is learned by the intermediate process, is the calculation result of the second step, The convolution kernel with the current coordinate as the center point is used to extract the input feature map and obtain the intermediate output. .

[0067] Finally, the intermediate output Use 1 1 Convolution restores the feature dimension to the original number of channels, and its formula is expressed as:

[0068] (6);

[0069] Where, Represents the parameters of the dimensionality-raising convolution kernel itself. The whole process is the inverse process of the dimensionality reduction process. After the above steps, the final output of the module is obtained. Compared with the C2f module before the improvement, the C2f module after the dynamic snake convolution introduced in step 303 can better fit the tubular defects for feature extraction, effectively improving the feature expression ability of the model and thus improving the model detection accuracy.

[0070] The depth-wise separable convolution module includes a depth-wise convolution unit and a point-by-point convolution unit; the depth-wise convolution unit includes three groups of first depth-wise convolution layers, which respectively extract features from the three color channels of the input feature map of the depth-wise separable convolution module to obtain a first intermediate output feature map, a second intermediate output feature map and a third intermediate output feature map; the point-by-point convolution unit includes several second depth-wise convolution layers, each of which performs weighted fusion on the first intermediate output feature map, the second intermediate output feature map and the third intermediate output feature map to obtain a cross-channel fused high-order feature map corresponding to each second depth-wise convolution layer.

[0071] Step 303: Depthwise separable convolution processes the first and fifth convolution modules of the feature extraction backbone network of the YOLOv8 model. Specifically, the number of input channels and output channels of the original standard convolution module structure of the YOLOv8 model are , The original standard convolution module is now split into depth-wise convolution (processing spatial features) and point-wise convolution (processing channel relationships).

[0072] Step 3031: Depth-by-depth convolution unit: Taking the first convolution module of the feature extraction backbone network of the YOLOv8 model as an example, the input it receives is an RGB three-channel image, that is, =3. Let the input image be Y, and use 3 groups (number of groups = ) 3 3 convolution kernels are used to extract features from each color channel C, and each color channel uses an independent 3 The convolution kernel of 3 ensures the number of channels of the intermediate output feature map of this part and Keep it consistent. Then the convolution kernel size of the first depth convolution layer in the depth-wise convolution unit is 3 3.

[0073] Step 3032: Point-by-point convolution unit: The point-by-point convolution part uses 64 groups (number of groups = )1 1 convolution kernel, each group of convolution kernels processes each channel of the 3 channels output in step 3031. Specifically, suppose the outputs of the 3 channels in step 3031 are C1, C2, and C3 respectively, and each group of 1 1's convolution kernels are randomly assigned weights W1, W2, and W3 to C1, C2, and C3, respectively, with 1 for each group. The final output of the convolution kernel of 1 is 64 groups 1 The convolution kernel of 1 finally obtains 64 sets of different high-order feature maps of cross-channel fusion. The convolution kernel size of the second deep convolution layer is 1 1.

[0074] Depthwise convolution is used to split the feature extraction process (step 3031) and the channel number control process (step 3032) in the standard convolution process into two independent steps. 3, the computational complexity is reduced by 60.4% compared to standard convolution. Furthermore, since the first and fifth convolutional modules are located in the non-core feature part of the backbone network, using depthwise separable convolution in these two modules can minimize the adverse impact on model accuracy while significantly reducing model complexity and computational complexity, making the model suitable for deployment on edge devices with limited computing power.

[0075] Step 304: Reconstruct upsampling module and cross-level feature fusion architecture: Specifically: 20, 40 40, 80 Based on the prediction layers of three scales, a new 160 The 160-scale prediction layer is specifically designed for predicting small object defects. Its specific steps are divided into the following three steps.

[0076] The improved YOLOv8 model also adds a fusion convolution module after the last C2f module of the Neck network part of the YOLOv8 model, and adds a small defect prediction layer to the detection head module; the fusion convolution module includes a Conv module and a C2f module, and the Conv module is used to upsample the output of the last C2f module of the Neck network part of the YOLOv8 model to obtain shallow features; the shallow features and the output of the first C2f module of the feature extraction backbone network are spliced ​​along the channel dimension to obtain a fusion feature map; the C2f module in the fusion convolution module is used to perform cross-scale feature interaction on the fusion feature map to obtain an optimized fusion feature map, and the optimized fusion feature map is the input of the small defect prediction layer.

[0077] 1) Add a new set of fused convolution modules after the last set of convolution modules in the backbone of the YOLOv8 model. This set of fused convolution modules consists of a Conv module and a C2f module. The Conv module uses 3 The convolution kernel size is 3, and the step size S is set to 2 to control the resolution of the output feature map to 160 160.

[0078] 2) First, build the upward fusion path. The specific operation is: convert the original 20 20 The deep semantic feature map of 512 is upsampled for the first time to obtain a scale of 40 40 512 feature maps, and then upsample them a second time to obtain a scale of 160 160 512 feature maps, followed by the upsampled deep semantic feature maps (160 160 512) and the shallow features constructed in this application (160 160 256) along the channel dimension to obtain a scale of 160 160 The fused feature map of 768 is obtained by splicing the output of the first C2f module of the feature extraction backbone network and the shallow features along the channel dimension. The fused feature map is subjected to cross-scale feature interaction by the C2f module and then passes through 1 1 convolution compresses the channel dimension and outputs the optimized 160 160 256 features, that is, the optimized fusion feature map. (Same as step 3032 The second step is to build the downward fusion path. The specific operation is: the shallow features (160 160 256) After a downsampling operation, the scale size is 80 80 256 feature maps, and then compare the downsampled feature maps with the original mid-scale feature maps of the Neck network (80 80 256) element-by-element addition, integrating shallow details with mid-level semantics, and outputting the enhanced 80 80 256 feature maps.

[0079] 3) The 160 obtained in the upward fusion path 160 The 256 feature maps are fed into the newly added detection head module (this module is the original module of the YOLOv8 model, which originally had only 3 detection heads. This application uses the detection head that comes with the YOLOv8 model to detect 160 160-scale defects), specially used for the detection of tiny defects.

[0080] This application constructs 160 160 The optimized fusion feature map of 256 scale can not only be used to detect tiny defects (millimeter level), but also to build a downward fusion path to improve the original 80 The defects with 80-degree resolution are supplemented with features to improve the detection accuracy of the model.

[0081] Step 305: Frequency adaptive expansion factor optimization detection head. The improved YOLOv8 model also introduces a frequency adaptive expansion factor in the detection head module of YOLOv8. The frequency adaptive expansion factor is used to adjust the expansion rate of the convolution kernel. Specifically: After the work of the previous steps is completed, the frequency adaptive expansion factor is introduced in the detection head module of the YOLOv8 model. Its implementation is divided into three steps: First, a lightweight frequency domain analysis module is used to perform local spectral decomposition on the input features. For the feature map X with the number of channels C, width and height W and H respectively, 8 8-size window, perform block frequency energy extraction on the feature map and calculate each 8 The DCT coefficient matrix of 8 image blocks has the following formula:

[0082] (7);

[0083] Where, The coordinates of the upper left corner of the window are The DCT coefficient matrix of the image block, represents the local discrete cosine transform, is the coordinate of the upper left corner of the window, and the output on the left side of the equation is 8 8 coefficient matrix.

[0084] Then, the high-frequency component and low-frequency component bands of the generated DCT coefficient matrix are used to quantize the entire frequency band, and the output result is the energy of the corresponding frequency band:

[0085] (8);

[0086] Where: is the quantized high-frequency band energy, is the quantized low-frequency band energy; , Represent high-frequency components and low-frequency components respectively; are the components of the DCT coefficient matrix, are the components of the DCT coefficient matrix.

[0087] Finally, normalization operation is performed on each channel C to generate the spectrum energy distribution matrix :

[0088] (9);

[0089] Where, For coordinates The spectrum energy distribution at .

[0090] The second step is to perform frequency-space gated attention mapping. Specifically: first, 1 1. Convolution converts the spectrum energy distribution matrix with C channels in the step Perform channel dimension compression to generate intermediate results :

[0091] (10);

[0092] Where, Indicates 1 1 is the weight of the convolution kernel; S is the output result of the previous step, that is, the spectrum energy distribution matrix.

[0093] Using intermediate results Generate spatial weights , and its calculation process is:

[0094] (11);

[0095] Where, is the Sigmod function, represents point-by-point multiplication, Perform average pooling operation on the initial input feature map.

[0096] The third step is to use spatial weight Dynamically adjust the expansion rate of the convolution kernel at each position. The specific operation is:

[0097] (12);

[0098] Where, is the initial dilation rate of the convolution kernel, is the preset minimum expansion rate, is the preset maximum expansion rate.

[0099] First, according to the preset maximum expansion rate and the preset minimum expansion rate And the spatial weight calculated in the second step Calculate the initial expansion rate Then the initial expansion rate The rounded expansion rate is obtained by rounding, and the receptive field of the convolution kernel at each position is dynamically adjusted using the rounded expansion rate, and the optimal representation of the multi-scale signal is finally output.

[0100] This method is expected to significantly improve the model's ability to represent multi-scale targets in complex scenes, accurately capture high-frequency texture differences in fine-grained image classification, and enhance the model's robustness to interference such as noise and occlusion through the collaborative optimization of frequency-space domain parameters, thereby achieving a balanced expression of details and semantics.

[0101] Step 306: Model training. Specifically, upload the prepared dataset to the server where the model is deployed, select the detect task, set the initial learning rate to 0.005, the final learning rate to 0.01, use the default SGD optimizer in training, set the optimizer momentum to 0.930, and the number of training epochs to 300. Considering the possible lack of computing power, set the number of images per batch (batch size) to 8. Record the average accuracy of each round during training. Stop training if the accuracy does not improve within 50 rounds. Disable mosaic data augmentation in the last 10 rounds of training to stabilize training.

[0102] The steel defect detection model training method provided in this embodiment further includes: evaluating the performance of the trained steel defect detection model.

[0103] Step 307: Construct an evaluation index system: This application uses the average detection precision (AP) to evaluate the detection ability of the model for each defect category, and the mean average detection precision (mAP) of all categories to evaluate the model performance, where mAP refers to mAP. 0.5. In addition, frames per second (FPS) is used to measure the detection speed of the model. The number of parameters (params) and the model computation time (FLOGS) are used to evaluate the deployment difficulty of the model.

[0104] Average precision (AP) is a commonly used evaluation metric in the field of object detection. It is calculated by precision (P) and recall (R). Precision indicates the proportion of samples predicted by the model to be positive that are actually positive, while recall indicates the proportion of all actual positive samples correctly predicted by the model to be positive.

[0105] (13);

[0106] (14);

[0107] (15);

[0108] in, Represents the number of correctly detected targets, represents the number of falsely detected targets, Represents the number of undetected targets, Indicates the precision value at different recall rates. The mean precision at different recall rates can be obtained by integrating all the precisions with a recall rate of 0 to 1. .

[0109] mAP 0.5 (i.e., mAP used in this application) refers to the mean of the average detection accuracy of all categories when the IOU threshold is 0.5, that is:

[0110] (16);

[0111] in is the number of defect categories detected, Refers to Average detection accuracy for each defect type.

[0112] FPS refers to the number of images detected by the model per second, which can well represent the detection speed of the model. Its calculation formula is:

[0113] (17);

[0114] in, is the total number of sample steel surface images detected, The total time spent on the test.

[0115] Step 308: Experimental verification. This embodiment compares the Faster R-CNN, SSD, YOLOv3, YOLOv5, YOLOv7, YOLOvX, YOLOv8, and YOLOv9-c models with the steel defect detection model proposed in this embodiment. The comparative experimental results are shown in Table 1.

[0116] Table 1 Comparative experimental results of different models

[0117]

[0118] Ours in Table 1 is the improved YOLOv8 model proposed in this application. Experimental analysis shows that the improved YOLOv8 model proposed in this application has achieved significant performance improvement in steel defect detection tasks, with its mAP 0.5 reaches 78.8%, which is 4.4% higher than the 74.4% of the original YOLOv8 model. At the same time, the number of model parameters is reduced by 24.5%, the amount of calculation is reduced by 35.4%, and the detection speed is increased by 2.8%. This breakthrough is mainly due to five technical innovations: First, through the multi-branch residual topology reconstruction of the backbone network, while retaining the original feature information, the complementarity of cross-level features is strengthened, so that the efficiency of texture feature extraction in the defect area is improved; Second, the introduction of the dynamic snake convolution module innovatively solves the problem of capturing the curvature features of tubular defects. Through the learnable sampling point offset mechanism, the recall rate of linear defects such as scratches is significantly improved; Third, the application of deep separable convolution in non-core layers achieves a balance between computational complexity and model accuracy, while maintaining the feature extraction capability of the backbone network. , the floating-point operations of the first and fifth convolution modules are reduced to a certain extent respectively; Fourth, the construction of a 160×160 high-resolution prediction layer combined with a two-way feature fusion strategy significantly improves the detection accuracy of millimeter-level pitting defects. At the same time, through the cross-scale interaction of shallow detail features and deep semantic features, the problem of missed detection of small targets is effectively alleviated; Fifth, the adaptive expansion factor mechanism based on frequency domain energy analysis dynamically adjusts the receptive field of the convolution kernel to reduce the false positive rate of the model under complex background interference, especially in high-frequency texture areas, to achieve more accurate boundary positioning. This embodiment significantly improves the detection accuracy and robustness of multi-scale defects in complex industrial scenarios while maintaining the real-time detection advantages of YOLOv8 through systematic model architecture optimization, providing an innovative solution for high-precision defect detection in edge computing environments.

[0119] like Figure 6 The following shows the visual comparison of the defect detection results of various categories by the improved YOLOv8 model proposed in this application and the original YOLOv8 model. Figure 6 (a) and (b) are the visual comparison results of two defect categories. This solution refers to the use of the improved YOLOv8 model for steel surface defect detection. Figure 6 It can be seen from the figure that the improved YOLOv8 model proposed in this application has achieved a certain degree of accuracy improvement in each defect category, especially the crack category. Figure 6 As can be seen in the figure, the problem of missed detection of some defects has been solved. Overall, the solution proposed in this application can achieve high-precision, high-response real-time detection. At the same time, the model parameters and computational complexity are greatly reduced, and it can also be deployed and run on edge devices with limited computing power.

[0120] Through multi-dimensional collaborative optimization, this application demonstrates significant advantages in detection efficiency, accuracy, and robustness, specifically including the following four points.

[0121] 1) Efficiency Improvement: Lightweight Structure and Computational Decoupling: Step 301 (Optimizing the Backbone Network Topology Using Structural Reparameterization): During training, a multi-branch residual structure is employed to enhance feature representation. During inference, parameter merging is used to simplify multiple branches into a single path, reducing computational redundancy and achieving a balance between accuracy and speed. Step 303 (Reconstructing the Backbone Network's Basic Convolutional Units Using Depthwise Separable Convolution): Traditional convolution is decomposed into depthwise convolution (spatial filtering) and pointwise convolution (channel fusion), significantly reducing computational complexity while preserving cross-channel semantic relevance. Step 304 (Reconstructing the Upsampling Module and Cross-Level Feature Fusion Architecture): 1×1 convolutions are introduced in the dynamic snake convolution (step 302) and FADC (step 305) for channel compression, controlling module complexity and avoiding the computational burden of high-resolution features. While maintaining high accuracy, this algorithm significantly reduces computational resource consumption and improves real-time detection speed, making it suitable for edge device deployment or large-scale industrial inspection scenarios.

[0122] 2) Improved Detection Accuracy: Multi-Scale Feature Adaptive Enhancement: Step 302 (Dynamic Snake Convolution Improved Feature Extraction Network) utilizes the curvature constraint mechanism of the deformable convolution kernel to adaptively track the continuous boundaries and curvature changes of tubular defects, enhancing feature extraction for long-range nonlinear structures and addressing the inadequate response of traditional convolution to blurred edges. Step 304 (Reconstructed Upsampling Module and Cross-Level Feature Fusion Architecture) adds a 160×160 scale prediction layer, combining shallow, fine-grained features with deep semantic features for cross-scale fusion, improving the localization accuracy of small defects (millimeter-level cracks). Step 305 (Frequency-Adaptive Dilated Convolution) dynamically adjusts the dilation rate based on frequency domain energy, focusing on details in high-frequency regions and expanding context in low-frequency regions, achieving balanced local-global feature representation. In complex industrial scenarios (such as high-speed rail tracks), the algorithm significantly improves the detection rate of tubular defects and small flaws, maintaining stable performance in particularly low-contrast conditions and strong noise interference.

[0123] 3) Robustness Enhancement: Geometric Continuity and Spectral Co-Optimization: Step 304 (Reconstructing the Upsampling Module and Cross-Level Feature Fusion Architecture): In the neck region, high-resolution details are integrated with deep semantic information to enhance the algorithm's contextual awareness of multi-scale objects and reduce missed detections due to scale variations. Step 305 (Frequency-Adaptive Dilated Convolution): Through end-to-end joint training of frequency domain analysis and spatial convolution parameters, spectral aliasing and spatial discontinuities are suppressed, improving the algorithm's generalization to illumination variations and background interference. The algorithm maintains stability in complex environments (such as uneven illumination, object deformation, and background clutter), reduces false detections and missed detections, and adapts to the diverse needs of industrial inspection scenarios.

[0124] 4) Domain Adaptability: Customized Design for Industrial Defects: Step 302 (Medical Image Migration of Dynamic Snake Convolution): Drawing on the morphological priors of vascular structure detection, a convolutional deformation strategy is designed for the topological characteristics of industrial tubular defects (such as steel rolling cracks) to achieve domain knowledge-driven feature enhancement.

[0125] Step 304 (Fine-grained optimization of the high-resolution prediction layer in the reconstructed upsampling module and cross-level feature fusion architecture): This algorithm preserves the pixel-level texture of tiny defects through a shallow network, addressing the problem of small object information loss caused by downsampling in traditional detectors. This algorithm is customized and optimized for the specific characteristics of industrial defects (such as long-range continuity and small size), resulting in greater domain adaptability in similar scenarios and reducing performance degradation caused by feature bias.

[0126] High-speed railway track condition monitoring and defect prediction play a key role in ensuring train operation safety, suppressing the multi-physics coupling effects caused by dynamic wheel-rail interaction, and reducing lifecycle maintenance costs. Its implementation directly depends on the stable and reliable service performance of track structural materials. Therefore, developing high-precision and efficient steel surface defect detection methods is a prerequisite for ensuring the quality, safety, and service reliability of rail transit infrastructure. After comprehensively considering factors such as detection speed, accuracy, model deployment difficulty, and algorithm robustness in industrial scenarios, a prediction model, namely a steel defect detection model, is proposed that integrates multi-dimensional feature enhancement integration (MFEI), multi-dimensional collaborative optimization (MCO), and multiple detection heads. Based on the existing YOLOv8 model, this application addresses issues such as insufficient detection accuracy, missed detection of small defects, and difficulty deploying the model on edge devices. First, based on the characteristics of high-speed rail track defects, the upsampling module and cross-level feature fusion architecture were redesigned to construct a new prediction branch containing shallow detail features, effectively improving the module's feature representation ability for various defects. Secondly, to meet the real-time requirements of detection, the structural reparameterization technology was used to optimize the network topology, achieving equivalent fusion of multi-module computing paths while maintaining the model inference speed, significantly improving the model operation efficiency. The dynamic snake convolution (DSC) was innovatively introduced to reconstruct the C2f module, and adaptive topological learning of tubular defects was achieved through flexible deformation convolution kernels to ensure the continuity of feature extraction in key areas. At the same time, considering the computing power resources of edge devices, the basic units of the backbone network were reconstructed through depthwise separable convolution to achieve synchronous compression of parameter quantity and computational density, and construct a lightweight feature extraction base. Then, a frequency-aware dilated convolution operator was introduced in the detection head module to achieve adaptive fusion of local texture features and global context information by dynamically adjusting the receptive field. Finally, rigorous experimental verification of the proposed improvement scheme was conducted. The experimental results show that the improved model reduces computational complexity by 60.4%, reduces the number of parameters by 25%, and improves mean average precision (MAP) by 3.9%, while maintaining essentially unchanged detection speed. Compared to traditional detection methods, the detection method proposed in this application demonstrates advantages in detection accuracy, detection speed, and model deployment difficulty, providing a reliable technical solution for industrial-grade steel quality testing.

[0127] The present application also provides an application scenario, which applies the above-mentioned steel defect detection model training method. Specifically: the steel defect detection model training method provided in this embodiment can be applied in the steel surface defect detection scenario. The steel surface defect detection scenario includes an image acquisition link and a steel surface defect detection link; the target steel surface image enters the steel surface defect detection link from the image acquisition link, and the corresponding defect detection result is obtained through human-machine collaboration. The steel defect detection model training method provided in this embodiment belongs to the machine marking link in the steel surface defect detection link. Specifically, in the steel surface defect detection link process for the target steel surface image, the target steel surface image can be input into the trained steel defect detection model to obtain the defect detection result corresponding to the target steel surface image.

[0128] Example 2.

[0129] The steel defect detection model application method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, or integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the target steel surface image to the server 104. After the server 104 receives the target steel surface image, for the target steel surface image, the server 104 inputs the target steel surface image into the trained steel defect detection model to obtain the defect detection result corresponding to the target steel surface image. The server 104 can feed back the obtained defect detection result corresponding to the target steel surface image to the terminal 102. In addition, in some embodiments, the steel defect detection model training method can also be implemented separately by the server 104 or the terminal 102, such as the terminal 102 can directly perform steel defect detection on the target steel surface image, or the server 104 can obtain the target steel surface image from the data storage system and perform steel defect detection on the target steel surface image.

[0130] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0131] like Figure 7As shown, this embodiment provides a steel defect detection model application method including the following steps 401 to 402.

[0132] Step 401: Acquire a target steel surface image. The target steel surface image is an image captured of the high-speed rail to be inspected.

[0133] Step 402: Input the target steel surface image into a trained steel defect detection model to obtain a defect detection result corresponding to the target steel surface image; the trained steel defect detection model is a model trained using the steel defect detection model training method described in Example 1.

[0134] By implementing the above steps 401 and 402, the trained steel defect detection model is an improved YOLOv8 model. Using the improved YOLOv8 model for steel defect detection can improve detection efficiency and accuracy.

[0135] Example 3.

[0136] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store steel defect detection data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the steel defect detection model training method described in Example 1 or the steel defect detection model application method described in Example 2.

[0137] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0138] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0139] Example 4.

[0140] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0142] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0143] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0144] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A steel defect detection model training method, characterized in that: The steel defect detection model training method comprises: Acquire a data set; the data set includes a plurality of sample steel surface images and a defect category corresponding to each of the sample steel surface images; The sample steel surface image is used as input, and the defect category corresponding to the sample steel surface image is used as a label to train the steel defect detection model to obtain a trained steel defect detection model; the steel defect detection model is an improved YOLOv8 model; the improved YOLOv8 model replaces the second to fourth convolution modules of the feature extraction backbone network in the YOLOv8 model with residual convolution modules, replaces the first convolution module and the fifth convolution module of the feature extraction backbone network in the YOLOv8 model with depthwise separable convolution modules, adds a fusion convolution module after the last C2f module of the Neck network part of the YOLOv8 model, and adds a small defect prediction layer to the detection head module; the fusion convolution module includes Co The nv module and the C2f module, the Conv module, are used to upsample the output of the last C2f module of the Neck network part of the YOLOv8 model to obtain shallow features; the shallow features and the output of the first C2f module of the feature extraction backbone network are spliced ​​along the channel dimension to obtain a fused feature map; the C2f module in the fusion convolution module is used to perform cross-scale feature interaction on the fused feature map to obtain an optimized fused feature map, which is the input of the small defect prediction layer, and the dynamic snake convolution is introduced into the Bottleneck structure in the detection C2f module; the detection C2f module is a C2f module connected to the detection head; the residual convolution module is a convolution module that introduces a residual structure; the detection C2f module includes: First, the input image channel is reduced in dimension: 1×1 convolution is used to compress the number of input channels to 64 dimensions to reduce the amount of calculation: ; Where: is the intermediate output result, Represents the convolution kernel used for channel number compression, Represents the parameters of the convolution kernel itself, BN represents the batch normalization operation immediately after each convolution operation, and X represents the input image; The intermediate result Xmid is input to the 64-bit 3 The convolution layer consists of 3 convolution kernels. The ordinary convolution kernels in the convolution layer are replaced by dynamic snake convolution to better fit the defects for feature learning. The first step is to generate 9 dynamic sampling point coordinates for each position (p, q) of the feature map, with the center point (p, q) as the initial coordinate , generate the coordinates of the other 8 dynamic sampling points according to the following recursive formula: ; Where: and Respectively The horizontal and vertical coordinates of the dynamic sampling points; and Respectively The horizontal and vertical coordinates of the dynamic sampling points; , which means the The learnable parameters of the offset are: each dynamic sampling point is allowed to be 1 around the previous dynamic sampling point. Swing within the range of 1; Represents the curvature adjustment coefficient, which is used to control the curvature constraint strength. is the curvature constraint function, the curvature constraint function It is used to control the curvature of the convolution kernel sampling points so that it can adaptively fit the tubular structure. The mathematical expression of the curvature constraint function is: ; Where: Indicates that the input feature map I is at the coordinate point ( , ), which is used to quantify the local curvature; is the Frobenius norm of the Hessian matrix, reflecting the curvature intensity of the point; is an adjustable parameter used to control the sensitivity of the curvature constraint. When it belongs to the first setting value range, Stronger constraints are imposed on high curvature areas, forcing the sampling points to fit closely to the structure boundaries. When the sampling point belongs to the second set value range, it is allowed to deviate freely to adapt to the tubular structure in the flat area. The lower limit of the first set value range is greater than or equal to the upper limit of the second set value range. The exponential function Constrain the output to (0, 1], and the final output result is the constraint weight, which participates in the offset calculation of the coordinates of the above 8 dynamic sampling points; The second step is to calculate the coordinates of each dynamic sampling point ( , ) Obtain the actual sampling value through bilinear interpolation. The formula is as follows: ; Where, , is a dynamic coordinate, and the four nearest integer grid points (m, n) around it are the neighboring integer coordinate points; The third step is to aggregate the dynamic sampling value with the weight matrix, and the process is expressed as: ; in, For the The weight matrix of the sampling points is learned by the intermediate process, is the calculation result of the second step, is the bias constant, and after calculation, the specific coordinates of the convolution sum are finally obtained; For intermediate output Use 1 1 Convolution restores the feature dimension to the original number of channels, and its formula is expressed as: ; Where, Represents the parameters of the dimensionality-raising convolution kernel itself. The whole process is the inverse process of the dimensionality reduction process. After the above steps, the final output of the module is obtained. .

2. The steel defect detection model training method according to claim 1, characterized in that: The residual convolution module includes a first residual branch, a second residual branch, a third residual branch, a fusion layer and an activation function layer; the first residual branch includes a first convolution layer and a first normalization layer connected in sequence; the second residual branch includes a second convolution layer and a third normalization layer connected in sequence; the third residual branch includes a third normalization layer; the fusion layer is used to perform weighted fusion on the output of the first residual branch, the output of the second residual branch, and the output of the third residual branch to obtain a weighted fusion feature; The weighted fusion feature is the input of the activation function layer.

3. The steel defect detection model training method according to claim 1, characterized in that: The depth-wise separable convolution module includes a depth-wise convolution unit and a point-wise convolution unit; the depth-wise convolution unit includes three groups of first depth-wise convolution layers, which respectively extract features from the three color channels of the input feature map of the depth-wise separable convolution module to obtain a first intermediate output feature map, a second intermediate output feature map and a third intermediate output feature map; the point-wise convolution unit includes several second depth-wise convolution layers, each of which performs weighted fusion on the first intermediate output feature map, the second intermediate output feature map and the third intermediate output feature map to obtain a cross-channel fused high-order feature map corresponding to each second depth-wise convolution layer.

4. The steel defect detection model training method according to claim 1, characterized in that: The improved YOLOv8 model also introduces a frequency-adaptive expansion factor in the detection head module of YOLOv8, and the frequency-adaptive expansion factor is used to adjust the expansion rate of the convolution kernel.

5. The steel defect detection model training method according to claim 1, characterized in that: The steel defect detection model training method further includes: Evaluate the performance of the trained steel defect detection model.

6. A method for steel defect detection using a steel defect detection model, characterized in that: The detection method comprises: Acquire target steel surface image; The target steel surface image is input into a trained steel defect detection model to obtain a defect detection result corresponding to the target steel surface image; the trained steel defect detection model is a model trained using the steel defect detection model training method according to any one of claims 1 to 5.

7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steel defect detection model training method according to any one of claims 1 to 5 or the detection method according to claim 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steel defect detection model training method according to any one of claims 1 to 5 or the detection method according to claim 6 is implemented.

Citation Information

Patent Citations

  • Table identification method based on deep learning

    CN117636379A

  • Lightweight steel surface defect detection method based on improved YOLOv8n

    CN118196529A

  • Infant cry recognition nursing method and system and storage medium

    CN118298855A