Improved lightweight steel surface defect detection model for YOLOv9
By improving the lightweight steel surface defect detection model of YOLOv9 and utilizing the improved network structure and feature extraction technology, the problems of low detection efficiency and insufficient accuracy in traditional methods are solved, and efficient and accurate steel defect detection is achieved.
Patent Information
- Application Number
- CN202510907985.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing traditional and machine vision-based steel defect detection methods suffer from problems such as susceptibility to subjective human judgment, poor detection environment, low detection efficiency, and difficulty in adapting to large-scale data processing. Furthermore, deep learning algorithms have insufficient detection accuracy for small targets and complex backgrounds.
An improved lightweight steel surface defect detection model for YOLOv9 is proposed. By improving the backbone network and neck network, an effective network structure is generated. Combined with the Swin Transformer module and ODConv module, convolution processing and feature extraction are performed, and a multidimensional attention mechanism is used to identify the defect texture of steel.
It improves the accuracy of defect detection, enhances robustness to small targets, occluded targets and complex backgrounds, nearly doubles the detection accuracy, and maintains real-time inference speed.
Smart Images

Figure CN120747622B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of surface defect recognition technology, and in particular to an improved lightweight steel surface defect detection model based on YOLOv9. Background Technology
[0002] Steel, as a crucial raw material for industry, has always played a vital role in national economic development. The rapid development of new industries has placed increasingly higher demands on the surface quality of steel. Especially in industries such as automobiles, construction, and aerospace, the quality of steel directly affects the quality and performance of products. Defective steel can not only cause significant economic losses but also endanger personal safety, potentially leading to major accidents. Therefore, developing an efficient, accurate, and rapid algorithm for detecting defects on steel surfaces is of paramount importance.
[0003] Current defect detection methods can be broadly categorized into three types: manual defect detection, traditional machine vision-based defect detection, and deep learning-based defect detection. Traditional defect detection methods mainly include manual sampling, infrared detection, and magnetic flux leakage detection. These methods suffer from susceptibility to subjective human judgment, poor inspection environments, and high risk, making them unsuitable for efficient defect detection. Traditional machine vision detection methods have effectively improved the accuracy of defect feature extraction, demonstrating good stability and practical value. However, they lack adaptability when dealing with large-scale data processing, and their algorithmic structures cannot effectively address the challenges posed by changes in defect type or location. Advances in image processing and deep learning have effectively overcome the shortcomings of traditional detection methods in feature extraction, significantly improving the performance of detection systems.
[0004] Therefore, this invention provides a lightweight steel surface defect detection model that improves YOLOv9. Summary of the Invention
[0005] This invention improves the lightweight steel surface defect detection model of YOLOv9, enabling detailed positional perception analysis and having a significant impact on the development of computer vision.
[0006] This invention provides an improved lightweight steel surface defect detection model for YOLOv9, comprising:
[0007] The image acquisition module is used to acquire surface images of steel.
[0008] The network architecture module is used to improve the backbone and neck of YOLOv9 respectively, generating an effective network structure;
[0009] The convolution separation module is used to perform convolution processing on the surface image using the effective network structure to obtain several convolution textures of the surface image;
[0010] The sampling and repair module is used to upsample each of the convolutional textures to obtain several high-quality textures of the surface image.
[0011] The feature recognition module is used to extract features from the high-quality texture using a multidimensional attention mechanism, obtain several defect textures of the steel, generate defect information of the steel, and display it.
[0012] In one feasible approach
[0013] The image acquisition module includes:
[0014] An image acquisition unit is used to perform video monitoring on the steel to obtain several frames of video images;
[0015] An image optimization unit is used to perform color compensation on each frame of the video image to obtain grayscale saturation information corresponding to each video image.
[0016] The image selection unit is used to select a video image whose grayscale saturation information conforms to the color standard as the surface image of the steel.
[0017] In one feasible approach
[0018] The network architecture module includes:
[0019] The preprocessing unit is used to perform hierarchical processing on the YOLOv9 to obtain several layers of aggregation networks of the YOLOv9, and to identify the combination of modules contained in each aggregation network.
[0020] The backbone network improvement unit is used to locate the target aggregation network containing the RepNCSPELAN4 module in the YOLOv9 based on the module combination, and replace each of the RepNCSPELAN4 modules with the Swin Transformer module to generate the optimized backbone network of the YOLOv9.
[0021] The neck improvement unit is used to determine the module position of the upsampling module in YOLOv9 based on the module combination, and input the ultra-lightweight dynamic upsampling operator into the upsampling module to generate the optimized neck of YOLOv9;
[0022] The network construction unit is used to reorganize the aggregated network according to the improved information obtained for each layer of the aggregated network to generate an effective network structure.
[0023] In one feasible approach
[0024] Also includes:
[0025] The aggregation network has 9 layers, and the layer numbers of the target aggregation network are: layer 3, layer 5, layer 7, and layer 9.
[0026] In one feasible approach
[0027] The convolution separation module includes:
[0028] A depth convolutional unit is used to transmit the surface image to the target aggregation network containing the Swin Transformer module in YOLOv9 for depth convolution, so as to obtain the output information corresponding to each layer of the target aggregation network after the surface image passes through each layer of the target aggregation network.
[0029] The dilated convolutional unit is used to capture several global spatial information contained in the surface image in each output information to construct the dilated convolutional information of the surface image.
[0030] The texture localization unit is used to perform visual recognition on the surface image based on the output information and the dilated convolution information to obtain several convolutional textures of the surface image.
[0031] In one feasible approach
[0032] The sampling and repair module includes:
[0033] The network sampling unit is used to filter corresponding feature maps and sampling sets in big data based on the image specifications corresponding to each frame of the surface image, identify the pixel differences between different feature maps and each surface image in the sampling set, and construct network sampling parameters corresponding to each surface image.
[0034] The upsampling unit is used to perform pixel optimization on the corresponding surface image based on the network sampling parameters, input the optimized surface image into the effective network structure to upsample each convolutional texture respectively, and obtain the weighted result corresponding to each convolutional texture.
[0035] The texture optimization unit is used to input each weighted result into the corresponding surface image, identify the weighted smoothing information corresponding to each convolutional texture, and perform transpose convolution processing on the target convolutional texture with unqualified weighted smoothing information to obtain several optimized textures contained in each surface image.
[0036] The texture localization unit is used to fuse the surface image to obtain a multi-fusion image of the steel, identify the texture fusion results contained in the multi-fusion image, and perform texture enhancement according to the fusion number corresponding to the texture fusion result to obtain several high-quality textures of the steel.
[0037] In one feasible approach
[0038] The process of obtaining the weighted result corresponding to each of the convolutional textures includes:
[0039] Pixel information corresponding to each convolutional texture is constructed based on the upsampling result corresponding to each convolutional texture;
[0040] The pixel information is rearranged to obtain the texture edge corresponding to each convolutional texture;
[0041] When the texture clarity of the texture edges corresponding to all the convolutional textures is within a preset discrete range, a weighted result corresponding to each convolutional texture is generated based on the current convolutional feature corresponding to each convolutional texture.
[0042] Conversely, the pixel information corresponding to each convolutional texture is iteratively rearranged until the texture clarity of the texture edges corresponding to all convolutional textures is within a preset discrete range.
[0043] In one feasible approach
[0044] The feature recognition module includes:
[0045] The convolutional dilation unit is used to input each of the high-quality textures into the ODConv module of the effective network structure for depth-optimized convolution to obtain the wide field-of-view information corresponding to each of the high-quality textures.
[0046] The multi-dimensional recognition unit is used to perform multi-dimensional weighting on each of the high-quality textures using a multi-dimensional attention mechanism to obtain the spatial weight features, input weight features and output weight features corresponding to each high-quality texture, identify the field information corresponding to each weight feature in the corresponding wide field information, and draw the dimensional texture corresponding to each high-quality texture.
[0047] The feature generation unit is used to fuse the texture features corresponding to the same high-quality texture to obtain the fused texture features of the steel, identify the surface defect information corresponding to each fused texture feature, and construct the defect texture of the steel.
[0048] The information processing unit is used to count several defect textures corresponding to each type of steel, establish defect information of the steel according to the defect attributes and defect specifications corresponding to each defect texture, and display it.
[0049] In one feasible approach
[0050] Also includes:
[0051] The ODConv module optimization unit is used to collect the channel performance parameters corresponding to each single channel in the ODConv module, construct the corresponding sample channel based on the channel performance parameters, obtain the texture specifications of the high-quality texture, and construct the corresponding texture sample image based on the texture specifications.
[0052] The texture sample image is convolved using each of the sample channels to obtain the convolved sample output corresponding to each sample channel.
[0053] Based on the output specification corresponding to each convolutional sample output, derive the context weight coefficient of the single channel for the texture sample.
[0054] Convolutional output images corresponding to the single channel are constructed based on the convolutional sample outputs. The loss of each convolutional output image is evaluated using the texture sample images to obtain the convolutional accuracy corresponding to each sample channel.
[0055] The context weight coefficients corresponding to each sample channel are superimposed and adjusted to obtain the numerical trend of the convolution accuracy.
[0056] Based on the numerical change trend, the optimal context weight coefficient corresponding to the sample channel is derived, and the optimal context weight coefficient is used to optimize the corresponding single channel to generate the optimized single channel corresponding to the ODConv module.
[0057] In one feasible approach
[0058] Also includes:
[0059] The defect statistics module is used to identify the defect location corresponding to the defect texture of each steel material, and generate and display the defect location report of the steel material by combining the defect attributes and defect specifications of the steel material.
[0060] The beneficial effects of the above technical solution are as follows: Before defect detection, the backbone network and neck of YOLOv9 are improved to generate an effective network structure. This effective network structure is used to perform convolution processing on the surface image of the steel to obtain several convolutional textures of the steel. Then, the spatial clarity of the convolutional textures is improved by upsampling technology to obtain high-quality textures. Further feature extraction is performed on the high-quality textures to determine the defect textures of the steel, thereby generating the defect information of the steel. In this way, the accuracy of defect detection can be improved. The improved model is superior to similar classic algorithms in terms of detection accuracy, and the defect recognition performance is improved by nearly 100%. It effectively improves the robustness to small targets, occluded targets and complex backgrounds, while maintaining real-time inference speed.
[0061] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0062] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0063] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0064] Figure 1 This is a schematic diagram illustrating the composition of the lightweight steel surface defect detection model of the improved YOLOv9 in an embodiment of the present invention.
[0065] Figure 2 This is a schematic diagram of the network architecture module of the improved YOLOv9 lightweight steel surface defect detection model in an embodiment of the present invention. Detailed Implementation
[0066] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0067] Example 1
[0068] This embodiment provides an improved lightweight steel surface defect detection model for YOLOv9, such as... Figure 1 As shown, it includes:
[0069] The image acquisition module is used to acquire surface images of steel.
[0070] The network architecture module is used to improve the backbone and neck of YOLOv9 respectively, generating an effective network structure;
[0071] The convolution separation module is used to perform convolution processing on the surface image using the effective network structure to obtain several convolution textures of the surface image;
[0072] The sampling and repair module is used to upsample each of the convolutional textures to obtain several high-quality textures of the surface image.
[0073] The feature recognition module is used to extract features from the high-quality texture using a multidimensional attention mechanism, obtain several defect textures of the steel, generate defect information of the steel, and display it.
[0074] In this example, the effective network structure represents a network model that can perform surface defect detection;
[0075] In this example, the backbone network represents the base network for feature extraction in YOLOv9;
[0076] In this example, the neck represents the network that performs further processing on the features extracted from the backbone network;
[0077] In this example, convolutional texture represents the texture of the surface image obtained after convolution processing;
[0078] In this example, upsampling represents the process of improving the spatial texture edges of the convolutional texture;
[0079] In this example, a high-quality texture represents a texture that is sharper than a convolutional texture.
[0080] The working principle and beneficial effects of the above technical solution are as follows: Before defect detection, the backbone network and neck of YOLOv9 are improved to generate an effective network structure. This effective network structure is used to perform convolution processing on the surface image of the steel to obtain several convolutional textures of the steel. Then, the spatial clarity of the convolutional textures is improved by upsampling technology to obtain high-quality textures. Further feature extraction is performed on the high-quality textures to determine the defect textures of the steel, thereby generating the defect information of the steel. In this way, the accuracy of defect detection can be improved. The improved model is superior to similar classic algorithms in terms of detection accuracy, and the defect recognition performance is improved by nearly 100%.
[0081] Example 2
[0082] Based on Example 1, the improved YOLOv9 lightweight steel surface defect detection model includes an image acquisition module comprising:
[0083] An image acquisition unit is used to perform video monitoring on the steel to obtain several frames of video images;
[0084] An image optimization unit is used to perform color compensation on each frame of the video image to obtain grayscale saturation information corresponding to each video image.
[0085] The image selection unit is used to select a video image whose grayscale saturation information conforms to the color standard as the surface image of the steel.
[0086] In this example, the color standard is a pre-set value for the color saturation of the surface image by the inspectors.
[0087] The working principle and beneficial effects of the above technical solution are as follows: In order to ensure the effectiveness of defect detection, the clarity of the surface image must be guaranteed. Therefore, the steel is first monitored by video to obtain several frames of video images. Then, color compensation is performed on the video images to determine the color saturation of the video images. Finally, the surface image with qualified saturation is selected. In this way, the effectiveness of the surface image can be guaranteed, and the efficiency and quality of defect detection can be improved.
[0088] Example 3
[0089] Based on Example 1, the improved YOLOv9 lightweight steel surface defect detection model, such as Figure 2 As shown, the network architecture module includes:
[0090] The preprocessing unit is used to perform hierarchical processing on the YOLOv9 to obtain several layers of aggregation networks of the YOLOv9, and to identify the combination of modules contained in each aggregation network.
[0091] The backbone network improvement unit is used to locate the target aggregation network containing the RepNCSPELAN4 module in the YOLOv9 based on the module combination, and replace each of the RepNCSPELAN4 modules with the Swing TransformerR module to generate the optimized backbone network of the YOLOv9.
[0092] The neck improvement unit is used to determine the module position of the upsampling module in YOLOv9 based on the module combination, and input the ultra-lightweight dynamic upsampling operator into the upsampling module to generate the optimized neck of YOLOv9;
[0093] The network construction unit is used to reorganize the aggregated network according to the improved information obtained for each layer of the aggregated network to generate an effective network structure.
[0094] In this example, each aggregation network contains one or more modules;
[0095] In this example, module combination represents the statistical results of all modules contained in an aggregated network;
[0096] In this example, the RepNCSPELAN4 module is an innovative network module design in YOLOv9, mainly used for feature extraction and feature fusion;
[0097] The RepNCSPELAN4 module captures feature information at different scales through a spatial pyramid structure, enhancing the model's ability to detect targets of different sizes. It combines attention mechanisms and normalization techniques to improve feature representation capabilities, highlight important features, and utilizes RepVGG reparameterization technology to significantly reduce the amount of computation and parameters during inference without compromising performance. Furthermore, it can fuse features from multiple branches to comprehensively utilize feature information at different levels and scales.
[0098] The RepNCSPELAN4 module enables YOLOv9 to improve detection accuracy for small targets and complex scenes while maintaining high detection speed.
[0099] In this example, the Swin Transformer module represents a computer vision model based on the Transformer architecture, which has significant effects on image classification and segmentation tasks;
[0100] Replacing the RepNCSPELAN4 module with the Swing Transformer module can improve the accuracy of the optimized backbone network. When YOLOv9 processes targets of different sizes, it can improve detection accuracy through multi-scale feature fusion of the Swing Transformer module. Furthermore, the Swing Transformer module introduces a window attention mechanism, which can efficiently compute self-attention in local regions, effectively avoiding the problem of excessive computational complexity of the RepNCSPELAN4 module.
[0101] In this example, the upsampling module refers to the module used to perform the upsampling operation;
[0102] In this example, the ultralightweight dynamic upsampling operator represents a sampling operator that can improve upsampling efficiency and has low resource consumption.
[0103] The working principle and beneficial effects of the above technical solution are as follows: By performing layered processing on YOLOv9 and then optimizing the aggregation network of different layers, an effective network structure that is more suitable for defect detection can be obtained. This network structure can better capture global and local image features. By using this effective network structure, key features of small targets can be preserved during defect detection, thereby improving the quality of defect detection. This allows the model to automatically capture long-range dependency information and detailed features at different levels.
[0104] Example 4
[0105] Based on Example 3, the improved YOLOv9 lightweight steel surface defect detection model further includes:
[0106] The aggregation network has 9 layers, and the layer numbers of the target aggregation network are: layer 3, layer 5, layer 7, and layer 9.
[0107] In this example, after multiple experiments, the model using four Swin Transformer modules performed best. By using four Swin Transformer modules to reorganize the improved information, the local-global attention collaboration mechanism of the Swin Transformer modules can be used to change the sensitivity of the effective network structure to capturing local features.
[0108] The working principle and beneficial effects of the above technical solution are as follows: By using the target aggregation network to reorganize the improved information, the combination of low-level features and high-level semantic information can be enhanced through cross-layer information interaction, thereby optimizing the detection effect of small objects and complex backgrounds.
[0109] Example 5
[0110] Based on Example 1, the improved YOLOv9 lightweight steel surface defect detection model, the convolution separation module, includes:
[0111] A depth convolutional unit is used to transmit the surface image to the target aggregation network containing the Swin Transformer module in YOLOv9 for depth convolution, so as to obtain the output information corresponding to each layer of the target aggregation network after the surface image passes through each layer of the target aggregation network.
[0112] The dilated convolutional unit is used to capture several global spatial information contained in the surface image in each output information to construct the dilated convolutional information of the surface image.
[0113] The texture localization unit is used to perform visual recognition on the surface image based on the output information and the dilated convolution information to obtain several convolutional textures of the surface image.
[0114] In this example, the output information represents the result of the surface image passing through a target aggregation network layer. The working principle and beneficial effects of the above technical solution are as follows: Traditional convolutional neural networks struggle to effectively handle cross-scale information, and YOLOv9 often suffers from insufficient detection accuracy for small or distant objects. Therefore, a target aggregation network containing a SwingTransformer module is used to convolve the surface image. The layered design of the Transformer module enables effective convolution of multi-scale features. Specifically, the surface image is first input into the target aggregation network layer in YOLOv9 for depth convolution. Then, the global spatial information corresponding to the output information of each layer is captured to construct the dilated convolution information of the surface image. Finally, the output information and dilated convolution information are used to perform visual recognition of the surface image, determining the convolutional texture of the surface image. This method allows for four depth convolutions and one dilated convolution on the surface image, reducing the number of parameters in the convolutional texture and broadening the recognition field of the convolutional texture, resulting in realistic and effective steel texture. Furthermore, cross-layer information interaction enhances the combination of low-level features and high-level semantic information, optimizing the detection performance for small objects and complex backgrounds.
[0115] Example 6
[0116] Based on Example 1, the improved YOLOv9 lightweight steel surface defect detection model, the sampling and repair module, includes:
[0117] The network sampling unit is used to filter corresponding feature maps and sampling sets in big data based on the image specifications corresponding to each frame of the surface image, identify the pixel differences between different feature maps and each surface image in the sampling set, and construct network sampling parameters corresponding to each surface image.
[0118] The upsampling unit is used to perform pixel optimization on the corresponding surface image based on the network sampling parameters, input the optimized surface image into the effective network structure to upsample each convolutional texture respectively, and obtain the weighted result corresponding to each convolutional texture.
[0119] The texture optimization unit is used to input each weighted result into the corresponding surface image, identify the weighted smoothing information corresponding to each convolutional texture, and perform transpose convolution processing on the target convolutional texture with unqualified weighted smoothing information to obtain several optimized textures contained in each surface image.
[0120] The texture localization unit is used to fuse the surface image to obtain a multi-fusion image of the steel, identify the texture fusion results contained in the multi-fusion image, and perform texture enhancement according to the fusion number corresponding to the texture fusion result to obtain several high-quality textures of the steel.
[0121] In this example, the image specification refers to the specification of the pixel distribution in the image of the surface;
[0122] In this example, the feature map represents an existing image in the big data where the defect features have been identified;
[0123] In this example, the sampling set represents the set of samples that are related to defects in the feature map;
[0124] In this example, the weighted smoothing information represents the result of weighting the convolutional texture.
[0125] The working principle and beneficial effects of the above technical solution are as follows: By filtering feature maps and sampling sets related to the surface image from big data, network sampling parameters for the surface image are constructed. These parameters are then used to optimize the surface image pixels. An effective network is then used to upsample the convolutional texture, achieving weighting. Furthermore, transposed convolutions are performed on convolutional textures with unsatisfactory weighted smoothing information, generating high-quality textures for the surface image. This method enables depth recognition of convolutional textures on the surface image, determining the details of the convolutional textures, improving the clarity of the surface image, and generating high-quality textures for subsequent defect identification. Through multi-stage sampling, optimization, and texture enhancement, the model's accuracy in identifying subtle defects on the steel surface is improved while maintaining lightweight characteristics. This module, through the collaborative work of four core units, forms a complete processing chain from feature selection to high-quality texture generation.
[0126] Example 7
[0127] Based on Example 6, the process of obtaining the weighted result corresponding to each convolutional texture using the improved YOLOv9 lightweight steel surface defect detection model includes:
[0128] Pixel information corresponding to each convolutional texture is constructed based on the upsampling result corresponding to each convolutional texture;
[0129] The pixel information is rearranged to obtain the texture edge corresponding to each convolutional texture;
[0130] When the texture clarity of the texture edges corresponding to all the convolutional textures is within a preset discrete range, a weighted result corresponding to each convolutional texture is generated based on the current convolutional feature corresponding to each convolutional texture.
[0131] Conversely, the pixel information corresponding to each convolutional texture is iteratively rearranged until the texture clarity of the texture edges corresponding to all convolutional textures is within a preset discrete range.
[0132] In this example, the preset discrete range is [0, 0.5];
[0133] In this example, iterative rearrangement refers to the process of rearranging pixel information multiple times.
[0134] The working principle and beneficial effects of the above technical solution are as follows: In order to weight the convolutional texture, the pixel information of the convolutional texture is first constructed based on the upsampling result. At this time, in order to ensure that the convolutional texture remains unchanged, the pixel information is rearranged. Then, the pixel information of the convolutional texture with unqualified texture edges is rearranged again until the texture edges meet the requirements. Then, the convolutional texture is weighted according to the current convolutional features of the convolutional texture, thus realizing the weighting work.
[0135] Example 8
[0136] Based on Example 1, the improved YOLOv9 lightweight steel surface defect detection model includes a feature recognition module comprising:
[0137] The convolutional dilation unit is used to input each of the high-quality textures into the ODConv module of the effective network structure for depth-optimized convolution to obtain the wide field-of-view information corresponding to each of the high-quality textures.
[0138] The multi-dimensional recognition unit is used to perform multi-dimensional weighting on each of the high-quality textures using a multi-dimensional attention mechanism to obtain the spatial weight features, input weight features and output weight features corresponding to each high-quality texture, identify the field information corresponding to each weight feature in the corresponding wide field information, and draw the dimensional texture corresponding to each high-quality texture.
[0139] The feature generation unit is used to fuse the texture features corresponding to the same high-quality texture to obtain the fused texture features of the steel, identify the surface defect information corresponding to each fused texture feature, and construct the defect texture of the steel.
[0140] The information processing unit is used to count several defect textures corresponding to each type of steel, establish defect information of the steel according to the defect attributes and defect specifications corresponding to each defect texture, and display it.
[0141] In this example, the multiple dimensions are: spatial dimension, input dimension, and output dimension.
[0142] In this example, the wide field of view information represents the information presented by the high-quality texture in a wide field of view scene;
[0143] In this example, the ODConv module is a module that performs depth-optimized convolution. Its core mechanism lies in using a learnable deformation module to adaptively adjust the geometry and scale of the convolution kernel based on the feature information of the input data, thereby significantly improving the performance of the convolutional neural network.
[0144] The working principle and beneficial effects of the above technical solution are as follows: By using the ODConv module to perform convolution on high-quality textures, the wide field of view information corresponding to each high-quality texture is determined, as well as its weight features in space, and the weight features corresponding to the input and output. Then, the field of view information of each weight feature is identified to draw the corresponding dimensional texture. Furthermore, the texture features are fused to determine the defect texture of the steel. Finally, the defect information of the steel is established based on the defect attributes and defect specifications of each defect texture. In this way, the defect texture can be deeply processed, clarifying its performance in different dimensions, and thus constructing effective defect information for inspection personnel to refer to. At the same time, the convolution processing using the ODconv module can effectively improve the model's extraction of local defect information and improve the model's accuracy.
[0145] Example 9
[0146] Based on Example 8, the improved YOLOv9 lightweight steel surface defect detection model further includes:
[0147] The ODConv module optimization unit is used to collect the channel performance parameters corresponding to each single channel in the ODConv module, construct the corresponding sample channel based on the channel performance parameters, obtain the texture specifications of the high-quality texture, and construct the corresponding texture sample image based on the texture specifications.
[0148] The texture sample image is convolved using each of the sample channels to obtain the convolved sample output corresponding to each sample channel.
[0149] Based on the output specification corresponding to each convolutional sample output, derive the context weight coefficient of the single channel for the texture sample.
[0150] Convolutional output images corresponding to the single channel are constructed based on the convolutional sample outputs. The loss of each convolutional output image is evaluated using the texture sample images to obtain the convolutional accuracy corresponding to each sample channel.
[0151] The context weight coefficients corresponding to each sample channel are superimposed and adjusted to obtain the numerical trend of the convolution accuracy.
[0152] Based on the numerical change trend, the optimal context weight coefficient corresponding to the sample channel is derived, and the optimal context weight coefficient is used to optimize the corresponding single channel to generate the optimized single channel corresponding to the ODConv module.
[0153] In this example, a single channel includes all input channels and all output channels.
[0154] The working principle and beneficial effects of the above technical solution are as follows: Optimizing the channels of the ODConv module makes all spatial locations, all input channels, and all output channels of the ODConv module smoother, providing performance assurance for capturing rich defect clues. At the same time, optimizing the ODConv module can also enhance the model's attention to ordinary quality anchor boxes. It can also adaptively adjust the loss weights according to the target scale features, effectively improving the detection performance and localization accuracy of small targets. Furthermore, optimizing the channels for the actual specifications of high-quality textures can improve the convolution quality of the ODConv module, achieve fast convergence, and avoid overfitting.
[0155] Example 10
[0156] Based on Example 8, the improved YOLOv9 lightweight steel surface defect detection model further includes:
[0157] The defect statistics module is used to identify the defect location corresponding to the defect texture of each steel material, and generate and display the defect location report of the steel material by combining the defect attributes and defect specifications of the steel material.
[0158] In this example, one defect texture corresponds to one defect location, but one defect location can have one or more defect textures;
[0159] In this example, the defect attributes include: cracks, porosity, scratches, scale, and rust;
[0160] In this example, the defect specification represents the size of a defect.
[0161] The working principle and beneficial effects of the above technical solution are as follows: generating a defect location report for each piece of steel can help relevant personnel quickly locate the defect and improve the efficiency of sorting defective steel.
[0162] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. An improved lightweight steel surface defect detection model of YOLOv9, characterized in that, The method comprises the following steps: An image acquisition module is configured to acquire a surface image of a steel product; A network architecture module is configured to improve a backbone network and a neck of a YOLOv9 respectively to generate an effective network structure; A convolution separation module is configured to perform convolution processing on the surface image by using the effective network structure to obtain a plurality of convolution textures of the surface image; A sampling repair module is configured to perform up-sampling on each of the convolution textures to obtain a plurality of high-quality textures of the surface image; A feature recognition module is configured to perform feature extraction on the high-quality textures by using a multi-dimensional attention mechanism to obtain a plurality of defect textures of the steel product, generate defect information of the steel product, and display the defect information; The sampling repair module comprises: A network sampling unit is configured to filter corresponding feature maps and a sampling set in big data based on an image specification corresponding to each frame of the surface image, identify a pixel difference between different feature maps and each of the surface images in the sampling set, and construct a network sampling parameter corresponding to each of the surface images; An up-sampling unit is configured to perform pixel optimization on the corresponding surface image based on the network sampling parameter, input the optimized surface image into the effective network structure, and perform up-sampling on each of the convolution textures to obtain a weighted result corresponding to each of the convolution textures; A texture optimization unit is configured to input each of the weighted results into the corresponding surface image, identify weighted smoothing information corresponding to each of the convolution textures, and perform transpose convolution processing on a target convolution texture that does not meet the weighted smoothing information to obtain a plurality of optimized textures contained in each of the surface images; A texture positioning unit is configured to fuse the surface images to obtain a multi-fusion image of the steel product, identify a texture fusion result contained in the multi-fusion image, perform texture enhancement according to a fusion number corresponding to the texture fusion result, and obtain a plurality of high-quality textures of the steel product.
2. The lightweight steel surface defect detection model improved YOLOv9 of claim 1, wherein, The image acquisition module comprises: An image acquisition unit is configured to obtain a plurality of video images by video monitoring of the steel product; An image optimization unit is configured to perform color compensation on each of the video images to obtain gray saturation information corresponding to each of the video images; An image selection unit is configured to select a video image whose gray saturation information meets a color standard as the surface image of the steel product.
3. The lightweight steel surface defect detection model improved YOLOv9 of claim 1, wherein, The network architecture module comprises: A preprocessing unit is configured to perform hierarchical processing on the YOLOv9 to obtain a plurality of layer aggregation networks of the YOLOv9, and identify a module combination contained in each of the aggregation networks; A backbone network improvement unit is configured to locate a target aggregation network containing a RepNCSPELAN4 module in the YOLOv9 based on the module combination, replace each of the RepNCSPELAN4 modules by using a Swin Transformer module, and generate an optimized backbone network of the YOLOv9. The neck improvement unit is used for determining the module position of the up-sampling module in the YOLOv9 based on the module combination, inputting a super-light dynamic up-sampling operator into the up-sampling module to generate an optimized neck of the YOLOv9; The network construction unit is used for recombining the aggregation network according to the improvement information corresponding to each layer of the aggregation network to generate an effective network structure.
4. The lightweight steel surface defect detection model improved YOLOv9 of claim 3, wherein, Further comprising: The number of layers of the aggregation network is 9, wherein the layer numbers of the target aggregation network are: the 3rd layer, the 5th layer, the 7th layer and the 9th layer.
5. The lightweight steel surface defect detection model improved YOLOv9 of claim 1, wherein, The convolution separation module comprises: The deep convolution unit is used for transmitting the surface image to the target aggregation network containing a Swin Transformer module in the YOLOv9 for deep convolution to obtain output information corresponding to each target aggregation network layer after the surface image passes through the target aggregation network layer; The dilated convolution unit is used for capturing a plurality of global spatial information contained in the surface image in each output information respectively to construct dilated convolution information of the surface image; The texture positioning unit is used for performing visual identification on the surface image based on the output information and the dilated convolution information to obtain a plurality of convolution textures of the surface image.
6. The lightweight steel surface defect detection model improved YOLOv9 of claim 1, wherein, The process of obtaining the weighting result corresponding to each convolution texture comprises: constructing pixel information corresponding to each convolution texture according to the up-sampling result corresponding to each convolution texture; rearranging the pixel information to obtain a texture edge corresponding to each convolution texture; when the texture clarity of the texture edge corresponding to all the convolution textures is within a preset discrete range, generating the weighting result corresponding to each convolution texture according to the current convolution feature corresponding to each convolution texture; otherwise, iteratively rearranging the pixel information corresponding to each convolution texture until the texture clarity of the texture edge corresponding to all the convolution textures is within a preset discrete range.
7. The lightweight steel surface defect detection model improved YOLOv9 of claim 1, wherein, The feature recognition module comprises: The convolution dilating unit is used for inputting each high-quality texture into an ODConv module of the effective network structure respectively for deep optimization convolution to obtain wide field of view information corresponding to each high-quality texture; The multi-dimensional recognition unit is used for performing multi-dimensional weighting on each high-quality texture respectively by using a multi-dimensional attention mechanism to obtain multi-dimensional features corresponding to each high-quality texture, identifying the field of view information corresponding to each dimensional feature in the corresponding wide field of view information, and drawing a dimensional texture corresponding to each high-quality texture; The feature generation unit is used for fusing texture features corresponding to the same high-quality texture to obtain fusion texture features of the steel material, identifying surface defect information corresponding to each fusion texture feature respectively, and constructing defect textures of the steel material; The information arrangement unit is used for counting a plurality of defect textures corresponding to each steel material, establishing defect information of the steel material according to the defect attribute and defect specification corresponding to each defect texture, and displaying the defect information.
8. The lightweight steel surface defect detection model improved YOLOv9 of claim 7, wherein, Further comprising: An ODConv module optimization unit is configured to collect channel performance parameters corresponding to each single channel in the ODConv module respectively, construct a corresponding sample channel based on the channel performance parameters, obtain a texture specification of the high-quality texture, and construct a corresponding texture sample image based on the texture specification; Each sample channel is used to perform convolution test on the texture sample image respectively, and a convolution sample output corresponding to each sample channel is obtained; Context weight coefficients of the single channel corresponding to each convolution sample output are derived according to output features corresponding to the convolution sample output; A convolution output image corresponding to the single channel is constructed according to the convolution sample output, loss evaluation is performed on each convolution output image using the texture sample image respectively, and a context weight accuracy rate corresponding to each sample channel is obtained; The context weight coefficients corresponding to each sample channel are adjusted respectively, and a numerical change trend corresponding to the weight accuracy rate is obtained; Optimal context weight coefficients corresponding to the sample channel are derived according to the numerical change trend, the single channel corresponding to the optimal context weight coefficients is optimized, and an optimized single channel corresponding to the ODConv module is generated.
9. The improved YOLOv9-based lightweight steel surface defect detection model of claim 7, wherein, Further comprising: A defect statistical module is configured to identify a defect position corresponding to a defect texture of each steel product respectively, generate a defect positioning report of the steel product in combination with the defect attribute and the defect specification corresponding to the steel product, and display the defect positioning report.
Citation Information
Patent Citations
Lead screw module surface defect online detection method and device based on machine vision
CN119090862A
Breast image segmentation method based on dynamic sampling and multi-scale parallel attention
CN119810124A