Multidirectional inclined steel grade detection and identification method based on deep learning
Through the combination of the improved YOLOv11 model and the CRNN framework, the complex background and multi-directional inclination problems in multi-directional inclination steel number detection and identification are solved, and the precise positioning and automatic identification of metal sheet steel number is realized, which improves the detection accuracy and identification efficiency.
Patent Information
- Application Number
- CN202510472961.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has problems in the automatic detection and identification of steel numbers in metal sheets, where multi-directional inclination is inaccurate detection and positioning and incomplete identification are encountered. Especially when the steel numbers and background contrast are low, the character style is variable, the proportion is inconsistent, and multiple steel numbers are tilted in different directions, it is difficult to achieve efficient and accurate detection and identification.
Using a two-stage detection method based on deep learning, first coarse positioning and tilt correction are performed through the improved YOLOv11 model, and then a CRNN framework is constructed for precise positioning and character recognition. Combined with image processing technology, precise positioning and automatic recognition of multi-directional inclined steel numbers are achieved.
It improves the detection accuracy and identification efficiency of multi-directional inclined steel numbers, improves the automated detection and identification effect of metal sheet steel numbers, and enhances the accuracy and speed of material tracking and recording during the production process.
Smart Images

Figure CN120340014A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rolling automation and steel grade identification, and more particularly, to a multi-directional inclined steel grade detection and identification method based on deep learning. Background Art
[0002] With the development of industrial automation and intelligent manufacturing, the automatic detection and identification technology of metal sheet steel grades plays an increasingly important role in the metallurgical industry. As the unique identifier of metal sheets, the accurate identification of steel grades is of great significance for the production management, quality traceability, and logistics distribution of metallurgical enterprises. Currently, the detection and identification of metal sheet steel grades mainly rely on manual methods, which are not only inefficient but also prone to errors, making it difficult to meet the requirements of modern metallurgical enterprises for automated and intelligent production.
[0003] In recent years, with the rapid development of computer vision and deep learning technologies, the automatic detection and identification method of metal sheet steel grades based on image recognition has gradually become a research hotspot. Patent CN117173716A discloses a high-temperature slab ID character recognition method and system based on deep learning. This method builds a slab ID character recognition model, including an OCR text detection model, an OCR direction classifier, and an OCR text recognition model connected in sequence, and introduces strategies such as deep mutual learning during training, effectively improving the efficiency, accuracy, and stability of ID character recognition. Patent CN118298324A proposes a deep learning photovoltaic panel recognition method based on the multi-scale Retinex algorithm. This method performs data annotation and data augmentation on the remote sensing image dataset, and conducts target detection based on the trained photovoltaic panel recognition model, overcoming the phenomenon of unclear remote sensing images caused by weather factor limitations and improving the feature recognition ability and accuracy.
[0004] In the field of object detection, Patent CN118552958A introduces a multi-view pipe online recognition method and system based on machine vision. This method constructs a pipe recognition model based on the improved YOLOv5s algorithm. By collecting images of various types of pipes from multiple perspectives, performing data augmentation processing and annotation, the recognition of pipes is no longer limited by perspectives and occlusion problems, improving the recognition accuracy. Patent CN117894004A proposes a deep learning hot-cast billet number recognition method and system. This method performs denoising, shearing, and zero-mean processing on the original image, uses the VGG19 network model to detect the text orientation of the hot-cast billet number, calculates the inclination angle of the billet number and performs rotation correction, and finally constructs and trains an RS-RCNN model to recognize the hot-cast billet number, improving the recognition accuracy of the hot-cast billet number.
[0005] However, there are still some problems and challenges in the automatic detection and recognition of the steel grades of metal sheets in the existing technology. First, the image background of the steel grades of metal sheets is usually complex, and the contrast between the steel grades and the background is low, making it difficult to detect and recognize. Second, the number of characters contained in a group of steel grades is generally uncertain, the character styles are diverse, and the character quality is poor, increasing the difficulty of recognition. Third, the proportion of the steel grades in the image varies, and there is a problem of small target detection due to the relatively small proportion of the steel grades in some images. Finally, the steel grades are mostly marked on the side of the metal sheets, and multiple sheets are stacked untidily during transportation and storage. Therefore, the collected steel grade images have a complex situation where multiple steel grades are tilted in different directions in one image, resulting in difficulties in detecting and positioning the steel grades during the production process. At the same time, the recognition of the steel grades in the next process is greatly affected by the quality of the detection and positioning results of the steel grades of the metal sheets. Incomplete detection and positioning of the steel grades or detection and positioning results containing a large amount of redundant background information are likely to lead to incorrect recognition of the steel grades.
[0006] Most of the detection and recognition methods in the existing technology are designed for specific scenarios, and there is a lack of effective solutions for the detection and recognition of multi-directionally tilted steel grades of metal sheets. Especially when dealing with the complex situation of multiple steel grades tilted in different directions, the detection and positioning accuracy and recognition effect of the existing methods are often unsatisfactory. Therefore, there is an urgent need for a method that can effectively handle the detection and recognition of multi-directionally tilted steel grades to improve the accuracy and efficiency of the automatic detection and recognition of the steel grades of metal sheets. Summary of the Invention
[0007] In view of this, the present invention provides a method for detecting and recognizing multi-directionally tilted steel grades based on deep learning to solve the problems of incomplete recognition, low accuracy, and low automation in the detection and positioning of the steel grades of metal sheets during the rolling production process.
[0008] To achieve the above object, the technical means adopted by the present invention are as follows:
[0009] A method for detecting and recognizing multi-directionally tilted steel grades based on deep learning, comprising the following steps:
[0010] S1: Collect the original images of multi-directionally tilted steel grades of metal sheets and make the first-stage rough positioning data set;
[0011] S2: Improve YOLOv11 based on the problems existing in the first-stage rough positioning data set, construct the original model for detecting and positioning steel grades based on the improved YOLOv11, and train it with the first-stage rough positioning data set to obtain the first-stage rough positioning model for multi-directionally tilted steel grades;
[0012] S3: Use the first-stage rough positioning model of the multi-directionally inclined steel number to obtain the second-stage precisely positioned original image. After performing tilt correction processing using image processing technology, produce the second-stage precisely positioned dataset;
[0013] S4: Use the second-stage precisely positioned dataset to train the original steel number detection and positioning model to obtain the second-stage precisely positioning model of the multi-directionally inclined steel number;
[0014] S5: Use the second-stage precisely positioning model of the multi-directionally inclined steel number to obtain the steel number automatic recognition image data to produce the steel number automatic recognition dataset;
[0015] S6: Construct a steel number automatic recognition model based on the CRNN framework and use the steel number automatic recognition dataset to train it to obtain the multi-directionally inclined steel number automatic recognition model;
[0016] S7: Deploy the first-stage rough positioning model, the second-stage precisely positioning model, and the automatic recognition model of the multi-directionally inclined steel number on-site. Combine the image processing technology to construct a multi-directionally inclined steel number detection, positioning, and automatic recognition system, and according to the real-time collected steel number image, output the recognition result of the steel number character sequence in real time.
[0017] Further, the improved YOLOv11 includes:
[0018] (1) Design a multi-scale convolution module - residual mixed convolution module, that is, the RMC module, including those connected in sequence:
[0019] a. Primary feature extraction layer, using DWConv of the first preset size to perform spatial feature extraction on the original input features;
[0020] b. Multi-branch parallel processing layer, including those acting on the output of the primary feature extraction layer simultaneously:
[0021] b1. DWConv branch of the second preset size,
[0022] b2. PConv branch of the second preset size,
[0023] b3. Parameter compression type Adown branch;
[0024] c. Cross-layer residual processing layer, using DWConv of the third preset size to perform cross-layer feature extraction on the original input features;
[0025] d. Feature fusion layer, perform channel stacking on the output features of each branch of the multi-branch parallel processing layer and the output features of the cross-layer residual processing layer, and perform dimension matching through the channel adjustment convolution layer;
[0026] e. Post - processing layer, which performs BN normalization processing and SiLU activation function transformation on the features output by the feature fusion layer in sequence;
[0027] Among them, the third preset size is greater than the first preset size, and the parameter quantity of the Adown branch is less than that of the conventional convolutional structure;
[0028] (2) Improve the ordinary convolutions of CBSModule in the feature extraction network Backbone and the feature fusion network Neck in YOLOv11 to the RMC module, forming a multi - level enhanced feature processing architecture.
[0029] Furthermore, the improved YOLOv11 includes embedding a channel - priority - dual - spatial attention mechanism, namely the CP - DSA module, at the preset hierarchical nodes of the feature extraction network. This module includes:
[0030] a. Channel - priority processing branch, which generates the first attention feature map through the CPCA attention mechanism;
[0031] b. Spatial - coordinate attention branch, which generates the second attention feature map through the CA attention mechanism;
[0032] c. Feature fusion unit, which performs cross - dimensional splicing and fusion on the first attention feature map and the second attention feature map, and finally outputs an enhanced attention feature map;
[0033] Among them, the CP - DSA module is configured in the primary feature enhancement stage of the feature extraction network, and is specifically connected to the output end of the first multi - scale convolution module of the feature extraction network. By jointly optimizing the channel sensitivity and spatial position correlation, the multi - dimensional feature expression ability is enhanced.
[0034] Furthermore, the improved YOLOv11 includes improving the original CIoU loss function in the output prediction network Head of YOLOv11 to the WIoU loss function to more economically and accurately evaluate the positional relationship between the predicted target detection box and the marked detection box.
[0035] Furthermore, the principle of the tilt correction process is: judging the tilt angle of the steel number through the straight lines of the upper and lower edges of the metal plate, and performing rotation correction according to the tilt angle; the specific image processing technology is:
[0036] a. Input the original image containing the steel number and the upper and lower edges of the metal plate, and perform the following pre - processing operations on it in sequence:
[0037] a1. Grayscale processing, converting the image into a grayscale image;
[0038] a2. Logarithmic transformation processing to reduce the contrast difference between the steel grade characters and the background while retaining the edge information of the metal sheet;
[0039] a3. Bilateral filtering processing to perform edge-preserving denoising on the grayscale image;
[0040] b. Using the Canny edge detection algorithm to extract the edge features from the preprocessed image to generate an edge image;
[0041] c. Using the probabilistic Hough line detection algorithm to perform line detection on the edge image, and performing length screening on all detected line segments, and only retaining the single target line with the longest length;
[0042] d. Calculating the inclination angle θ between the target line and the horizontal direction; performing rotation correction on the original image according to the inclination angle θ, and the rotation angle is -θ to horizontally align the edge of the metal sheet.
[0043] Furthermore, the method for making the first-stage rough positioning dataset is: using the LabelImg software to annotate the original image, and annotating multiple rough positioning detection frames in each image, and each rough positioning detection frame encloses a complete set of steel grade strings and the upper and lower edges of the metal sheet where they are located; the method for making the second-stage precise positioning dataset is: using the LabelImg software to annotate the image after the inclination correction processing, and each precise positioning detection frame only encloses a complete set of steel grade strings, and completely segments the steel grade from the image background.
[0044] Furthermore, the specific method for obtaining the second-stage precise positioning original image is: inputting the image data in the first-stage rough positioning dataset into the first-stage rough positioning model for prediction, and intercepting the area of the rough positioning detection frame after the image result of the rough positioning of the steel grade is output; the specific method for obtaining the steel grade automatic recognition image data is: inputting the image data in the second-stage precise positioning dataset into the second-stage precise positioning model for prediction, and intercepting the area of the precise positioning detection frame after the image result of the precise positioning of the steel grade is output.
[0045] Furthermore, the method for making the steel grade automatic recognition dataset is: manually identifying the steel grade strings in the steel grade automatic recognition image data, recording the actual steel grade corresponding to each image, and organizing and making it into a txt. format file.
[0046] Furthermore, the method for constructing the steel grade automatic recognition model based on the CRNN framework is:
[0047] Adopting the CRNN framework to construct a three-cascade structure, including:
[0048] a. Convolutional feature extraction module: A variant of Resnet, ResNet-34, is used to extract a multi-dimensional feature sequence from the input image. Among them, ResNet-34 contains 33 convolutional layers and 1 global pooling layer, and the total layer depth is 34 layers;
[0049] b. Sequence prediction module: A BiLSTM is used to receive the feature sequence output by the convolutional feature extraction module and generate a sequence prediction result through forward and backward dual-path processing;
[0050] c. Sequence transcription module: The CTC algorithm is used to decode the sequence prediction result into the final steel number character sequence;
[0051] Furthermore, the multi-directionally inclined steel number detection, positioning and automatic recognition system includes a primary positioning module, an image processing module, a secondary positioning module, an identification module and a verification module; among them, the verification module can compare the recognition result with the steel number in the metal sheet scheduling system. If they are the same, it is determined that the recognition is accurate, and the recognition result is recorded in the system. Otherwise, an alarm is issued to remind manual verification.
[0052] The beneficial effects of the present invention are as follows: Through the two-stage metal sheet steel number detection and positioning algorithm combining deep learning and image processing technology and the improved metal sheet steel number detection and positioning model, the accurate and complete detection of multi-directionally inclined metal sheet steel numbers in one image is realized; Through the deep learning model and system, the steel numbers of metal sheets in the production process are detected and recognized in real time for material tracking and recording, improving the accuracy and speed of metal sheet steel number recognition and recording, which is beneficial to the process of steel production and the management of metal sheet warehousing.
[0053] The specific effects are as follows: The MAP of the improved YOLOv11 model for the characteristics of metal sheets on the first-stage rough positioning dataset has been increased to 94.6%, which is 3.6% higher than the original YOLOv11; The trained weight parameter file has decreased from 5.2MB to 5MB, and the model is more lightweight; Moreover, the MAP of the second-stage precise positioning model for multi-directionally inclined steel numbers on the test set has reached 99%, and the recognition accuracy of the multi-directionally inclined steel number automatic recognition model on the test set has also reached 92%. Description of the Drawings
[0054] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention, where:
[0055] Figure 1 It is a flow chart of a method for detecting and recognizing multi-directionally inclined steel numbers based on deep learning according to an embodiment of the present invention;
[0056] Figure 2 is the steel grade image of the metal sheet collected in real time in the embodiment of the present invention;
[0057] Figure 3 is the schematic diagram of the original model structure for steel grade detection and positioning based on the improved YOLOv11 in the embodiment of the present invention;
[0058] Figure 4 is the schematic diagram of the RMC module structure designed in the embodiment of the present invention;
[0059] Figure 5 is the schematic diagram of the CP-DSA module structure designed in the embodiment of the present invention;
[0060] Figure 6 is the schematic diagram of the process for steel grade image skew correction in the embodiment of the present invention;
[0061] Figure 7 is the detection effect diagram of the rough positioning of the first section of the steel grade and the precise positioning of the second section of the steel grade in the embodiment of the present invention. Detailed implementation manners
[0062] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; for those skilled in the art, it is understandable that some well-known structures or steps and their descriptions in the drawings may be omitted in order to better illustrate the embodiments of the present invention.
[0063] The following further elaborates on the present invention in conjunction with the drawings and embodiments:
[0064] As Figure 1 shown, a multi-directional skew steel grade detection and recognition method based on deep learning disclosed in the embodiment of the present invention includes the following steps:
[0065] S1: Collect the original images of the steel grades of multi-directionally skewed metal sheets and make the first-stage rough positioning data set.
[0066] In this embodiment, first, the original images of the steel grades are collected by an industrial camera on the metal sheet production line. During the collection process, the industrial camera is fixed above the production line and takes pictures of the passing metal sheets. Since the metal sheets may have position offsets and angular skews during the transmission process, the collected steel grade images mostly show a multi-directionally skewed state, as Figure 2As shown, there are many problems with the original image of the steel grade of the collected metal sheet, such as complex image background, small contrast between the steel grade string and the background, small proportion of some steel grades in the image, problems in detecting small targets, variable styles of steel grade characters and different string lengths, and multiple steel grades with multi-directional inclinations in the same image. Moreover, since the detection boxes of the object detection algorithm based on deep learning are mostly horizontal detection boxes, when the steel grade text information is completely detected and located, the background near the steel grade is easily included in the horizontal detection box at the same time. The resulting localization detection result image has poor image quality for subsequent automatic recognition of the steel grade, and the inclination angle of the steel grade in the image also affects the automatic recognition effect of the steel grade. Therefore, based on the above problems, the basic idea of the two-stage steel grade localization detection algorithm provided by the present invention is to first perform a rough localization of the steel grade on the original metal sheet steel grade image with a complex background, separate it from the image with multiple steel grades, perform inclination correction on the image containing only one string of steel grade after the first-stage localization detection, and segment the steel grade from the background near it through the second-stage localization detection, ultimately achieving the precise localization of the steel grade.
[0067] After obtaining the original image of the steel grade, use the LabelImg software to annotate it to make a data set for the first-stage rough localization of multi-directionally inclined steel grades. During the annotation process, multiple rough localization detection boxes are annotated in one image, and each rough localization detection box encloses a complete string of steel grade characters and the upper and lower edges of the metal sheet where it is located. This annotation method ensures that the upper and lower edge information of the metal sheet can be completely retained after rough localization, providing a basis for subsequent inclination correction. After the annotation is completed, the obtained annotated samples constitute a data set for the first-stage rough localization of multi-directionally inclined steel grades.
[0068] S2: Improve YOLOv11 based on the problems existing in the first-stage rough localization data set, construct an original model for steel grade detection and localization based on the improved YOLOv11, and train it using the first-stage rough localization data set to obtain a first-stage rough localization model for multi-directionally inclined steel grades.
[0069] Regarding the problems existing in the first-stage rough detection and localization data set, such as Figure 3 As shown in the schematic diagram of the model structure, in this embodiment, the following specific improvements are made to the YOLOv11 model:
[0070] 1) Aiming at the model lightweight requirement for real-time steel grade detection and localization and the problems of complex image background, variable styles of steel grade characters and different string lengths in the steel grade detection and localization data set, a new multi-scale convolution module - Residual Mixed Convolution Block (RMC) is designed, and its structural schematic diagram is as Figure 4 shown. The structural design of the RMC module is as follows: the input feature F inputInitial feature extraction is performed through DWConv convolution with a 3x3 convolutional kernel to capture the spatial dimension information of the input features. The expression is:
[0071] F D3 = Dwconv 3×3 (F input )
[0072] In the formula, F D3 is the feature map after 3×3 depthwise separable convolution, and Dwconv 3×3 is the 3×3 depthwise separable convolution operation. The depthwise separable convolution consists of two parts: depthwise convolution and pointwise convolution. Compared with ordinary convolution, it has fewer parameters and computational complexity.
[0073] Then, the RMC module passes through three parallel DWConv with 3x3 convolutional kernels, PConv with 3x3 convolutional kernels, and the Adown module to further extract features from the feature map F D3 . The PConv convolution can reduce redundant calculations while retaining the ability of conventional convolution to extract spatial features. The Adown module reduces the complexity of the model by reducing the number of parameters. At the same time, the RMC also has a residual structure. The residual structure branch inputs the feature F input and performs feature extraction through DWConv convolution with a 5x5 convolutional kernel. Finally, these features are fused together to generate a richer feature representation. The output feature maps of these parts are concatenated in the channel dimension. The expression is:
[0074] F ′ D3 = Dwconv 3×3 (F D3 )
[0075] F P3 = Pconv 3×3 (F D3 )
[0076] F A3 = Adown(F D3 )
[0077] F D5 = Dwconv 5×5 (F input )
[0078] F C = concat(F ′ D3 , F P3 , F A3 , F D5 )
[0079] In the formula, F′ D3 To transform the feature map F D3 after further 3×3 depthwise separable convolution, F P3 To transform the feature map F D3 after further 3×3 partial convolution, F A3 To transform the feature map F D3 after 3×3 downsampling, F C To transform the feature map F ′ D3 and F P3 and F A3 and F D5 into the multi-scale feature map after splicing and mixing, where Pconv 3×3 is the 3×3 partial convolution operation, Adown is the downsampling operation, Dwconv 5×5 is the 5×5 depthwise separable convolution operation, concat() is the splicing operation. The 3x3 convolution kernel is suitable for capturing details, while the 5x5 convolution kernel is better at capturing large-scale context information. This multi-scale convolution can capture multiple patterns of the input features and improve the feature expression ability of the model. And the residual structure helps to solve the gradient vanishing problem in deep networks and allows information to flow more easily in the network.
[0080] Finally, the RMC module adjusts the number of output channels through a standard convolution with a convolution kernel of 1. The expression is:
[0081] F output = Conv 1×1 (F C )
[0082] In the formula, F output is the final output feature map, Conv 1×1 is the 1×1 ordinary convolution operation, which rearranges the spliced features to form new features. The new features continuously update the weights under the constraint of the loss function and continuously generate more compliant features, and the new features fuse multiple features concatenated by concat.
[0083] Finally, it passes through the BN normalization layer and the SiLU activation function;
[0084] The expression formula of the batch normalization layer (BN) is:
[0085]
[0086]
[0087] where γ and β are two parameters to be learned, and α is a batch of data {x1, x2,..., x m}, the samples {y1, y2,..., y after batch normalization m}; The BN layer is used to readjust the data distribution after the convolutional layer and perform normalization processing, which can reduce the risk of gradient vanishing or gradient explosion.
[0088] The expression formula of the SiLU activation function is:
[0089] f(x) = x sigmoid(x)
[0090]
[0091] The SiLU function is continuously differentiable, and its gradient propagation has the advantages of being smooth and continuous. Moreover, SiLU is a non-linear activation function, which can help the neural network learn complex patterns and features.
[0092] 2) Improve the ordinary convolutions in the CBSModule of the feature extraction network Backbone and the feature fusion network Neck in the YOLOv11 model to RMC modules. This lightweight multi-scale convolution can capture various patterns of the input features while reducing the number of model parameters and improving the feature expression ability of the model. At the same time, the CBSModule in the YOLOv11 model consists of a 3×3 ordinary convolution, a BN layer, and a SiLU activation function. By improving the convolutional layer to a multi-scale residual convolution structure, depthwise separable convolution, partial convolution, and downsampling modules are used to extract more and deeper features while reducing the parameters of the model.
[0093] 3) Aiming at the small target detection problem caused by the small contrast between the steel grade string and the background, its different proportions in the image, and the small proportion of steel grade characters in some images, design a Channel Prior-Dual Spatial Attention mechanism (CP-DSA). Its structural schematic diagram is as Figure 5 shown. The CP-DSA attention mechanism combines the CPCA attention mechanism and the CA attention mechanism, dynamically allocates attention weights in the channel and spatial dimensions, pays attention to and extracts features from aspects such as spatial coordinates at the same time, and splices and fuses their outputs.
[0094] Among them, the CPCA attention mechanism can comprehensively consider channel and spatial information and dynamically allocate weights in the spatial dimension. First, it takes the given feature map X as input, and the channel attention module will calculate a one-dimensional channel attention map M C , and then multiply M C element-wise with the input feature X to obtain the channel attention refined feature X C , and the expression is:
[0095]
[0096] Its channel attention module collects the spatial information of the feature map by performing average pooling and max pooling, and transmits it to a shared multi-layer perceptron for processing. The expression is:
[0097] CA(X) = σ(MLP(Avgpool(X)) + MLP(Maxpool(X)))
[0098] In the formula, MLP represents multi-layer perceptron processing, Avgpool() represents average pooling operation, Maxpool represents max pooling operation, and σ represents the Sigmoid function; then X is processed through spatial attention C to generate a spatial attention map M S , and the final result X is obtained by multiplying M S and X C . Its expression is: CPCA
[0099]
[0100] Among them, spatial attention captures the spatial relationship between features by using depth convolution, and uses multi-scale results to further enhance the ability of convolution operation to capture spatial relationships. A 1×1 convolution layer is introduced to mix channels and fuse the feature information obtained from each scale. The expression is:
[0101]
[0102] In the formula, Dconv represents depth convolution operation, Branch i represents the i-th branch.
[0103] Among them, the CA attention mechanism can reassign and combine the weights in the horizontal and vertical directions, focus on the feature information of the steel number along one spatial direction, and retain the exact position information of the steel number along the other spatial direction. It can effectively enhance the feature channels containing steel number information and suppress the channels containing complex information. Finally, a feature map is generated for separate encoding to form a feature map X that is sensitive to direction perception and position CA .
[0104] The output features of the CPCA attention mechanism and the CA attention mechanism are spliced and fused. The expression is:
[0105] X CP-DSA = concat(X CPCA , X CA )
[0106] Two attention mechanisms run in parallel, and then the CP-DSA module adjusts the number of output channels through a standard convolution with a convolution kernel of 1. The expression is as follows:
[0107] X output = Conv 1×1 (X CP-DSA )
[0108] In the formula, X output is the final output feature map.
[0109] 4) Add the CP-DSA attention mechanism module after the first C3K2 module in the feature extraction network Backbone of the YOLOv11 model. Through the fusion of the features output by the CPCA attention mechanism and the CA attention mechanism, improve the feature extraction ability for the target, and then improve the detection performance for the target.
[0110] 5) Change the CIoU loss function in the output prediction network Head of the YOLOv11 model to the WIoU loss function. The calculation of the CIoU loss function is relatively complex, which may lead to a large computational overhead during the training process, and the WIoU loss function can more accurately evaluate the positional relationship between the predicted target detection box and the marked detection box.
[0111] Wise-IOU calculates the IOU loss in the category prediction loss using a dynamic method. The specific formula is as follows:
[0112]
[0113] L IOU = 1 - IOU
[0114] L WIOU = R WIOU L IOU
[0115]
[0116] In the formula, A represents the area of the predicted detection box, B represents the area of the ground truth detection box, IoU represents the intersection over union, L IOU represents the intersection over union loss. The ground truth detection box is B = [x y w h], and the target detection box is B gt = [x gt y gt w gt h gt , and W g and H g represent the size of the smallest bounding box.
[0117] In an embodiment, the specific process of training the original model for steel grade detection and positioning based on the improved YOLOv11 is as follows: The labeled first-stage rough positioning dataset is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1. 482 samples are used for the training of the original model for steel grade detection and positioning, 61 samples are used for the validation of the model, and 61 samples are used for the testing of the model. Finally, a first-stage rough positioning model for multi-directionally inclined steel grades is obtained.
[0118] The YOLOv11 model is trained using the same dataset and parameters, compared with the improved YOLOv11 model, and the performance of the model is evaluated. The MAP of the YOLOv11 model on the first-stage rough positioning dataset is 91%, while the MAP of the improved YOLOv11 model has increased to 94.6%, an increase of 3.6%. Moreover, the weight parameter file after training has also decreased from 5.2MB to 5MB, and the FPS gap is relatively small, changing from the original 54.6 to 53.7.
[0119] S3: Use the first-stage rough positioning model for multi-directionally inclined steel grades to obtain the original image for the second-stage precise positioning. After performing tilt correction processing using image processing techniques, a second-stage precise positioning dataset is produced.
[0120] When this step is implemented, the specific method for obtaining the original image for the second-stage precise positioning of multi-directionally inclined steel grades is as follows: The image data in the first-stage rough positioning dataset is input into the first-stage rough positioning model for prediction. After the image result of the rough positioning of the steel grade is output, the area of the rough positioning detection box is intercepted. Among them, the rough positioning detection box in the image result of the first-stage rough positioning contains the complete steel grade and the upper and lower edge information of the metal plate where it is located. In this embodiment, the basic principle of the tilt correction processing is: the tilt angle of the steel grade is judged by the upper and lower edge straight lines of the steel plate on the steel sheet, and rotation correction is performed according to the tilt angle. The process of the image processing technique is as follows: First, the image is grayscale and logarithmically transformed to reduce the contrast between the character information and the background in the image while retaining the upper and lower edge information of the steel plate. Then, the image is denoised by bilateral filtering and edge detection is performed using the Canny edge detection algorithm. After that, probabilistic Hough line detection is used to detect the straight lines in the image. The lengths of all the straight lines detected in one image are screened, and only the longest straight line is retained. Finally, the angle of this straight line is detected, and the image is rotationally corrected according to the detected straight line angle to make the edges of the metal plate horizontally aligned. At the same time, the method for producing the second-stage precise positioning dataset is: Use the LabelImg software to annotate the precise positioning detection boxes for the images after the tilt correction processing. Each precise positioning detection box only encloses a group of complete steel grade strings, completely separating the steel grade from the image background.
[0121] S4: Use the second-stage precise positioning dataset to train the original steel grade detection and positioning model to obtain a multi-directionally inclined steel grade second-stage precise positioning model.
[0122] When implementing this step, the specific process of using the second-stage precise positioning dataset to train the original steel grade detection and positioning model is as follows: Divide the labeled multi-directionally inclined steel grade second-stage precise positioning dataset into a training set, a validation set, and a test set at a ratio of 8:1:1. Use 2347 samples for training the original steel grade detection and positioning model, 293 samples for validating this model, and 293 samples for testing this model. Finally, obtain a multi-directionally inclined steel grade second-stage precise positioning model, and the MAP of this model on the test set reaches 99%.
[0123] S5: Use the multi-directionally inclined steel grade second-stage precise positioning model to obtain steel grade automatic recognition image data to make a steel grade automatic recognition dataset.
[0124] When implementing this step, the specific process of obtaining steel grade automatic recognition image data is as follows: Input the image data in the second-stage precise positioning dataset into the second-stage precise positioning model for prediction, and intercept the area of the precise positioning detection frame after the image result of the steel grade precise positioning is output. Among them, the detection frame in the image result of the steel grade precise positioning contains complete steel grade string information. At the same time, the method of making a steel grade automatic recognition dataset is: Manually recognize the steel grade strings in the steel grade automatic recognition image data, record the actual steel grade corresponding to each image, and organize and make it into a txt format file.
[0125] S6: Build a steel grade automatic recognition model based on the CRNN framework and use the steel grade automatic recognition dataset to train it to obtain a multi-directionally inclined steel grade automatic recognition model.
[0126] When implementing this step, the structure of the steel grade automatic recognition model based on the CRNN framework is divided into three layers, namely the convolutional layer, the recurrent layer, and the transcription layer. The convolutional layer extracts the feature map and feature sequence of the image, the recurrent layer practices processing sequence information and makes predictions on the sequence, and the transcription layer integrates and transforms the predicted sequence into a final sequence result; the convolutional layer uses a variant of Resnet, Resnet-34. The Resnet-34 network has a depth of 34 layers, including 33 convolutional layers and 1 global pooling layer. The recurrent layer uses BILSTM, and the transcription layer uses CTC.
[0127] The training of the steel grade automatic recognition dataset using the above-mentioned text recognition model based on the CRNN framework is specifically as follows: The steel grade automatic recognition dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. The number of images in the training set is 2,347, and the number of images in the validation set and the test set is 293. The dataset is augmented by randomly rotating and adding noise, and a total of 7,041 images are used as the training set. The dataset is also normalized in terms of size. The labeled dataset is trained to obtain a multi-directionally inclined steel grade automatic recognition model, and the recognition accuracy of this model on the test set reaches 92%.
[0128] S7: Deploy the multi-directionally inclined steel grade first-stage rough positioning model, the second-stage precise positioning model, and the automatic recognition model on-site. Combine the image processing technology to construct a multi-directionally inclined steel grade detection, positioning, and automatic recognition system, and according to the steel grade images collected in real time, output the recognition results of the steel grade character sequence in real time.
[0129] When this step is implemented, deploy the multi-directionally inclined steel grade first-stage rough positioning model, the second-stage precise positioning model, and the automatic recognition model on-site, and combine the image processing technology to run the real-time steel grade detection and positioning model and the recognition model on the workstation computer to construct a multi-directionally inclined steel grade detection, positioning, and automatic recognition system. This system includes a primary positioning module, an image processing module, a secondary positioning module, a recognition module, and a verification module, and can accurately detect and recognize the steel grade images collected in real time. When the system receives the steel grade images collected in real time, it stores the images and inputs the images into the primary positioning module, and inputs the output result into the image processing module. Similarly, input the output result of the previous step into the secondary positioning module, and then input the output result of the secondary positioning module into the recognition module. Finally, compare the recognition result with the steel grade in the metal sheet scheduling system through the verification module. If they are consistent, the recognition is accurate, and the recognition result is directly recorded into the system. If they are inconsistent, an alarm is popped up to remind manual verification, and the manual determination and correction are carried out. After that, the system records the verification result into the system again.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A multi-directional inclined steel number detection and recognition method based on deep learning, characterized in that, It includes the following steps: S1: Collect the original images of the steel grades of multi-directionally inclined metal sheets and make the first-stage rough positioning dataset; S2: Improve YOLOv11 based on the problems existing in the first-stage rough positioning dataset, construct the original model for steel grade detection and positioning based on the improved YOLOv11, and train it with the first-stage rough positioning dataset to obtain the first-stage rough positioning model for multi-directionally inclined steel grades; S3: Use the first-stage rough positioning model for multi-directionally inclined steel grades to obtain the original images for the second-stage precise positioning. After performing tilt correction processing using image processing techniques, make the second-stage precise positioning dataset; S4: Train the original model for steel grade detection and positioning with the second-stage precise positioning dataset to obtain the second-stage precise positioning model for multi-directionally inclined steel grades; S5: Use the second-stage precise positioning model for multi-directionally inclined steel grades to obtain the steel grade automatic recognition image data to make the steel grade automatic recognition dataset; S6: Construct a steel grade automatic recognition model based on the CRNN framework and train it with the steel grade automatic recognition dataset to obtain the multi-directionally inclined steel grade automatic recognition model; S7: Deploy the first-stage rough positioning model, the second-stage precise positioning model, and the automatic recognition model for multi-directionally inclined steel grades on-site. Combine the image processing techniques to construct a multi-directionally inclined steel grade detection, positioning, and automatic recognition system, and according to the real-time collected steel grade images, output the recognition results of the steel grade character sequence in real time.
2. The multi-directional inclined steel number detection and recognition method based on deep learning according to claim 1, characterized in that The improved YOLOv11 includes: (1) Design a multi-scale convolution module - residual hybrid convolution module, that is, the RMC module, including the following connected in sequence: a. The primary feature extraction layer uses DWConv with the first preset size to perform spatial feature extraction on the original input features; b. The multi-branch parallel processing layer includes the following acting on the output of the primary feature extraction layer simultaneously: b1. The DWConv branch with the second preset size, b2. The PConv branch with the second preset size, b3. The parameter compression type Adown branch; c. The cross-layer residual processing layer uses DWConv with the third preset size to perform cross-layer feature extraction on the original input features; d. The feature fusion layer performs channel stacking on the output features of each branch of the multi-branch parallel processing layer and the output features of the cross-layer residual processing layer, and performs dimension matching through the channel adjustment convolution layer; e. The post-processing layer sequentially performs BN normalization processing and SiLU activation function transformation on the output features of the feature fusion layer; Among them, the third preset size is larger than the first preset size, and the number of parameters of the Adown branch is less than that of the conventional convolution structure; (2) Improve the ordinary convolutions in the feature extraction network Backbone and the feature fusion network Neck CBSModule in YOLOv11 to the RMC module to form a multi-level enhanced feature processing architecture.
3. A multi-directional inclined steel number detection and recognition method based on deep learning according to claim 1, characterized in that The improved YOLOv11 includes embedding a channel priority - dual spatial attention mechanism, that is, the CP-DSA module, at the preset hierarchical nodes of the feature extraction network. This module includes: a. Channel priority processing branch, generating a first attention feature map through the CPCA attention mechanism; b. Spatial coordinate attention branch, generating a second attention feature map through the CA attention mechanism; c. Feature fusion unit, performing cross-dimensional splicing and fusion on the first attention feature map and the second attention feature map, and finally outputting an enhanced attention feature map; Among them, the CP-DSA module is configured in the primary feature enhancement stage of the feature extraction network, specifically connected to the output end of the first multi-scale convolution module of the feature extraction network, and enhancing the multi-dimensional feature expression ability by jointly optimizing channel sensitivity and spatial position correlation.
4. A method for multi-directional inclined steel number detection and recognition based on deep learning according to claim 1, characterized in that, The improved YOLOv11 includes improving the original CIoU loss function in the YOLOv11 output prediction network Head to the WIoU loss function to more economically and accurately evaluate the positional relationship between the predicted target detection box and the marked detection box.
5. A multi-directional inclined steel number detection and recognition method based on deep learning according to claim 1, characterized in that, The principle of the tilt correction process is: judging the tilt angle of the steel number through the straight lines of the upper and lower edges of the metal plate, and performing rotation correction according to the tilt angle; the specific image processing technology is: a. Input the original image containing the steel number and the upper and lower edges of the metal plate, and sequentially perform the following preprocessing on it: a1. Grayscale processing, converting the image into a grayscale image; a2. Logarithmic transformation processing, reducing the contrast difference between the steel number characters and the background, and at the same time retaining the edge information of the metal plate; a3. Bilateral filtering processing, performing edge-preserving denoising on the grayscale image; b. Extracting edge features from the preprocessed image using the Canny edge detection algorithm to generate an edge image; c. Performing line detection on the edge image using the probabilistic Hough line detection algorithm, screening the lengths of all detected line segments, and only retaining the single target line with the longest length; d. Calculating the tilt angle θ between the target line and the horizontal direction; performing rotation correction on the original image according to the tilt angle θ, and the rotation angle is -θ to horizontally align the edges of the metal plate.
6. A method for multi-directional inclined steel number detection and recognition based on deep learning according to claim 1, characterized in that, The method for making the first-stage rough positioning dataset is: using the LabelImg software to annotate the original image, annotating multiple rough positioning detection boxes in each image, and each rough positioning detection box encloses a complete set of steel number strings and the upper and lower edges of the metal plate where they are located; the method for making the second-stage precise positioning dataset is: using the LabelImg software to annotate the image after the tilt correction process, and each precise positioning detection box only encloses a complete set of steel number strings, completely segmenting the steel number from the image background.
7. A method for multi-directional inclined steel number detection and recognition based on deep learning according to claim 6, characterized in that, The specific method for obtaining the original image of the second-stage precise positioning is: inputting the image data in the first-stage rough positioning dataset into the first-stage rough positioning model for prediction, and intercepting the area of the rough positioning detection box after the image result of the rough positioning of the steel number is output; the specific method for obtaining the image data of the automatic steel number recognition is: inputting the image data in the second-stage precise positioning dataset into the second-stage precise positioning model for prediction, and intercepting the area of the precise positioning detection box after the image result of the precise positioning of the steel number is output.
8. A method for multi-directional inclined steel number detection and recognition based on deep learning according to claim 1, characterized in that The method for making the automatic steel grade recognition dataset is as follows: Manually recognize the steel grade strings in the automatic steel grade recognition image data, record the actual steel grade corresponding to each image, and organize them into a txt format file.
9. A method for multi-directional inclined steel number detection and recognition based on deep learning according to claim 1, characterized in that The method for constructing the automatic steel grade recognition model based on the CRNN framework is as follows: Use the CRNN framework to construct a three-cascade structure, including: a. Convolutional feature extraction module: Use a variant of Resnet, ResNet-34, to extract a multi-dimensional feature sequence from the input image. Among them, ResNet-34 contains 33 convolutional layers and 1 global pooling layer, and the total layer depth is 34 layers; b. Sequence prediction module: Use BiLSTM to receive the feature sequence output by the convolutional feature extraction module, and generate a sequence prediction result through forward and backward dual-path processing; c. Sequence transcription module: Use the CTC algorithm to decode the sequence prediction result into the final steel grade character sequence.
10. A multi-directional inclined steel number detection and recognition method based on deep learning according to claim 1, characterized in that The multi-directionally inclined steel grade detection, positioning and automatic recognition system includes a primary positioning module, an image processing module, a secondary positioning module, an identification module and a proofreading module; among them, the proofreading module can compare the recognition result with the steel grade in the metal sheet scheduling system. If they are consistent, it is determined that the recognition is accurate, and the recognition result is recorded in the system. Otherwise, an alarm is issued to remind manual proofreading.
Citation Information
Patent Citations
High-temperature slab ID character recognition method and system based on deep learning
CN117173716A
Hot casting blank number identification method and system based on deep learning
CN117894004A
Deep learning photovoltaic panel identification method based on multi-scale Retinex algorithm
CN118298324A
Cited By
Deep learning-based method and system for inspecting embedding box before batch dehydration
CN120877307A