Natural Scene Text Detection and Recognition Method Based on YOLOV5

The YOLOV5-based method addresses the challenge of large model parameters and deep semantic understanding limitations by using lightweight features and attention mechanisms, enabling efficient natural scene text detection and recognition on mobile devices.

CN115205839BActive Publication Date: 2025-07-15FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210785742.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-07-15
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

In the prior art, traditional visual model parameters are too large and cannot effectively understand the deep semantic information of the image, cannot be deployed on the mobile side, and requires a large number of artificially synthesized fake data sets for training.

Method used

A lightweight feature extractor based on YOLOV5 is adopted, combining cross-layer connections and spatial pyramid pooling layer, the length-width ratio of the anchor box is fitted using the Kmeans algorithm, deformation convolution processing features are added, and text features are aligned using bidirectional LSTM and hierarchical attention mechanism to predict text sequences.

Benefits of technology

It realizes lightweight natural scene text detection and recognition, can be deployed on the mobile side, improves the accuracy of detection and recognition, reduces model parameters, and solves the problem that traditional models cannot understand deep semantic information in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205839B_ABST
    Figure CN115205839B_ABST
Patent Text Reader

Abstract

The present invention proposes a natural scene text detection and recognition method based on YOLOV5, including: Step S1: Obtain a natural scene text image dataset and convert the corresponding labels into the format required by YOLOV5; Step S2: Use the lightweight feature extractor of YOLOV5 to extract the position information and deep semantic information of the image text; Combine the shallow features and deep features by using cross-layer connection and spatial pyramid pooling layer; Add deformable convolution in the cross-layer connection so that the network can better handle the change of the feature map scale; Step S3: Use the anchor boxes aggregated by the Kmeans algorithm to fit the aspect ratio of the real text box and predict the deviation between the anchor box and the real box; Use long convolution to process the features to make the aspect ratio of the anchor box more conform to the real text box; Step S4: Use bidirectional LSTM and attention mechanism to align the text features and predict the text sequence; It can realize the detection and recognition of natural scene text by using deep learning, and is lightweight enough to be deployed on mobile devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision understanding, and in particular to a natural scene text detection and recognition method based on YOLOV5. Background Art

[0002] In recent years, artificial intelligence technology has developed rapidly. Using deep learning to process some natural scene texts in our lives, that is, natural scene text detection and recognition has become a popular technology. Natural scene text detection and recognition is a very important research field in the fields of computer vision and artificial intelligence. It mainly studies whether a machine can correctly understand a picture, so as to complete the detection and recognition of the targets in the picture. Summary of the Invention

[0003] The present invention proposes a natural scene text detection and recognition method based on YOLOV5. The present invention can realize the detection and recognition of natural scene texts by using deep learning, and the lightweight of this method is sufficient to realize deployment on mobile devices.

[0004] The present invention specifically adopts the following technical solutions:

[0005] A natural scene text detection and recognition method based on YOLOV5, comprising the following steps;

[0006] Step S1: Obtain a natural scene text image data set, and convert the corresponding labels into the format corresponding to YOLOV5;

[0007] Step S2: Use the lightweight feature extractor of YOLOV5 to extract the position information and deep semantic information of the image text; combine the shallow features and deep features by using cross-layer connection and spatial pyramid pooling layer; add deformable convolution in the cross-layer connection so that the network can better process the change of the feature map scale;

[0008] Step S3: Use the anchor boxes aggregated by the Kmeans algorithm to fit the aspect ratio of the real text box, and predict the deviation between the anchor box and the real box; use long convolution to process the features to make the aspect ratio of the anchor box more conform to the real text box;

[0009] Step S4: Use bidirectional LSTM and attention mechanism to align text features and predict text sequences.

[0010] Furthermore, step S1 specifically includes the following steps;

[0011] Step S11: Obtain a public natural scene text data set;

[0012] Step S12: Convert all the label formats in the data set into the format required by YOLOV5;

[0013] Step S13: Record the corresponding text in the text area of the dataset into the json file for convenient subsequent recognition. Further, step S2 specifically includes the following steps;

[0014] Step S21: Input the images in batches into a feature extractor composed of multiple Conv modules and multiple BottleneckCSP modules, where the Conv module includes a convolutional layer with a convolution kernel size of 3×3, a batch normalization layer BN, and a SiLU activation function; as shown in Formula 1:

[0015] F Conv_out = SiLU(BN(Conv 3×3 (F Conv_in )))

[0016] Formula 1;

[0017] where F Conv_in is the input feature of the Conv module, Conv 3×3 is the convolutional layer with a convolution kernel size of 3×3;

[0018] The BottleneckCSP module is composed of Bottleneck plus CSP; Bottleneck passes the input feature through a convolutional layer with a convolution kernel size of 1×1, then through a convolutional layer with a convolution kernel size of 3×3, and then adds the input feature to it; as shown in Formula 2, where F Bottleneck is the output of the Bottleneck module, F Bottleneck_in is the input feature of the Bottleneck module, Conv 3×3 is the convolutional layer with a convolution kernel size of 3×3, Conv 1×1 is the convolutional layer with a convolution kernel size of 1×1;

[0019] F Bottleneck = F Bottleneck_in + Conv 3×3 (Conv 1×1 (F Bottleneck_in ))

[0020] Formula 2;

[0021] CSP divides the original input into two branches, performs convolutional operations on each branch to halve the number of channels, then one branch performs Bottleneck×N operations, where N is a custom parameter, and then Concat the two branches so that the input and output of BottlenneckCSP are of the same size; as shown in Formula 3:

[0022] F Concat = Concat(N×Bottleneck(Conv 1×1 (Fin_c / 2_1 ))), Conv 3×3 (F in_c / 2_2 ))) Formula 3;

[0023] Where F Concat is the result of concatenating two branches, Concat is the feature concatenation operation, Bottleneck refers to the operation of Formula 2, F in_c / 2_1 and F in_c / 2_2 represent two branches of the input features, and the number of channels is half of the original input features;

[0024] Then, pass F Concat through the batch normalization layer BN, the LekyReLU activation function, and Conv 1×1 to obtain the output F BottleneckCSP of BottlenneckCSP, as shown in Formula 4:

[0025] F BottleneckCSP = Conv 1×1 (LekyReLU(BN(F Concat ))))

[0026] Formula 4;

[0027] Step S22: Input the features that have been downsampled 32 times by the Conv module and BottleneckCSP into the SPP spatial pyramid pooling layer module, perform max pooling operations on the feature maps of different sizes, and then concatenate the pooled features as the output of the feature extractor; as shown in Formula 5:

[0028] F SPP_out = DeformableConv(Concat(F SPP_in , MaxPooling 13×13 (F SPP_in ),

[0029] MaxPooling 9×9 (F SPP_in ), MaxPooling 5×5 (F SPP_in ))))

[0030] Formula 5;

[0031] Where F SPP_in is the input feature of the SPP module, F SPP_out is the output of the SPP module, MaxPooling 13×13 , MaxPooling 9×9 , MaxPooling 5×5They represent the maximum pooling layers with sampling kernel sizes of 13×13, 9×9, and 5×5 respectively, and DeformableConv is the deformable convolution module.

[0032] Furthermore, step S3 specifically includes the following steps:

[0033] Step S31: Use the Kmeans algorithm to fit the aspect ratios of the true text boxes, input the ratios of all true text boxes into Kmeans to cluster the aspect ratios of multiple anchor boxes;

[0034] Step S32: Predict the deviation between the anchor box and the true text box using the features extracted by the feature extractor; First, pass the features through a 1×7 long convolutional network to extract semantic features suitable for long texts; Then divide the processed features into gridn×gridn grids, where gridn is a custom parameter; The network will predict four offsets t x1 , t y1 , t h1 , t w1 , and the calculation methods are shown in Formulas 6, 7, 8, and 9:

[0035] t x1 = log((bbox x2 - c x3 ) / (1 - (bbox x2 - c x3 ))) Formula 6;

[0036] t y1 = log((bbox y2 - c y3 ) / (1 - (bbox y2 - c y3 ))) Formula 7;

[0037] t h1 = log(gt h4 / p h5 ) Formula 8;

[0038] t w1 = log(gt w4 / p w5 ) Formula 9;

[0039] Where bbox x2 , bbox y2 represent the horizontal and vertical coordinates of the center point of the true text box respectively; c x3 , c y3 represent the horizontal and vertical coordinates of the upper left corner of the grid corresponding to the true text box; gt h4 , gt w4 represent the height and width of the true text box; ph5 , p w5 represent the height and width of the anchor box; the network predicts the position of the text box by predicting these 4 offsets.

[0040] Further, step S4 specifically includes the following steps;

[0041] Step S41: Process the long semantic features using a hierarchical attention mechanism. The hierarchical attention mechanism is implemented through three matrices, including a query matrix Q, a key matrix K, and a value matrix V; and load the word embeddings of the predicted text features into matrix E, and linearly map matrix E into the query matrix Q, the key matrix K, and the value matrix V; multiply the query matrix Q by the key matrix K to perform a score evaluation for each pixel in the feature map; where the high or low score represents the degree of tightness of the association between two feature pixels; then divide the obtained score by the square root of the dimension dim of the key vector to enhance the stability of the gradient; then use the softmax function to make the scores of all words positive and their sum equal to 1; finally, multiply the obtained LekyReLU scores by the value matrix V to obtain the output of the attention layer, denoted here as matrix O; as shown in Equation ten:

[0042]

[0043] Step S42: Input O into a bidirectional LSTM to align the text features with the text and predict the final text result.

[0044] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the above-mentioned YOLOV5-based natural scene text detection and recognition method when executing the program.

[0045] A computer-readable storage medium has a computer program stored thereon, wherein the program implements the above-mentioned YOLOV5-based natural scene text detection and recognition method when executed by a processor.

[0046] Compared with the prior art, the present invention and its preferred solutions have the following beneficial effects:

[0047] 1. The constructed YOLOV5-based natural scene text detection and recognition method has a very lightweight model and fast inference speed compared with other existing methods, and can be deployed on mobile devices.

[0048] 2. The dataset does not require a large number of annotation files. Using the pre-trained model provided by YOLOV5 official, a text detection model with good performance can be trained.

[0049] 3. The hierarchical attention mechanism can mimic the phenomenon of attention concentration caused by humans observing things, enabling it to understand the hidden relationships in images and texts by associating local features or ignoring some useless features, and solving the problem that ordinary attention cannot focus on long text features.

[0050] 4. Using methods such as data augmentation, data enhancement, and model integration can further optimize the performance of our detection and recognition models, and the accuracy rate can be further improved.

[0051] In view of the problems that the traditional vision model contains too many parameters and cannot understand the deep semantic information of images, etc., the present invention proposes a method based on YOLOV5. Using the idea of the pre-trained model, it effectively solves the problem that a large number of artificially synthesized false data sets are required for training the model, and due to the idea of its lightweight model, the model can be deployed on mobile devices.

[0052] The present invention utilizes the hierarchical attention mechanism, which can mimic the phenomenon of attention concentration caused by humans observing things, effectively extracts the hidden connections inside images and texts, reduces the model parameters, and uses bidirectional LSTM to align text features and text content. Brief Description of the Drawings

[0053] The following further elaborates on the present invention in conjunction with the drawings and specific embodiments;

[0054] Figure 1 It is a schematic diagram of the process and working principle of the embodiment of the present invention. Detailed Embodiments

[0055] To make the features and advantages of this patent more obvious and understandable, specific embodiments are hereby given and detailed descriptions are as follows:

[0056] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0057] It should be noted that the terms used here are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to this application. As used here, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0058] As Figure 1 shown, the natural scene text detection and recognition method based on YOLOV5 provided in this embodiment includes the following steps;

[0059] Step S1: Obtain a natural scene text image dataset and convert the corresponding labels into the format required by YOLOV5;

[0060] Step S2: Use the lightweight feature extractor of YOLOV5 to extract the location information and deep semantic information of the image text; Combine the shallow features and deep features using cross-layer connections and spatial pyramid pooling layers; Add deformable convolutions in the cross-layer connections so that the network can better handle changes in the feature map scale;

[0061] Step S3: Use the Kmeans algorithm to aggregate the aspect ratios of the anchor boxes to fit the aspect ratios of the real text boxes, and predict the deviation between the anchor boxes and the real boxes; Use long convolutions to process the features to make the aspect ratios of the anchor boxes more consistent with the real text boxes;

[0062] Step S4: Use bidirectional LSTM and attention mechanism to align text features and predict text sequences.

[0063] The solution of this embodiment can achieve the detection and recognition of natural scene text using deep learning, and the lightweight of this method is sufficient to enable deployment on mobile devices.

[0064] Among them, step S1 specifically includes the following steps;

[0065] Step S11: Obtain public natural scene text datasets, such as ICDAR2013, ICDAR2015, ICDAR2019, RCTW, etc.;

[0066] Step S12: Convert all the label formats in the dataset into the format required by YOLOV5, that is, one picture corresponds to one txt file, and each line in the txt file corresponds to a text area in the image. The format of each line is (cls, x center / textw, y center / texth, imgw / textw, imgh / texth); where cls = 0, representing that this area is a positive sample, x center is the abscissa of the center point of the text area, y center is the ordinate of the center point of the text area, imgw and imgh represent the width and length of the image respectively, and textw and texth represent the width and length of the text area respectively.

[0067] Step S13: Record the corresponding text in the text areas in the dataset into a json file for subsequent recognition. The format of the json file is {'xxxjpg': {'points': [[coordinates of text area 1], [coordinates of text area 2],...]}, {'text'}: [[text of text area 1], [text of text area 2],...]}.

[0068] Step S2 specifically includes the following steps;

[0069] Step S21: Input the images batch by batch into a feature extractor composed of multiple Conv modules and multiple BottleneckCSP modules. The Conv module includes a convolutional layer with a convolution kernel size of 3×3, a batch normalization layer BN, and a SiLU activation function. As shown in Formula 1:

[0070] F conv_out = SiLU(BN(Conv 3×3 (F Conv_in ))) Formula 1;

[0071] where F Conv_in is the input feature of the Conv module, and Conv 3×3 is the convolutional layer with a convolution kernel size of 3×3.

[0072] The BottleneckCSP module is composed of Bottleneck plus CSP; Bottleneck passes the input feature through a convolutional layer with a convolution kernel size of 1×1, then through a convolutional layer with a convolution kernel size of 3×3, and then adds the input feature to it. As shown in Formula 2, where F Bottleneck is the output of the Bottleneck module, F Bottleneck_in is the input feature of the Bottleneck module, Conv 3×3 is the convolutional layer with a convolution kernel size of 3×3, and Conv 1×1 is the convolutional layer with a convolution kernel size of 1×1.

[0073] F Bottleneck = F Bottleneck_in + Conv 3×3 (Conv 1×1 (F Bottleneck_in )) Formula 2;

[0074] CSP divides the original input into two branches, respectively performs convolutional operations to halve the number of channels, then one branch performs Bottleneck×N operations, where N is a custom parameter, and then Concat the two branches so that the input and output of BottlenneckCSP are of the same size, so as to enable the model to learn more features. As shown in Formula 3:

[0075] F Concat = Concat(N×Bottleneck(Conv 1×1 (F in_c / 2_1 )),Conv 3×3 (F in_c / 2_2 ))) Formula 3;

[0076] Among them, F Concat is the result of concatenating two branches. Concat is an operation of feature concatenation. Bottleneck represents the operation of Formula 2. F in_c / 2_1 and F in_c / 2_2 represent two branches of the input features, and the number of channels is half of the original input features.

[0077] Then, F Concat is passed through the batch normalization layer BN, the LekyReLU activation function, and Conv 1×1 to obtain the output F BottleneckCSP of BottlenneckCSP, as shown in Formula 4:

[0078] F BottleneckCSP = Conv 1×1 (LekyReLU(BN(F Concat ))) Formula 4;

[0079] Step S22: Input the features that have been downsampled 32 times by the Conv module and BottleneckCSP into the SPP (Spatial Pyramid Pooling) layer module. Perform max pooling operations on feature maps of different sizes, and then concatenate the pooled features as the output of the feature extractor. As shown in Formula 5:

[0080] F SPP_out = DeformableConv(Concat(F SPP_in , MaxPooling 13×13 (F SPP_in) ,

[0081] MaxPooling 9×9 (F SPP_in ), MaxPooling 5×5 (F SPP_in ))) Formula 5;

[0082] Among them, F SPP_in is the input feature of the SPP module, F SPP_out is the output of the SPP module, MaxPooling 13×13 , MaxPooling 9×9 , MaxPooling 5×5 represent max pooling layers with sampling kernel sizes of 13×13, 9×9, and 5×5 respectively, and DeformableConv is a deformable convolution module.

[0083] Step S3 specifically includes the following steps;

[0084] Step S31: Use the Kmeans algorithm to fit the aspect ratio of the real text box. Input the ratios of all real text boxes into Kmeans, and the Kmeans algorithm can cluster the aspect ratios of multiple anchor boxes.

[0085] Step S32: Predict the deviation between the anchor box and the real text box using the features extracted by the feature extractor. First, pass the features through a 1×7 long convolutional network to extract semantic features suitable for long texts. Then divide the processed features into gridn×gridn grids, where gridn is a custom parameter. The network will predict four offsets t x1 , t y1 , t h1 , t w1 , and the calculation methods are shown in Formulas Six, Seven, Eight, and Nine:

[0086] t x1 = log((bbox x2 - c x3 ) / (1 - (bbox x2 - c x3 ))) Formula Six;

[0087] t y1 = log((bbox y2 - c y3 ) / (1 - (bbox y2 - c y3 ))) Formula Seven;

[0088] t h1 = log(gt h4 / p h5 ) Formula Eight;

[0089] t w1 = log(gt w4 / p w5 ) Formula Nine;

[0090] Where bbox x2 , bbox y2 represent the horizontal and vertical coordinates of the center point of the real text box respectively; c x3 , c y3 represent the horizontal and vertical coordinates of the upper left corner of the grid corresponding to the real text box; gt h4 , gt w4 represent the height and width of the real text box; P h5 , p w5 represent the height and width of the anchor box. The network predicts the position of the text box by predicting these 4 offsets.

[0091] Step S4 specifically includes the following steps;

[0092] Step S41: Process the long semantic features using a hierarchical attention mechanism. The hierarchical attention mechanism is implemented through three matrices, including a query matrix Q, a key matrix K, and a value matrix V. Embed the words of the predicted text features into matrix E, and linearly map E into the query matrix Q, the key matrix K, and the value matrix V. Multiply the query matrix Q by the key matrix K to evaluate the scores for each pixel in the feature map. The magnitude of the score represents the degree of tightness of the association between two feature pixels. Then divide the obtained score by the square root of the dimension dim of the key vector to enhance the stability of the gradient. Next, use the softmax function to make the scores of all words positive and their sum equal to 1. Finally, multiply the obtained LekyReLU scores by the value matrix V to obtain the output of the attention layer, denoted as matrix O. As shown in Equation Ten:

[0093]

[0094] Step S42: Input O into a bidirectional LSTM to align the text features with the text and predict the final text result.

[0095] This embodiment addresses the problems of the traditional vision model having excessive parameters and being unable to effectively understand the deep semantic information of images. A method based on YOLOV5 is proposed. Utilizing the idea of the pre-trained model, it effectively solves the problem of the need for a large number of artificially synthesized false datasets for training the model. Moreover, due to the idea of its lightweight model, the model can be deployed on mobile devices. The present invention utilizes a hierarchical attention mechanism, which can mimic the phenomenon of attention concentration when humans observe things, effectively extracts the hidden connections inside the image and the text, and solves the problem that ordinary attention cannot focus on long text features, reduces the model parameters, and uses a bidirectional LSTM to align the text features and the text content.

[0096] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] The present invention is described with reference to the flowcharts of methods, apparatuses (devices), and computer program products according to embodiments of the present invention. It should be understood that each process in the flowchart and the combination of processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 or multiple processes.

[0098] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 or multiple flowcharts.

[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 or multiple processes.

[0100] As mentioned above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

[0101] This patent is not limited to the above best implementation manner. Anyone can obtain various other forms of natural scene text detection and recognition methods based on YOLOV5 under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage of this patent.

Claims

1. A method for natural scene text detection and recognition based on YOLOV5, characterized in that, Including the following steps; Step S1: Obtain a natural scene text image dataset and convert the corresponding labels into the format corresponding to YOLOV5; Step S2: Use the lightweight feature extractor of YOLOV5 to extract the position information and deep semantic information of the image text; Combine the shallow features and deep features using cross-layer connections and spatial pyramid pooling layers; Add deformable convolutions in the cross-layer connections so that the network can better handle the changes in the feature map scale; Step S3: Use the Kmeans algorithm to aggregate the aspect ratios of the anchor boxes that fit the true text boxes and predict the deviation between the anchor boxes and the true boxes; Use long convolutions to process the features to make the aspect ratios of the anchor boxes more suitable for the true text boxes; Step S4: Use bidirectional LSTM and attention mechanisms to align text features and predict text sequences; Step S2 specifically includes the following steps; Step S21: Input the images batch by batch into a feature extractor composed of multiple Conv modules and multiple BottleneckCSP modules, where the Conv module contains a convolutional layer with a convolutional kernel size of 3×3, a batch normalization layer BN, and a SiLU activation function; As shown in Formula 1: F Conv_out = SiLU(BN(Conv 3×3 (F Conv_in ))) Formula 1; where F Conv_in is the input feature of the Conv module, and Conv 3×3 is a convolutional layer with a convolutional kernel size of 3×3; The BottleneckCSP module is composed of Bottleneck plus CSP; Bottleneck passes the input features through a convolutional layer with a kernel size of 1×1, then through a convolutional layer with a kernel size of 3×3, and then adds the input features to it; as shown in Formula 2, where F Bottleneck is the output of the Bottleneck module, F Bottleneck_in is the input feature of the Bottleneck module, Conv 3×3 is the convolutional layer with a kernel size of 3×3, Conv 1×1 is the convolutional layer with a kernel size of 1×1; F Bottleneck = F Bottleneck_in + Conv 3×3 (Conv 1×1 (F Bottleneck_in )) Formula 2; CSP divides the original input into two branches, performs convolutional operations on each branch to halve the number of channels, then one branch performs Bottleneck×N operations, where N is a custom parameter, and then Concat the two branches so that the input and output of BottlenneckCSP are of the same size; As shown in Formula 3: F Concat = Concat(N × Bottleneck(Conv 1×1 (F in_c / 2_1 )), Conv 3×3 (F in_c / 2_2 )) Formula 3; Among them, F Concat is the result of concatenating two branches. Concat is an operation of feature concatenation, and Bottleneck refers to the operation in Formula 2. F in_c / 2_1 and F in_c / 2_2 represent two branches of the input features, and the number of channels is half of the original input features; Then, F Concat passes through the batch normalization layer BN, the LeakyReLU activation function, and Conv 1×1 to obtain the output F of BottlenneckCSP BottleneckCSP , as shown in Equation 4: F BottleneckCSP = Conv 1×1 (LekyReLU(BN(F Concat ))) Formula 4; Step S22: Input the features downsampled 32 times by the Conv module and BottleneckCSP into the SPP spatial pyramid pooling layer module, perform max pooling operations on feature maps of different sizes, and then splice the pooled features as the output of the feature extractor; As shown in Formula 5: F SPP_out = DeformableConv(Concat(F SPP_in , Maxpooling 13×13 (F SPP_in ), Maxpooling 9×9 (F SPP_in ),MaxPooling 5×5 (F SPP_in ))) Formula 5; Among which F SPP_in is the input feature of the SPP module, and F SPP_out is the output of the SPP module. Maxpooling 13×13 , MaxPooling 9×9 , MaxPooling 5×5 respectively represent the maximum pooling layers with sampling kernel sizes of 13×13, 9×9, and 5×5. DeformableConv is the deformable convolution module; Step S3 specifically includes the following steps; Step S31: Use the Kmeans algorithm to fit the aspect ratios of the true text boxes, input the ratios of all true text boxes into Kmeans to cluster the aspect ratios of multiple anchor boxes; Step S32: Predict the deviation between the anchor box and the true text box using the features extracted by the feature extractor; first, pass the features through a 1×7 long convolutional network to extract semantic features suitable for long texts; then divide the processed features into gridn×gridn grids, where gridn is a custom parameter; the network will predict four offsets t x1 , t y1 , t h1 , t w1 , and the calculation methods are shown in Formulas Six, Seven, Eight, and Nine: t x1 = log((bbox x2 - c x3 ) / (1 - (bbox x2 - c x3 ))) Formula 6; t y1 = log((bbox y2 - c y3 ) / (1 - (bbox y2 - c y3 ))) Formula VII; t h1 = log(gt h4 / p h5 ) Formula VIII; t w1 = log(gt w4 / p w5 ) Formula IX; where bbox x2 , bbox y2 represent the horizontal and vertical coordinates of the center point of the true text box respectively; c x3 , c y3 represent the horizontal and vertical coordinates of the upper left corner of the grid corresponding to the true text box; gt h4 , gt w4 represent the height and width of the true text box; p h5 , p w5 represent the height and width of the anchor box; the network predicts the position of the text box by predicting these 4 offsets.

2. The method for natural scene text detection and recognition based on YOLOV5 according to claim 1, characterized in that: Step S1 specifically Including the following steps; Step S11: Obtain a public natural scene text dataset; Step S12: Convert all the label formats in the dataset into the format required by YOLOV5; Step S13: Record the corresponding text in the text area of the dataset into a json file for subsequent recognition.

3. The method for natural scene text detection and recognition based on YOLOV5 according to claim 1, wherein: Step S4 specifically includes the following steps; Step S41: Process the long semantic features using a hierarchical attention mechanism. The hierarchical attention mechanism is implemented through three matrices, including a query matrix Q, a key matrix K, and a value matrix V. Embed the words of the predicted text features into matrix E, and linearly map matrix E into the query matrix Q, the key matrix K, and the value matrix V. Multiply the query matrix Q by the key matrix K to evaluate the scores for each pixel in the feature map. The higher or lower the score represents the degree of tightness of the association between two feature pixels. Then divide the obtained score by the square root of the dimension dim of the key vector to enhance the stability of the gradient. Next, use the softmax function to make the scores of all words positive and their sum equal to 1. Finally, multiply the obtained LekyReLU scores by the value matrix V to obtain the output of the attention layer, denoted here as matrix O, as shown in Equation Ten: Step S42: Input O into a bidirectional LSTM to align the text features with the text and predict the final text result.

Citation Information

Patent Citations

  • A natural scene text detection method based on full convolution neural network

    CN109299274A

  • Patient assistance intelligent auditing system based on deep learning

    CN111353445A