Real-time Arbitrary Shape Text Detection Method Based on Bidirectional Information Transmission

By adopting a real-time arbitrary shape text detection method based on bidirectional information transmission in text detection in natural scenes, the problems of slow model inference speed and incomplete text kernel semantics in the prior art are solved, and the detection effect of high accuracy and high processing speed is achieved.

CN116092067BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310026346.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-06-27
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

In the text detection in natural scenarios, the model parameters are many and the calculation amount is large, which leads to slow inference speed and difficult to apply to mobile devices, and the text kernel semantics are incomplete, making it difficult to learn.

Method used

Using a real-time arbitrary shape text detection method based on bidirectional information transmission, by constructing a training data set, kernel area labels and gap area labels are generated, and supervised training is performed in the text detection network, and text instances are reconstructed using the predicted text kernel area.

Benefits of technology

It achieves high detection accuracy while maintaining high processing speed, solving the problem of difficult learning caused by incomplete text kernel semantics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092067B_ABST
    Figure CN116092067B_ABST
Patent Text Reader

Abstract

The present invention provides a real-time arbitrary shape text detection method based on bidirectional information transmission. First, a training data set is constructed, and kernel region labels and gap region labels are generated using text region labels, and data augmentation and normalization processing are performed on the sample images in the training set; then, a text detection network is constructed and network training is carried out. During the training, the text region and the text kernel region are supervised, and the gap prediction value is calculated using the two prediction results to achieve information flow; finally, the text instances are reconstructed using the predicted text kernel regions. The present invention can solve the problem that it is difficult to learn due to incomplete text kernel semantics, and can obtain high detection accuracy while maintaining a high processing speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image and graphic processing, and particularly relates to a real-time arbitrary shape text detection method based on bidirectional information transmission. Background Art

[0002] Text detection in natural scenes is an important step in text recognition. Although text detection in natural scenes has achieved high accuracy, due to the actual needs of intelligent scenarios, speed is also an important factor in applying text detection to real life. Most of the current models have many network parameters and large computational amounts, resulting in slow inference speed of the entire model and being difficult to apply to some mobile devices. For example, in the literature "Wang W, Xie E and Li X, et al., Shape robust text detection with progressive scale expansion network, in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 9336-9345.", a segmentation method of first shrinking and then expanding, called (progressive scale expansion, PSE), is proposed to solve the problem of adjacent instance edge adhesion in traditional segmentation. However, the speed of post-processing is too slow to be applied to mobile devices. The literature "M. Liao, Z. Wan, C. Yao, K. Chen, and X. Bai, Real-time scene text detection with differentiable binarization, in Proceedings of the AAAI Conference on Artificial Intelligence. 2020: 11474–11481." further trains the binarization post-processing process of the model in the network by means of an adaptive threshold, solving the problem that the post-processing is not connected with the network training. However, no additional attention is paid to the text kernel. The literature "Wang W, Xie E and Li X, et al. Efficient and accurate arbitrary-shaped text detection with pixel aggregation network, in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2019: 8440–8449." uses the clustering method to cluster the pixels that belong to the text region but not the text kernel region, solving the problem of adjacent text instance adhesion. However, the pixel-level post-processing reduces the inference speed.The above three methods all utilize the text kernel region for detection. However, the text kernel region is an artificially defined concept that is difficult to understand and learn. Summary of the Invention

[0003] To overcome the deficiencies of the prior art, the present invention provides a real-time arbitrary-shaped text detection method based on bidirectional information transfer. First, a training data set is constructed. Kernel region labels and gap region labels are generated using text region labels, and data augmentation and normalization are performed on the sample images in the training data set. Then, a text detection network is constructed and network training is carried out. During training, the text region and the text kernel region are supervised, and the gap prediction value is calculated using the two prediction results to achieve information flow. Finally, the predicted text kernel region is used to reconstruct text instances. The present invention can solve the problem that it is difficult to learn due to the incomplete semantics of the text kernel, and can obtain high detection accuracy while maintaining a high processing speed.

[0004] A real-time arbitrary-shaped text detection method based on bidirectional information transfer, characterized by the following steps:

[0005] Step 1: Select the publicly available CTW1500 line-level curved text annotation data set as the training data set, generate kernel region labels using the publicly available text region labels, and subtract the text region labels from the kernel region labels to generate gap region labels;

[0006] Step 2: Select positive and negative samples from the training data set in a ratio of 1:3, where positive samples refer to pixels belonging to the foreground part and negative samples refer to pixels belonging to the background part. Then, perform data augmentation processing and normalization operations on all sample images. The data augmentation processing includes random rotation, random cropping, and random scaling operations, so that the image size in the data set is unified to 3×H×W, where H represents the image height and W represents the image width, and both are set to 640 during training;

[0007] Step 3: Input the sample images processed in Step 2 into the text detection network for network training to obtain a trained network. During training, the SGD optimizer is used, and the learning rate decay is calculated using the following formula:

[0008]

[0009] where l represents the learning rate of the current training epoch, l ini is the initial learning rate, epoth is the current training epoch, and max_epoch is the total number of training epochs;

[0010] During training, set the text kernel prediction probability map loss function L k as follows:

[0011] L k = ∑i∈S [-y i *log(x i )-(1-y i )*log(1-x i )] (2)

[0012] where \(i\) represents the \(i\)-th pixel sample, \(S\) represents the sample set selected according to the positive and negative sample ratio of 1:3, \(x\) i is the predicted value of the text kernel probability, and \(y\) i is the label of the text kernel;

[0013] During training, the loss function \(L\) of the text region prediction probability map t and the loss function \(L\) of the gap prediction probability map g are as follows:

[0014]

[0015] where \(S'\) is the set of all pixels in the image, \(\epsilon\) is a very small value to avoid a denominator of 0, and \(\epsilon\) is set to 0.000001; \(p\) gt represents the true label of the pixel, and \(p\) pre represents the predicted value of the pixel;

[0016] The total loss function \(L\) adopted during training is:

[0017]

[0018] where \(\mu\) is the coefficient of the loss term of the text kernel prediction probability map, is the coefficient of the loss term of the text region prediction probability map, \(\omega\) is the coefficient of the loss term of the gap prediction probability map, and \(\mu = 6\), \(\omega = 1\) are set respectively;

[0019] The specific processing process of the described text detection network is as follows:

[0020] Step a: Input the image into the ResNet18 network with deformable convolution for feature extraction, and output four groups of features with different resolutions. The resolutions are respectively

[0021] Step b: Use a convolutional layer with a convolution kernel of 1×1 to adjust the resolutions of the four groups of features obtained in step a to

[0022] Step c: Upsample the feature with a resolution of by 2 times and then add it to the feature with a resolution of . Add the newly obtained feature with a resolution of upsampled by 2 times to the feature with a resolution of The features are added together, and the newly obtained resolution is The features are upsampled by a factor of 2 and then added to the features with a resolution of The features are added together, and four new groups of features are formed with the features with a resolution of in step b;

[0023] Step d: Use a 3×3 convolutional layer to reduce the number of channels of the four groups of features obtained in step c from 256 to 64, and then use upsampling to adjust these four groups of features to the same resolution and then concatenate them to obtain a feature map F with a resolution of 256× ;

[0024] Step e: Use a 3×3 convolutional layer with normalization and ReLU activation function to process the feature map F to obtain a feature map F1 with a resolution of ; Use a 1×1 transposed convolutional layer with normalization and ReLU activation function to process the feature map F1 to obtain a feature map F2 with a resolution of ; Finally, process the feature map F2 through a 1×1 transposed convolutional layer and a Sigmoid function to obtain a text kernel region prediction probability map P with a size of 1×H×W K ;

[0025] Step f: Use a new group of 3×3 convolutional layers with normalization and ReLU activation function to process the feature map F to obtain a feature map F3 with a resolution of ; Use a 1×1 transposed convolutional layer with normalization and ReLU activation function to process the feature map F3 to obtain a feature map F4 with a resolution of ; Finally, process the feature map F4 through a 1×1 transposed convolutional layer and a Sigmoid function to obtain a text region prediction probability map P with a size of 1×H×W t ;

[0026] Step g: Input the text kernel region prediction probability map P K obtained in step e and the text region prediction probability map P t obtained in step f into the bidirectional information transfer module, and output a gap region prediction probability map P with a size of 1×H×W g ;

[0027] The described bidirectional information transfer module is processed according to the following function:

[0028]

[0029] where α is the first network hyperparameter to be learned, and β is the second network hyperparameter to be learned;

[0030] Step 4: Input the text image dataset to be processed into the text detection network trained in Step 3. Binarize the predicted probability map of the text kernel region output by the network, and then perform dilation processing using the functions in the OpenCV library to obtain the final detection result map.

[0031] Further, Step 4 adopts the following processing method: Discard the text region prediction and gap region prediction in the text detection network trained in Step 3, only perform text kernel region prediction, then input the text image dataset to be processed into the network to obtain the predicted probability map of the text kernel region, binarize the predicted probability map, and then perform dilation processing using the functions in the OpenCV library to obtain the final detection result map.

[0032] The beneficial effects of the present invention are as follows: Since a deep neural network is used to extract text information in natural scenes, the position of the text can be segmented; considering the guidance of text information for text kernel information while constructing the deep neural network can improve the model's cognitive ability for the text kernel region, and can well detect an entire sentence as a text instance. At the same time, this guidance can be removed during the final processing, so as to maintain a fast inference speed while improving the accuracy. Brief Description of the Drawings

[0033] Figure 1 is a flowchart of the real-time arbitrary shape text detection method based on bidirectional information transmission of the present invention;

[0034] Figure 2 is a text detection result map in different scenarios using the method of the present invention. Detailed Embodiments

[0035] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.

[0036] As Figure 1 shown, the present invention provides a real-time arbitrary shape text detection method based on bidirectional information transmission. The main process is as follows: Select a natural scene text dataset, generate text kernel region and gap region labels according to the publicly available text region labels; construct a scene text detection network, and use the training dataset to train the model; input the data to be detected into the optimal model for text detection to obtain the text detection result. The specific steps are as follows:

[0037] 1. Select the publicly available CTW1500 line-level curved text annotation dataset as the training dataset, use the publicly available text region labels to generate kernel region labels, and subtract the text region labels from the kernel region labels to generate gap region labels. Specifically, the Vatti clipping algorithm described in the literature "Vatti B R. A generic solution to polygon clipping[J]. Communications of the ACM, 1992, 35(7): 56-63." can be used to generate kernel region labels.

[0038] 2. Select positive and negative samples from the training dataset in a ratio of 1:3, where positive samples refer to pixels belonging to the foreground part and negative samples refer to pixels belonging to the background part; perform data augmentation on the images, including random rotation, random cropping, and random scaling operations, so that the images in the dataset are uniformly sized to 3×H×W, and then perform normalization operations; where, H represents the image height and W represents the image width, both of which are set to 640 during training.

[0039] 3. Input the sample images processed in step 2 into the text detection network for network training to obtain the detection results.

[0040] Among them, the specific processing process of the text detection network is as follows:

[0041] Step a: Input the image into the ResNet18 network with deformable convolution for feature extraction, and output four groups of features with different resolutions, and the resolutions are

[0042] Step b: Use a convolutional layer with a convolution kernel of 1×1 to adjust the resolutions of the four groups of features obtained in step a to

[0043] Step c: Upsample the feature with a resolution of by 2 times and add it to the feature with a resolution of , add the newly obtained feature with a resolution of by 2 times and add it to the feature with a resolution of , add the newly obtained feature with a resolution of by 2 times and add it to the feature with a resolution of , and form four groups of new features with the feature with a resolution of in step b;

[0044] Step d: Use a 3×3 convolutional layer to reduce the number of channels of the four groups of features obtained in step c from 256 to 64, and then use upsampling to adjust these four groups of features to the same resolution Then perform splicing to obtain a feature map F with a resolution of 256× ;

[0045] Step e: Process the feature map F using a 3×3 convolutional layer with normalization and ReLU activation function to obtain a feature map F1 with a resolution of ; Process the feature map F1 using a 1×1 transposed convolutional layer with normalization and ReLU activation function to obtain a feature map F2 with a resolution of ; Finally, process the feature map F2 through a 1×1 transposed convolutional layer and the Sigmoid function to obtain a text kernel region prediction probability map P of size 1×H×W K ;

[0046] Step f: Process the feature map F using a new set of 3×3 convolutional layers with normalization and ReLU activation function to obtain a feature map F3 with a resolution of ; Process the feature map F3 using a 1×1 transposed convolutional layer with normalization and ReLU activation function to obtain a feature map F4 with a resolution of ; Finally, process the feature map F4 through a 1×1 transposed convolutional layer and the Sigmoid function to obtain a text region prediction probability map P of size 1×H×W t ;

[0047] Step g: Input the text kernel region prediction probability map P K obtained in step e and the text region prediction probability map P t obtained in step f into the bidirectional information transfer module, and output a gap region prediction probability map P of size 1×H×W g ;

[0048] The described bidirectional information transfer module is processed according to the following function:

[0049]

[0050] where α is the first network hyperparameter to be learned, and β is the second network hyperparameter to be learned;

[0051] During training, the specific settings include: The SGD optimizer is used during training, and the learning rate decay is calculated using the following formula:

[0052]

[0053] where l represents the learning rate of the current training epoch, l ini is the initial learning rate, epoch is the current training epoch, and max_epoch is the total number of training epochs;

[0054] Set the loss function \(L\) of the text kernel prediction probability map during training k as follows:

[0055] \(L\) k =\(\sum\) i∈S [-y i *\(\log(x\) i )-(1 - y i )*\(\log(1 - x\) i )] (8)

[0056] where \(i\) represents the \(i\)-th pixel sample, \(S\) represents the sample set selected according to the positive-negative sample ratio of 1:3, \(x\) i is the text kernel probability prediction value, and \(y\) i is the label of the text kernel;

[0057] Set the loss function \(L\) of the text region prediction probability map t and the loss function \(L\) of the gap prediction probability map g as follows:

[0058]

[0059] where \(S'\) is the set of all pixels in the image, \(\varepsilon\) is a very small value to avoid a denominator of 0, and \(\varepsilon\) is set to 0.000001; \(p\) gt represents the true label of the pixel, and \(p\) pre represents the predicted value of the pixel;

[0060] The total loss function \(L\) adopted during training is:

[0061]

[0062] where \(\mu\) is the coefficient of the text kernel prediction probability map loss term, is the coefficient of the text region prediction probability map loss term, \(\omega\) is the coefficient of the gap prediction probability map loss term, and \(\mu = 6\), \(\omega = 1\) are set respectively;

[0063] 4. Input the text image dataset to be processed into the text detection network trained in step 3, binarize the text kernel region prediction probability map output by the network, and then perform dilation processing using the functions in the opencv library to obtain the final detection result map.

[0064] To improve the processing speed and efficiency, it is also possible to discard the text region prediction and gap region prediction in the trained text detection network, only perform text kernel region prediction, and then perform the above binarization processing and dilation processing respectively to obtain the final detection result map.

[0065] Figure 2The predicted result images of text detection for different scenarios using the method of the present invention are given. Among them, the leftmost column is the final detection result image, the middle column is the corresponding kernel region prediction image, and the rightmost column is the corresponding text region prediction image. It can be seen that the method of the present invention can well detect the text in the figure and can well detect an entire sentence as one target.

Claims

1. A real-time arbitrary shape text detection method based on bidirectional information transmission, characterized in that The steps are as follows: Step 1: Select the publicly available CTW1500 line-level curved text annotation dataset as the training dataset. Use the publicly available text region labels to generate kernel region labels, and subtract the kernel region labels from the text region labels to generate gap region labels; Step 2: Select positive and negative samples from the training dataset in a ratio of 1:

3. Among them, positive samples refer to pixels belonging to the foreground part, and negative samples refer to pixels belonging to the background part; then, perform data augmentation processing and normalization operations on all sample images. The data augmentation processing includes random rotation, random cropping, and random scaling operations, so that the image size in the dataset is unified to 3×H×W, where H represents the image height and W represents the image width, and both are set to 640 during training; Step 3: Input the sample images processed in Step 2 into the text detection network for network training to obtain a trained network. During training, use the SGD optimizer, and the learning rate decay is calculated using the following formula: Among them, l represents the learning rate of the current training round, and l ini is the initial learning rate, epoch is the current training round, and max_epoch is the total number of training rounds; During training, set the text kernel prediction probability map loss function L k as follows: L k = ∑ i∈S [-y i * log(x i ) - (1 - y i ) * log(1 - x i )] (2) where i represents the i-th pixel sample, S represents the sample set selected according to the positive-negative sample ratio of 1:3, x i is the predicted value of the text kernel probability, and y i is the label of the text kernel; Set the loss function \(L\) of the text region prediction probability map during training t and the loss function \(L\) of the gap prediction probability map g as follows: Among them, S′ is the set of all pixels in the image, ε is a very small value to avoid a denominator of 0, and ε is set to 0.000001; p gt represents the true label of the pixel, and p pre represents the predicted value of the pixel; The total loss function L used during training is: Among them, μ is the coefficient of the loss term of the text kernel prediction probability map, is the coefficient of the loss term of the text region prediction probability map, ω is the coefficient of the loss term of the gap prediction probability map, and μ = 6 and ω = 1 are set respectively; The specific processing process of the text detection network is as follows: Step a: Input the image into the ResNet18 network with deformable convolution for feature extraction, and output four groups of features with different resolutions, and the resolutions are respectively Step b: Use a convolutional layer with a convolution kernel of 1×1 to adjust the resolution of the four groups of features obtained in step a to Step c: Upsample the feature with a resolution of by 2 times and then add it to the feature with a resolution of . Upsample the newly obtained feature with a resolution of by 2 times and then add it to the feature with a resolution of . Upsample the newly obtained feature with a resolution of by 2 times and then add it to the feature with a resolution of . Combine the features with the features with a resolution of in step b to form four groups of new features; Step d: Use a 3×3 convolutional layer to reduce the number of channels of the four groups of features obtained in step c from 256 to 64, and then use upsampling to adjust these four groups of features to the same resolution. Then perform concatenation to obtain a feature map F with a resolution of ​ Step e: Process the feature map F using a 3×3 convolutional layer with normalization and ReLU activation function to obtain a feature map F1 with a resolution of ; Process the feature map F1 using a 1×1 transposed convolutional layer with normalization and ReLU activation function to obtain a feature map F2 with a resolution of ; Finally, process the feature map F2 through a 1×1 transposed convolutional layer and the Sigmoid function to obtain a text kernel region prediction probability map P of size 1×H×W K ; Step f: Process the feature map F using a new set of 3×3 convolutional layers with normalization and ReLU activation functions to obtain a feature map F3 with a resolution of ; Process the feature map F3 using a 1×1 transposed convolutional layer with normalization and ReLU activation functions to obtain a feature map F4 with a resolution of ; Finally, process the feature map F4 through a 1×1 transposed convolutional layer and the Sigmoid function to obtain a text region prediction probability map P of size 1×H×W t ; Step g: Input the text kernel region prediction probability map P obtained in step e K and the text region prediction probability map P obtained in step f t into the bidirectional information transfer module, and output the gap region prediction probability map P with a size of 1×H×W g ; The two-way information transfer module is processed according to the following function: where α is the first network hyperparameter to be learned, and β is the second network hyperparameter to be learned; Step 4: Input the text image dataset to be processed into the text detection network trained in Step 3, perform binarization processing on the text kernel region prediction probability map output by the network, and then use the functions in the opencv library for dilation processing to obtain the final detection result map.

2. The real-time arbitrary shape text detection method based on bidirectional information transmission according to claim 1, wherein: Step 4 adopts the following processing method: Discard the text region prediction and gap region prediction in the text detection network trained in Step 3, only perform text kernel region prediction, then input the text image dataset to be processed into the network to obtain the text kernel region prediction probability map, perform binarization processing on this prediction probability map, and then use the functions in the opencv library for dilation processing to obtain the final detection result map.

Citation Information

Patent Citations

  • Multidirectional natural scene text detection method based on character segmentation

    CN111753714A

  • Text detection method based on attention feature fusion and hole residual feature enhancement

    CN113486890A