A Deep Learning-Based Text Detection Method for Complex Scenarios

By combining the GCCM module and the Transformer encoder ResNet50 network, the problem of text detection in complex backgrounds is solved, accurate distinction between text and non-text areas and adaptive detection of variable text shapes is achieved, and the detection accuracy of small text instances is improved.

CN118968491BActive Publication Date: 2025-07-11SICHUAN JISU POWER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411172626.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-07-11
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

The prior art has difficulty distinguishing text from non-text areas in complex contexts, difficult to deal with variable text shapes and directions, and difficult to detect small text instances.

Method used

The GCCM module and Transformer encoder combined with the ResNet50 network are used to build a multi-layer global context convolution module through global average pooling and depth separation of convolution and sigmoid activation functions, thereby enhancing the context utilization ability of the text detection model.

Benefits of technology

It improves the accuracy and robustness of text detection, can better handle text detection in complex backgrounds, adapt to variable text shapes and directions, and enhances the detection accuracy of small text instances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968491B_ABST
    Figure CN118968491B_ABST
Patent Text Reader

Abstract

The present invention discloses a complex scene text detection method based on deep learning, belonging to the technical field of specific computer models, including obtaining a text detection data set; constructing a GCCM module; constructing a multi-layer global context convolution module based on 2 Comodule modules, 3 GCCM modules, a splitting unit and a Conca unit, constructing a ResNet50-TransGC network based on the backbone network of Mask R-CNN and the multi-layer global context convolution module, and training with the text detection data set until convergence to obtain a text detection model; S6, obtaining a text image to be recognized, performing target recognition with the text detection model, and outputting a predicted text box. The present invention can more easily distinguish text and non-text regions in complex backgrounds, cope with variable text shapes and directions, improve the detection accuracy of small text instances, enhance the modeling ability for long-distance dependencies, and thus achieve accurate text detection in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a specific computing model, and particularly to a method for detecting text in complex scenarios based on deep learning. Background Art

[0002] Intelligent character recognition technology refers to the process of automatically recognizing characters from images containing text by a computer, which includes two parts: text detection and character recognition. The present invention mainly focuses on text detection technology. There are various existing technical solutions for text detection technology, such as:

[0003] 1. Traditional methods: Based on classical computer vision technologies such as edge detection, connected component analysis, and Hough transform to detect lines and boundaries in images, thereby locating text regions. However, traditional methods have poor performance in the face of complex backgrounds or lighting conditions, and have limited detection capabilities for non-standard fonts or handwritten texts.

[0004] 2. Morphological methods: Using mathematical morphological operations, such as dilation and erosion, to enhance text edges and remove noise, which is suitable for structured text detection. However, morphological methods are mainly suitable for structured text detection and have poor detection effects for free-form or unstructured texts.

[0005] 3. Deep learning-based methods: Using deep learning models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs). These models can learn features from a large amount of labeled data to detect text. CRNN (Convolutional Recurrent Neural Network) combines CNN and RNN for sequence modeling, such as text line detection and recognition. Object detection frameworks such as Mask R-CNN and Faster R-CNN. These methods take images with text as input, regard the text as the target, and output images with target detection boxes. These methods can accurately locate and segment text in images. However, there are still defects: (1) It is difficult to accurately detect and process text in complex backgrounds: In complex backgrounds, text may be mixed with other graphic elements, making it difficult for the model to distinguish text and non-text regions. (2) It is difficult to handle variable text shapes and orientations: Text may present various shapes, such as curved, tilted, or irregular arrangements. (3) Small texts may only occupy a small part of the image, so small text instances are often difficult to detect.

[0006] The Transformer generally refers to the Transformer encoder-decoder, which includes a Transformer encoder and a Transformer decoder. The Transformer encoder is a part of the Transformer architecture and is composed of multiple identical layers stacked together. Each layer has two sub-layers. The first sub-layer is the multi-head self-attention aggregation. The multi-head self-attention allows the model to simultaneously focus on different positions of the input, process information in parallel from multiple perspectives, thereby enhancing the ability to model global dependencies. This mechanism enables the model to effectively distinguish and aggregate relevant information from different parts of the input sequence, even if this information is physically far apart. The second sub-layer is the positionwise feed-forward network. Its main component is a multi-layer perceptron (MLP). The MLP consists of a series of fully connected layers, which are used to further process and transform the feature representation to make it more abstract and complex. By introducing non-linear transformations, the MLP can learn higher-level feature combinations and provide the model with more powerful expressive capabilities. To further improve the model performance, residual connections are embedded within each sub-layer. The residual connection allows the input to directly skip one or more layers and be directly passed to subsequent layers, which helps to alleviate the problem of gradient vanishing / explosion and makes the network more stable when training deep structures. By combining multi-head attention and the MLP, the Transformer self-attention structure not only strengthens the capture of different local information but also can deeply explore features through the self-attention mechanism. The self-attention mechanism enables the features at each position to interact with the features at all other positions in the sequence, thereby revealing the internal correlations between features and promoting the comprehensive understanding and utilization of features.

[0007] ResNet50 is a commonly used deep convolutional neural network, which is divided into 5 stages from a structural perspective and is named Stage 0, Stage 1, Stage 2, Stage 3, and Stage 4 in sequence, which can be translated as Stage 0, Stage 1, Stage 2, Stage 3, and Stage 4. The structure of Stage 0 is simple and only contains 1 convolution calculation operation. The convolution kernel is 7×7 and the stride is 2, which can be regarded as the preprocessing of the input. The structures of Stages 1 to 4 are relatively similar and are all composed of several Bottlenecks (bottleneck blocks). For example, Stage 1 contains 3 bottleneck blocks, and Stages 2 to 4 include 4, 6, and 3 bottleneck blocks respectively. Bottleneck is abbreviated as BTNK, which is a special type of residual module used to increase the depth of the network and reduce the computational amount. It is further divided into two specific implementation methods, BTNK1 and BTNK2. BTNK1 refers to the Bottleneck structure with different numbers of input and output channels, while BTNK2 refers to the Bottleneck structure with the same number of input and output channels. BTNK2 has 2 variable parameters C and W, that is, C and W in the input shape (C, W, W). Let the input with the shape (C, W, W) be x, and let the 3 convolution blocks (and related BN and RELU) on the left side of BTNK2 be the function F(x). After adding the two (F(x)+x) and then passing through 1 Reu activation function, the output of BTNK2 is obtained, and the shape of this output is still (C, W, W). BTNK1 has 4 variable parameters C, W, C1, and S. Compared with BTNK2, BTNK1 has 1 additional convolution layer on the right, and let it be the function G(x). BTNK1 corresponds to the case where the number of input x and output F(x) channels is different. It is exactly this added convolution layer that changes x to G(x), playing the role of matching the dimensional difference between the input and output (the number of channels of G(x) and F(x) is the same), and then the summation F(x)+G(x) can be performed. Stage 4 of ResNet50 is composed of 1 BTNK1 and 2 BTNK2s.

[0008] Mask R-CNN is a deep learning model for object detection and instance segmentation, which is extended based on Faster R-CNN. It mainly includes a backbone network, a Region Proposal Network (RPN), a Roni Head, a classification and bounding box regression, and a segmentation branch.

[0009] ConvModule Module: This module is commonly used in object detection networks from the YOLOv5 series to the YOLOv8 series. This module includes three functions: convolution, normalization, and activation. Specifically, it consists of 1 layer of 3×3 Conv2d + BatchNorm + SiLU activation function. The conx2d function is a function in the important deep learning library TensorFlow, which is used to implement the convolution operation of a convolutional neural network (CNN). Its input parameters include the input matrix, filter (weights), stride size, and padding method, and its output is generally the result after convolution. BatchNorm (Batch-Normalization), Chinese for batch normalization, is a commonly used technique in deep learning, which is used to accelerate the training of neural networks and improve the performance of the model. It normalizes the input at each hidden layer in the network, making the input distribution of each neuron more stable, thus contributing to the convergence and training effect of the network. Summary of the Invention

[0010] The object of the present invention is to provide a deep learning-based complex scene text detection method that solves the defects of difficult to distinguish text and non-text regions in complex backgrounds, difficult to cope with changing text shapes and directions, and difficult to detect small text instances.

[0011] To achieve the above object, the technical solution adopted by the present invention is as follows: A deep learning-based complex scene text detection method includes the following steps;

[0012] S1. Obtain a text detection data set, and the samples in the text detection data set are text images with text box annotations;

[0013] S2. Construct a GCCM module, including a global average pooling layer, a first 1×1 convolutional layer, a 1×11 depthwise separable convolutional layer, an 11×1 depthwise separable convolutional layer, a second 1×1 convolutional layer, and a sigmoid activation function layer connected in sequence from top to bottom;

[0014] S3. Construct a multi-layer global context convolution module, including 2 ConvModule modules, 3 GCCM modules, a splitting unit, and a Concat unit. The 2 ConvModule modules are respectively marked as ConvModule1 and ConvModule2, and the 3 GCCM modules are respectively marked as GCCM1, GCCM2, and GCCM3;

[0015] For the feature F input to the multi-layer global context convolution module, after passing through ConvModule1 and the splitting unit, it is split into a first feature and a second feature , Output the first global context feature through GCCM1, GCCM2, and GCCM3 in sequence and the second global context feature and the third global context feature , Output the fourth global context feature through GCCM2 and GCCM3 in sequence and the fifth global context feature , and add to get the sixth global context feature ; , and are concatenated by the Concat unit to form a concatenated feature , and then output multi-layer global context features through ConvModule1 ;

[0016] S4. Construct a ResNet50-TransGC network;

[0017] Select a Mask R-CNN network, whose backbone network is Resnet50. The stage 4 of the Resnet50 includes BTNK1 and two BTNK2 connected in sequence. Replace BTNK1 with a Transformer encoder and the first BTNK2 with a multi-layer global context convolution module to obtain a ResNet50-TransGC network;

[0018] S5. Train the ResNet50-TransGC network with a text detection dataset until convergence to obtain a text detection model;

[0019] S6. Obtain a text image to be recognized, perform object recognition with the text detection model, and output a predicted text box.

[0020] Preferably, the text detection dataset includes the Total-Text dataset, the ICDAR 2015 dataset, and the CTW1500 dataset.

[0021] Preferably, in the global context convolution module, if the input image size is C×W×H, where C is the number of channels, and W and H are the width and height of the input image respectively;

[0022] then the output size of the global average pooling layer is C×1×1;

[0023] The output feature size of the first 1×1 convolutional layer is K1×1×1, where K1 is the number of output channels of the first 1×1 convolutional layer;

[0024] The output feature size of the 1×11 depthwise separable convolutional layer is K2×W×H, where K2 is the number of output channels of the 1×11 depthwise separable convolutional layer;

[0025] The output feature size of the 11×1 depthwise separable convolutional layer is K3×W×H, where K3 is the number of output channels of the 11×1 depthwise separable convolutional layer;

[0026] The output feature size of the second 1×1 convolutional layer is K4×W×H, where K4 is the number of output channels of the second 1×1 convolutional layer;

[0027] The sigmoid activation function layer is used to compress the values of its input features into the interval [0,1], and its output feature size is K4×W×H.

[0028] Preferably: K1 = C, K2 = C / 2, K3 = C / 2, K4 = C.

[0029] Compared with the prior art, the advantages of the present invention are as follows:

[0030] (1) In the present invention, a Transformer encoder is introduced in stage 4 of the backbone network. The multi-head attention mechanism in the Transformer allows the model to view the input data from different perspectives or feature spaces, strengthening the utilization of context. This design allows the model to process sequences in parallel and dynamically allocate attention weights, so that when processing tasks such as natural language, it can instantaneously and comprehensively utilize context information, which is particularly useful for text detection because it can simultaneously consider multiple attributes of the text, such as font, size, direction, and layout, thereby improving the detection accuracy and robustness.

[0031] (2) The GCCM module is designed. In the present invention, GCCM stands for Global Context Convolution Module in English and Global Context Convolution Module in Chinese. In this module, first, global average pooling operation is performed on the input image to map high-dimensional features to a low-dimensional space, retaining the overall information of the image; then a 1×1 convolutional layer is used for dimensionality reduction to reduce the number of channels while keeping the input spatial size unchanged for subsequent operations; an 11×1 depthwise separable convolution is performed, which is an efficient convolution operation that can extract more feature information while reducing the computational cost; again, a 1×1 convolutional layer is used for further feature extraction or adjustment of the number of channels to meet the requirements of the model. Finally, the output is passed through the sigmoid function to limit it within the range of [0, 1] for facilitating the representation of attention weights.

[0032] (3) A multi - layer global context convolution module is designed based on the GCCM module. This module uses three GCCM modules to form a recursive structure. The input feature map undergoes initial feature extraction and processing through ConvModule1. The output feature map is split into two parts, and the split feature maps are respectively processed through their respective local context convolution modules. This process is repeated multiple times to deepen the level of feature extraction. Then, the processed feature maps are recombined together to form a new feature map. Finally, the multi - layer global context convolution module of the present invention can capture long - distance context information, use global average pooling and 1D bar convolution to enhance the features in the central region, and adapt to text features of different scales. This enables the model to finely perceive text elements and establish deeper semantic associations at multiple scales, especially when dealing with challenging text instances, such as those with dense background interference or distorted shapes.

[0033] (4) Integrate the above - mentioned Transformer encoder and multi - layer global context convolution module into stage 4 of Mask R - CNN. The Transformer encoder can process these long - distance dependencies in parallel without sacrificing computational efficiency, thus improving the processing speed of the entire system. The GCCM module is applied recursively, which can capture long - distance context information, which is crucial for understanding the relationships between text instances. By integrating these two technologies, the model can better adapt to a variety of text instances, including texts with different fonts, sizes, orientations, and layouts. This enables the model to more effectively capture the context information of text instances, thereby improving the accuracy of text detection, especially when facing challenging text instances such as texts in dense backgrounds and distorted - shaped texts, and achieving better detection results.

[0034] In summary, the present invention can more easily distinguish text and non - text regions in complex backgrounds, cope with variable text shapes and orientations, improve the detection accuracy of small text instances, and enhance the ability to model long - distance dependencies. Thus, accurate text detection in complex scenarios is achieved. Brief Description of the Drawings

[0035] Figure 1 It is the flow chart of the present invention;

[0036] Figure 2 It is the structural diagram of the global context convolution module;

[0037] Figure 3 It is the structural diagram of the multi - layer global context convolution module;

[0038] Figure 4 It is the structural diagram of the ResNet50 - TransGC network structure. Detailed Implementation Manner

[0039] The present invention will be further described below in conjunction with the accompanying drawings.

[0040] Example 1: Refer to Figures 1 to 4 , a complex scene text detection method based on deep learning, characterized in that it includes the following steps;

[0041] S1. Obtain a text detection data set, where the samples in the text detection data set are text images with text box annotations;

[0042] S2. Construct a GCCM module, including a global average pooling layer, a first 1×1 convolutional layer, a 1×11 depthwise separable convolutional layer, an 11×1 depthwise separable convolutional layer, a second 1×1 convolutional layer, and a sigmoid activation function layer connected in sequence from top to bottom;

[0043] S3. Construct a multi-layer global context convolutional module, including 2 ConvModule modules, 3 GCCM modules, a splitting unit, and a Concat unit. The 2 ConvModule modules are respectively labeled as ConvModule1 and ConvModule2, and the 3 GCCM modules are respectively labeled as GCCM1, GCCM2, and GCCM3;

[0044] For the feature F input to the multi-layer global context convolutional module, after passing through ConvModule1 and the splitting unit, it is split into a first feature , a second feature , successively passes through GCCM1, GCCM2, and GCCM3 to output a first global context feature , a second global context feature , a third global context feature , successively passes through GCCM2 and GCCM3 to obtain a fourth global context feature , a fifth global context feature , and is added to to obtain a sixth global context feature ; , and are concatenated by the Concat unit to form a concatenated feature , and then output a multi-layer global context feature through ConvModule1;

[0045] S4. Construct a ResNet50-TransGC network;

[0046] Select a Mask R-CNN network with Resnet50 as its backbone network. The fourth stage of the Resnet50 includes BTNK1 and two BTNK2s connected in sequence. Replace BTNK1 with a Transformer encoder and the first BTNK2 with a multi-layer global context convolution module to obtain the ResNet50-TransGC network;

[0047] S5. Train the ResNet50-TransGC network with a text detection dataset until convergence to obtain a text detection model;

[0048] S6. Obtain the text image to be recognized, perform object recognition with the text detection model, and output the predicted text box.

[0049] In this embodiment, the text detection dataset includes the Total-Text dataset, the ICDAR 2015 dataset, and the CTW1500 dataset.

[0050] In the global context convolution module, if the input image size is C×W×H, where C is the number of channels, and W and H are the width and height of the input image respectively;

[0051] Then the output size of the global average pooling layer is C×1×1;

[0052] The output feature size of the first 1×1 convolutional layer is K1×1×1, where K1 is the number of output channels of the first 1×1 convolutional layer;

[0053] The output feature size of the 1×11 depthwise separable convolutional layer is K2×W×H, where K2 is the number of output channels of the 1×11 depthwise separable convolutional layer;

[0054] The output feature size of the 11×1 depthwise separable convolutional layer is K3×W×H, where K3 is the number of output channels of the 11×1 depthwise separable convolutional layer;

[0055] The output feature size of the second 1×1 convolutional layer is K4×W×H, where K4 is the number of output channels of the second 1×1 convolutional layer;

[0056] The sigmoid activation function layer is used to compress the values of its input features into the interval [0,1], and its output feature size is K4×W×H.

[0057] In addition, K1 = C, K2 = C / 2, K3 = C / 2, K4 = C.

[0058] Regarding the multi-layer global context convolution module: It includes at least three GCCM modules, and the GCCM modules are connected by means of feature splitting and merging to recursively enhance the feature representation. The feature splitting unit divides the output feature map of the GCCM module into two parts, one part is passed to the next GCCM module, and the other part is passed to the Concat unit. The multi-layer global context convolution module can capture long-distance context information, enhance the features in the central region, and is applicable to text features of different scales. By iterating the GCCM module multiple times, it can capture context information in different ranges and improve the quality and diversity of feature representation. After the output feature map of the multi-layer global context convolution module is processed by the final convolution module, it can be used for subsequent text detection tasks, such as text instance segmentation and bounding box regression.

[0059] Embodiment 2: To illustrate the effects of the present invention, in this embodiment, the Total-Text dataset is used, and a comparative experiment is conducted between the algorithms in 4 existing technologies and the present invention under the same experimental conditions, and the following Table 1 is obtained:

[0060] Table 1. Comparative table of experimental results of each algorithm

[0061] algorithm Precision Recall F-measure TextSnake 82.70 74.50 78.40 PSENet-1s 81.77 75.11 78.30 TextSpotter 82.50 75.60 78.60 MaskR-CNN 72.83 79.83 76.17 the method of the present invention 79.81 79.95 79.88

[0062] In Table 1, Precision is the precision rate, Recall is the recall rate, and F-measure is the weighted harmonic mean of Precision and Recall. It can be seen from Table 1 that the present invention is superior to the existing algorithms in both precision rate and recall rate.

[0063] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A text detection method for complex scenes based on deep learning, characterized in that: It includes the following steps; S1. Obtain a text detection dataset, where the samples in the text detection dataset are text images with text box annotations; S2. Construct a GCCM module, including a global average pooling layer, a first 1×1 convolutional layer, a 1×11 depthwise separable convolutional layer, an 11×1 depthwise separable convolutional layer, a second 1×1 convolutional layer, and a sigmoid activation function layer connected in sequence from top to bottom; S3. Construct a multi-layer global context convolutional module, including 2 ConvModule modules, 3 GCCM modules, a splitting unit, and a Concat unit. The 2 ConvModule modules are respectively labeled as ConvModule1 and ConvModule2, and the 3 GCCM modules are respectively labeled as GCCM1, GCCM2, and GCCM3; For the feature F input to the multi-layer global context convolution module, after passing through ConvModule1 and the splitting unit, it is split into a first feature , a second feature , and successively passes through GCCM1, GCCM2, and GCCM3 to output a first global context feature , a second global context feature , a third global context feature , successively passes through GCCM2 and GCCM3 to obtain a fourth global context feature , a fifth global context feature , and is added to to obtain a sixth global context feature ; , and are concatenated by the Concat unit into a concatenated feature , and then pass through ConvModule1 to output a multi-layer global context feature ; S4. Construct a ResNet50-TransGC network; Select a Mask R-CNN network with a Resnet50 backbone network. The stage 4 of the Resnet50 includes BTNK1 and 2 BTNK2s connected in sequence. Replace BTNK1 with a Transformer encoder and the multi-layer global context convolutional module with the first BTNK2 to obtain a ResNet50-TransGC network; S5. Train the ResNet50-TransGC network with the text detection dataset until convergence to obtain a text detection model; S6. Obtain the text image to be recognized, perform object recognition with the text detection model, and output the predicted text box.

2. The method for detecting complex-scene text based on deep learning according to claim 1, characterized in that: The text detection dataset includes the Total-Text dataset, the ICDAR 2015 dataset, and the CTW1500 dataset.

3. A method for detecting text in complex scenarios based on deep learning according to claim 1, characterized in that: In the global context convolutional module, if the input image size is C×W×H, where C is the number of channels, and W and H are the width and height of the input image respectively; then the output size of the global average pooling layer is C×1×1; The output feature size of the first 1×1 convolutional layer is K1×1×1, where K1 is the number of output channels of the first 1×1 convolutional layer; The output feature size of the 1×11 depthwise separable convolutional layer is K2×W×H, where K2 is the number of output channels of the 1×11 depthwise separable convolutional layer; The output feature size of the 11×1 depthwise separable convolutional layer is K3×W×H, where K3 is the number of output channels of the 11×1 depthwise separable convolutional layer; The output feature size of the second 1×1 convolutional layer is K4×W×H, where K4 is the number of output channels of the second 1×1 convolutional layer; The sigmoid activation function layer is used to compress the values of its input features into the interval [0,1], and its output feature size is K4×W×H.

4. A method for detecting text in complex scenes based on deep learning according to claim 3, characterized in that: K1 = C, K2 = C / 2, K3 = C / 2, K4 = C.

Citation Information

Patent Citations

  • Scene text detection method, system and equipment based on deep convolutional neural network

    CN114724155A

  • Efficient and accurate ambiguous scene character detection method and system

    CN117765520A