An Arbitrary Shape Scene Text Detection Method Based on Contour Feature Enhancement
By combining Swin Transformer and convolutional neural network, the perception ability of text areas is enhanced, and the orthogonal text attention module is used to enhance feature differences, the problems of edge contour missing and misdetection in text detection are solved, and more accurate and efficient text detection is achieved.
Patent Information
- Application Number
- CN202411519578.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-29
AI Technical Summary
Existing text detection methods can easily lead to missing edge contours and mis-checking when processing text with arbitrary shapes.
Using a contour feature enhancement method, by combining Swin Transformer and convolutional neural network, the perception ability of text areas is enhanced, the loss of edge information is compensated, and the orthogonal text attention module is used to enhance the feature difference between text and background.
Effectively compensate for the lack of text edge information, reduce error detection, and improve the accuracy and efficiency of text detection.
Smart Images

Figure CN119296094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and specifically to a method for detecting arbitrary-shaped scene text based on contour feature enhancement. Background Art
[0002] The core content of computer vision is to process and analyze images or videos to complete deeper tasks such as classification, detection, localization, and recognition. As a classic task in the field of computer vision, the purpose of text detection is to detect text regions in natural images and locate them with bounding boxes or curves.
[0003] The text information widely distributed in natural scene images is a unique and important information source in the images. Extracting the text information in the images has important research value and application value for the analysis and understanding of the images. In recent years, text detection has been widely applied in fields such as autonomous driving, image understanding / retrieval, and machine translation. In the field of autonomous driving: The detection and recognition of traffic signs are the key to accurately judging the road conditions. Detecting traffic signs, license plates, and any text information in the scene helps the vehicle accurately navigate and locate the destination. In the field of image understanding / retrieval: By detecting and recognizing the text information in the scene image, it helps to understand the surrounding environment and objects, find the uniqueness between different objects, and further retrieve the desired object from the environment, improving the efficiency and accuracy of image retrieval. In the field of machine translation: In recent years, various machine translation software have emerged in an endless stream. Among them, reliably detecting and recognizing text from smartphone images can be more conveniently input into the translation software, bringing convenience to people's lives. Thus, it can be seen that the natural scene text detection technology has profound research significance and practical value.
[0004] In recent years, with the development of deep learning in the field of computer vision, more and more scholars have paid attention to the field of text detection and achieved excellent results. The text detection methods based on deep learning can be divided into regression-based text detection algorithms, transformer-based text detection algorithms, and segmentation-based text detection algorithms. Limited by the shape of the fixed text box, it is difficult for the regression-based text detection algorithm to obtain excellent results for detecting arbitrary-shaped text; the segmentation-based text detection algorithm locates the text region through pixel-level prediction and is not restricted by the text shape, and has become one of the mainstream algorithms in the field of text detection. Subsequently, with the development and application of transformers in the field of object detection, considering that text detection is a special object detection, the transformer-based text detection algorithm has also become the focus in the field of text detection.
[0005] The text detection algorithm based on segmentation is mainly inspired by semantic segmentation. This algorithm first uses the network to generate pixel-level predictions, discriminates each pixel point, and then reconstructs text instances through pixel-level predictions and specific post-processing. To better process the features extracted by the backbone network and avoid the adhesion of adjacent text regions, in 2019, the Progressive Scale Expansion Network (PSENet) proposed by Wang et al. expanded the detection region from a small kernel to a large kernel, and finally fused to form a complete text region, effectively separating text instances close to each other. The Pixel Aggregation Network (PAN) designed an efficient and accurate arbitrary-shaped text detector through a lightweight segmentation network. This method uses an efficient segmentation head to refine the features, predicts the similarity vector between the text kernel and the surrounding pixels, and then aggregates the pixels belonging to the same text kernel through a learnable pixel aggregation algorithm. In 2020, the Differentiable Binarization Network (DBNet) proposed by Liao et al. inserted a differentiable binarization module into the segmentation network. The binarization module can be trained along with the network, which not only has a simple network structure and simplified post-processing, but also achieves good performance in real-time scene text detection. The Text Detection via Segmentation with Probability Maps (TextPMs) proposed a detection method for arbitrary-shaped text based on probability maps. This method uses the probability maps of a set of functions to describe the possible probability distributions as comprehensively as possible, optimizes the prediction of the probability maps by iteratively learning the mapping relationship between continuous probability distributions, and then aggregates the probability maps to generate complete text instances.
[0006] The text detection algorithm based on Transformer is mainly inspired by the application of Transformer in object detection. The initial framework of Transformer in object detection, DETR (Detection Transformer), regards object detection as a direct set prediction problem and proposes a simple end-to-end framework without anchor generation and complex post-processing. The text detection task can be regarded as a special object detection problem. In 2021, the text detection method based on Transformer introduced a loss function for multi-directional text detection, which mainly uses quadrilateral-based prediction to represent the text region. The Feature Sampling and Grouping for Scene Text Detection (FSG) proposed a scene text detection framework based on feature sampling and grouping, avoiding background interference and reducing computational complexity. In 2023, Ye et al. designed an efficient dynamic point text detection Transformer network (Text Detection with Dynamic Points in Transformer, DPText-DETR), which uses explicit point coordinates to generate position queries and updates them dynamically in a progressive manner to enhance training convergence.
[0007] Although the above text detection algorithms have achieved good performance, these methods still face some problems. The text detection method based on Transformer is affected by the self-attention mechanism and position encoding, usually having a high computational complexity, which affects the detection efficiency of the entire model. The text detection method based on segmentation usually faces two problems. The first problem is that using multiple convolutional layers will result in unclear text edge information, thus causing the loss of the text boundary contour. Another difficulty is that when using the Feature Pyramid Network (FPN) to fuse multi-level features from top to bottom, it will blur the significant features between the text region and the background, which may cause the network to misclassify non-text elements as text instances, resulting in false detections.
[0008] To solve the above problems, the present invention proposes an arbitrary-shaped scene text detection method based on contour feature enhancement. Summary of the Invention
[0009] The object of the present invention is to propose a method for detecting arbitrary-shaped scene text based on contour feature enhancement to solve the problems of missing edge contours and misdetection in text detection. First, the present invention combines the swin transformer and the convolutional neural network, uses the rich detail information in the low-level feature map to enhance the network's perception of the text area, and compensates for the missing edge contours. In addition, the present invention also uses the gradient information enhancement of the text edge to distinguish the text area from the text-like area, thereby avoiding the occurrence of misdetection. The present invention provides a research solution for the deep learning method of natural scene text detection, and promotes the development of visual tasks such as image understanding to a certain extent.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] A method for detecting arbitrary-shaped scene text based on contour feature enhancement, the method is implemented by an arbitrary-shaped scene text detection system based on contour feature enhancement, the system is composed of a backbone network, a multi-scale feature fusion module and an output module, wherein, the multi-scale fusion module includes a feature compensation module (FCM) and an orthogonal text attention module (OTAM);
[0012] The method includes the following steps:
[0013] S1. Use the backbone network to extract pyramid features, and input the outputs of each stage into the multi-scale feature fusion module;
[0014] S2. Adopt a top-down strategy to fuse features from different levels, input the features of the lowest level into the feature compensation module (FCM), combine the lightweight swin transformer and the convolutional neural network to enhance the model's perception of the text area with lower computational complexity, and compensate for the missing text edge information;
[0015] S3. Use the attention module to adjust the information to enhance the interaction of the feature map information in the channel dimension;
[0016] S4. Perform top-down feature fusion on the newly obtained multi-layer pyramid features, then perform upsampling through 3×3 convolution and bilinear interpolation of different ratios, adjust the resolution of the multi-scale feature map to the same size, and adjust the number of channels to 64; then splice the multi-level feature maps along the channel dimension to obtain the feature map F;
[0017] S5. Use two parallel branches of the orthogonal text attention module (OTAM) to respectively learn the attention weights in the orthogonal directions of the fused features, and then adopt residual connection to avoid the loss of detail information;
[0018] S6. Obtain a pair of orthogonal edge texture feature maps based on the operation described in S5, concatenate them along the channel dimension to obtain a new feature map F C , and then adopt a non-linear activation module to assist in gradient propagation to obtain feature F*;
[0019] S7. For the orthogonal feature maps obtained by the processing in S6, first divide them into two parts along the channel, namely α times the channel and (1 - α) times the channel, and then process them through an activation module to generate attention weights; use a convolution with a kernel size of 1×1 to compress the number of channels to improve computational efficiency; finally, multiply the attention weights by the input feature map and perform a residual connection to obtain a feature map with enhanced edge information;
[0020] S8. Input the feature map obtained in S7 into an output module for post-processing. The output module consists of two parallel branches with the same structure, and each branch consists of two transposed convolution functions for predicting the text probability map f P and the threshold map f T , and then obtain a binary map through a differentiable binarization function, and thus obtain the predicted bounding box.
[0021] Preferably, the S1 specifically includes the following content:
[0022] Adopt ResNet-50 as the backbone network to extract pyramid features, and input the outputs {F 2 , F 3 , F 4 , F 5} of stage 2, stage 3, stage 4, and stage 5 into a multi-scale feature fusion module. The resolution sizes of the multi-scale features are {1 / 4, 1 / 8, 1 / 16, 1 / 32} of the input image, and the number of channels is 256.
[0023] Preferably, the S2 specifically includes the following content:
[0024] First pass the input high-resolution feature {F 2} through two serial swin transformer blocks (STB). The swin transformer block (STB) uses window attention mechanism and multi-head self-attention mechanism to enhance the expression ability of the model;
[0025] In the second block, based on the shifted window method, further introduce boundary embedding (BE) to encode the initial pixels of the feature map before segmentation and embed the encoded information into the model to make the model more easily perceive text information. The process of a single swin transformer block (STB) for processing features is shown as follows:
[0026] F M = F E + f MSA (f LN (F E )
[0027] F S = F M + f LN (f MLP (F M )
[0028] Among them, f MSA represents multi-head self-attention; f LN represents layer normalization; F E represents the input feature after layer normalization and the multi-head self-attention module; F M represents the output feature after layer normalization and the multi-head self-attention module; f MLP represents the multi-layer perceptron module; F S represents the output of the swin transformer block.
[0029] Preferably, the function of adjusting the information by using the attention module in S3 is as follows:
[0030] F' 2 = F S × σ(Conv k×k (f GAP (F S )))
[0031] Among them, σ represents the Sigmoid activation function; f GAP represents global average pooling; Conv k×k (·) represents a two-dimensional convolution with a convolution kernel size of k×k, and "×" represents element-wise multiplication.
[0032] Preferably, the newly obtained multi-layer pyramid features in S4 are denoted as {F' 2 , F 3 , F 4 , F 5}; the fused features are denoted as {P 2 , P 3 , P 4 , P 5}; the functions of feature fusion and feature splicing are as follows:
[0033]
[0034] F = [f up×8 (P 5 ), f up×4 (P 4 ), fup×2 (P 3 ),Conv 3×3 (P 5 )]
[0035] where f up×N (·) represents the bilinear interpolation upsampling operation with a scaling factor of N; Conv 3×3 (·) represents a 3×3 convolutional layer; [·] represents a special concatenation along the channel dimension.
[0036] Preferably, the S5 specifically includes the following:
[0037] In the left branch of the Orthogonal Text Attention Module (OTAM), global average pooling is used for feature aggregation in the horizontal direction, only focusing on the texture features along the horizontal direction, and then the feature map is expanded to the original size; in the right branch of the Orthogonal Text Attention Module (OTAM), the same operation is taken to process the feature map in the vertical direction; the functional representation of the above feature processing process is as follows:
[0038] F h = f ex (f GAP_h (F))
[0039] F v = f ex (f GAP_v (F))
[0040] where f ex (·) represents the dilation operation; f GAP_h (·) and f GAP_v (·) respectively represent global average pooling in the horizontal and vertical directions; F h and F v represent the output features in the orthogonal direction.
[0041] Preferably, the functional representation of the S6 is:
[0042] F C = [F h , F v
[0043] F * = ReLU(f BN (Conv 1×1 (F C )))
[0044] where [·] represents feature concatenation along the channel dimension; Conv 1×1 (·) represents a convolution with a kernel size of 1×1; f BN (·) represents the batch normalization operation; ReLU(·) represents the ReLU activation function.
[0045] Preferably, the function of S7 is expressed as:
[0046]
[0047] F O = F + F × F S1 × F S2
[0048] where slice(·) represents a slicing operation; σ(·) represents a sigmoid activation function; Conv 1×1 (·) represents a convolution with a kernel size of 1×1; "×" represents element-wise multiplication.
[0049] Preferably, the differentiable binarization function in S8 is specifically:
[0050]
[0051] where (i,j) represents the coordinate point of a pixel in the feature map.
[0052] Compared with the prior art, the present invention provides a method for detecting arbitrary-shaped scene text based on contour feature enhancement, having the following beneficial effects:
[0053] The present invention proposes a method for detecting arbitrary-shaped scene text based on contour feature enhancement, aiming to solve the problem of blurred contours in text detection. The present invention includes a Feature Compensation Module (FCM) and an Orthogonal Text Attention Module (OTAM); wherein, the Feature Compensation Module FCM uses a lightweight swin transformer combined with boundary embedding to enrich the detailed information of the feature map, and then combines the attention module to compensate the text contour information in the spatial and channel dimensions of the feature map respectively; then, the Orthogonal Text Attention Module OTAM processes the fused feature map to enhance the feature difference between the text and the background, thereby reducing the occurrence of false positives and false negatives, enabling the network to accurately locate the text boundary. A large number of experiments on four text detection benchmark datasets have demonstrated the effectiveness and advantages of the method proposed by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is the overall network structure diagram of a method for detecting arbitrary-shaped scene text based on contour feature enhancement mentioned in Embodiment 1 of the present invention;
[0055] Figure 2This is the detailed structure diagram of the Feature Compensation Module (FCM) mentioned in Embodiment 1 of the present invention;
[0056] Figure 3 This is the detailed structure diagram of the Orthogonal Text Attention Module (OTAM) mentioned in Embodiment 1 of the present invention;
[0057] Figure 4 This is the comparison chart of the detection results of the present method and other methods mentioned in Embodiment 2 of the present invention. Detailed implementation manners
[0058] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0059] First, the abbreviations and key terms defined in the present invention are described as follows:
[0060] PSENet: Progressive Scale Expansion Network
[0061] PAN: Pixel Aggregation Network
[0062] DBNet: Differentiable Binarization Network
[0063] TextPMs: Text Detection via Segmentation with Probability Maps
[0064] FSG: Feature Sampling and Grouping for Scene Text Detection
[0065] DPText-DETR: Text Detection with Dynamic Points in Transformer
[0066] FCM: Feature Compensation Module
[0067] OTAM: Orthogonal Text Attention Module, the orthogonal text attention module
[0068] Based on the above content, specific examples are described below.
[0069] Example 1:
[0070] The present invention proposes a method for detecting arbitrary-shaped scene text based on contour feature enhancement. Its overall network structure is as Figure 1 shown. It consists of three parts: a backbone network, a multi-scale feature fusion module, and an output module. Among them, the multi-scale feature fusion module includes a feature compensation module (Feature Compensation Module, FCM) and an orthogonal text attention module (Orthogonal Text Attention Module, OTAM).
[0071] Specifically, first, ResNet-50 is used as the backbone network to extract pyramid features. The outputs {F 2 , F 3 , F 4 , F 5} of stages 2, 3, 4, and 5 are input into the multi-scale feature fusion module. The resolution sizes of the multi-scale features are {1 / 4, 1 / 8, 1 / 16, 1 / 32} of the input image respectively, and the number of channels is 256 for all. Then, a top-down strategy is used to fuse features from different levels. However, simply using upsampling for top-down feature fusion usually leads to the loss of text edge information and affects the positioning of text boundaries. Therefore, in the present invention, the lowest-level feature is input into the feature compensation module (FCM), and a lightweight swin transformer and a convolutional neural network are combined to enhance the model's perception of the text area with a relatively low computational complexity and compensate for the lack of text edge information.
[0072] The detailed structure of the feature compensation module FCM is as Figure 2 shown. The input high-resolution feature {F 2}First, it passes through two serial Swin Transformer blocks (STBs). The Swin Transformer utilizes the window attention mechanism and the multi-head self-attention mechanism to enhance the model's expressive power. In the second block, the shifted window method is adopted, and our model sets the window size to 4×4 to reduce the computational complexity. However, the method of window partitioning may lead to discontinuity of the text within the same instance, which is not conducive to the network's long-range modeling and thus affects the effectiveness of text detection. Although stacking more Swin Transformer blocks can partially alleviate this problem, it will greatly increase the computational complexity of the network. To solve this problem, Boundary Embedding (BE) is introduced, which encodes the initial pixels of the feature map before segmentation and embeds the encoded information into the model, making it easier for the model to perceive text information. The process of a single STB processing features is as follows:
[0073] F M =F E +f MSA (f LN (F E ))
[0074] F S =F M +f LN (f MLP (F M ))
[0075] Among them, f MSA represents multi-head self-attention, f LN represents layer normalization, F E and F M represent the input features and output features passing through the layer normalization and multi-head self-attention modules respectively, f MLP represents the multi-layer perceptron module, and F S represents the output of the STB.
[0076] After that, to enhance the interaction of the feature map information in the channel dimension, the attention module is used to adjust the information. The specific steps can be expressed as follows:
[0077] F' 2 =F S ×σ(Conv k×k (f GAP (F S )))
[0078] Among them, σ represents the Sigmoid activation function, f GAP represents global average pooling, and Conv k×k(·) represents a two-dimensional convolution with a convolution kernel size of k×k, and "×" is element-wise multiplication.
[0079] Then, for the newly obtained multi-layer pyramid features {F′ 2 , F 3 , F 4 , F 5}, top-down feature fusion is performed to obtain {P 2 , P 3 , P 4 , P 5}. After that, through 3×3 convolution and bilinear interpolation at different scales, the resolution of the multi-scale feature map is adjusted to the same size, and the number of channels is adjusted to 64. Subsequently, the multi-level feature maps are concatenated along the channel dimension to obtain the feature map F. Feature fusion and feature concatenation are as follows:
[0080]
[0081] F = [f up×8 (P 5 ), f up×4 (P 4 ), f up×2 (P 3 ), Conv 3×3 (P 5 )]
[0082] Among them, f up×N (·) represents the bilinear interpolation upsampling operation with a scaling factor of N, and Conv 3×3 (·) represents the 3×3 convolutional layer, and [·] represents feature concatenation along the channel dimension.
[0083] As is well known, the feature difference between text edge information and complex background is the key to accurately locating the text region. However, using a convolution with a kernel size of k×k when fusing multi-level features in the top-down path will lead to the blurring of edge gradient information. In addition, objects with similar texture features to the text region will lead to false positive situations. Therefore, it is necessary to enhance the text contour features. The orthogonal text attention module (OTAM) proposed in the present invention adopts two parallel branches to separately learn the attention weights in the orthogonal directions of the fused features, and then uses residual connections to avoid the loss of detailed information.
[0084] The detailed structure of the orthogonal text attention module OTAM is as Figure 3As shown, in the left branch, global average pooling is adopted in the horizontal direction for feature aggregation, only focusing on the texture features along the horizontal direction. Subsequently, the feature map is expanded to the original size. In the right branch, the same operation is performed on the feature map in the vertical direction. In this way, the model can perceive the text boundary information from two directions, thereby improving the ability of the present invention to locate text instances. The entire feature processing process is as follows:
[0085] F h = f ex (f GAP_h (F))
[0086] F v = f ex (f GAP_v (F))
[0087] wherein, f ex (·) represents the dilation operation; f GAP_h (·) and f GAP_v (·) respectively represent global average pooling in the horizontal and vertical directions; F h and F v represent the output features in the orthogonal directions.
[0088] After obtaining a pair of orthogonal edge texture feature maps, first, they are concatenated along the channel dimension to obtain the feature map
[0089] F C , and subsequently, a non-linear activation module is adopted to assist in gradient propagation. The process of obtaining the feature F * is as follows:
[0090] F C = [F h , F v
[0091] F * = ReLU(f BN (Conv 1×1 (F C )))
[0092] wherein, [·] represents feature concatenation along the channel dimension, Conv 1×1 (·) represents convolution with a kernel size of 1×1, f BN (·) is the batch normalization operation, and ReLU(·) represents the ReLU activation function.
[0093] For the obtained orthogonal feature maps, first divide them into two parts along the channels, namely α times the channels and (1 - α) times the channels, and then process them through an activation module to generate attention weights. In the method of the present invention, α is set to 0.5. In this part, in order to improve the computational efficiency, a convolution with a kernel size of 1×1 is used to compress the number of channels. Finally, by multiplying the attention weights with the input feature map and performing a residual connection, a feature map with enhanced edge information is obtained. By processing the fused features in the orthogonal direction, the feature distinctiveness between the text region and the background can be effectively enhanced, which helps to more accurately predict the text region, thereby improving the detection accuracy of the network. The overall process of this part is as follows:
[0094]
[0095] F O = F + F × F S1 × F S2
[0096] where slice(·) represents the slicing operation, σ(·) represents the sigmoid activation function, Conv 1×1 (·) represents the 1×1 convolution, and "×" is the element-wise multiplication.
[0097] Finally, input
[0098] F O into the output module for post-processing. The output module of the present invention consists of two parallel branches with the same structure. Each branch is mainly composed of two deconvolution functions, which are used to predict the text probability map f P and the threshold map f T , and then obtain the binary map through a differentiable binarization function, and thus obtain the predicted bounding box. The binarization process is as shown in the following formula:
[0099]
[0100] where (i, j) represents the coordinate point of the pixel in the feature map.
[0101] Example 2:
[0102] Based on Example 1 but with differences. Specifically, the following design of specific experiments is carried out to verify the performance of the method for detecting arbitrary-shaped scene text based on contour feature enhancement proposed by the present invention. The specific content is as follows.
[0103] 1. Dataset
[0104] The method for detecting arbitrary-shaped scene text based on contour feature enhancement proposed in Example 1 is experimentally verified on four benchmark datasets. The datasets used include:
[0105] (1) Total-Text is a dataset mainly for arbitrary-shaped text detection. This dataset is annotated with polygon bounding boxes and consists of 1,255 training images and 300 test images. These images contain horizontal, multi-directional, and curved text.
[0106] (2) CTW1500 is a dataset mainly composed of curved text, which covers multiple languages mainly in Chinese and English. This dataset contains 1,000 training images and 500 test images, and the text instances are annotated with polygons having 14 vertices.
[0107] (3) MSRA-TD500 is a multi-language dataset with complex backgrounds and multi-directions. It includes 300 training images and 200 test images. Referring to previous methods, 400 training images from HUST-TR400 were also used in the training phase of the present invention.
[0108] (4) The ICDAR2015 dataset was collected by Google Glass and contains blurred text instances with different scales and directions. This dataset contains 1,000 training images and 500 test images, and the text instances are annotated with rectangular bounding boxes.
[0109] (5) SynthText is a large dataset that contains approximately 850,000 synthetic text images. This dataset includes different background images from different scenarios and is annotated with characters, words, and lines. The method proposed in the present invention is only pre-trained on this dataset.
[0110] 2. Experimental Setup
[0111] The entire experimental process is divided into a pre-training phase and a fine-tuning phase. First, pre-training is carried out on the synthetic dataset SynthText for 2 epochs with a fixed learning rate of 0.015 and a batch size set to 8. The fine-tuning phase is trained on four real-scene datasets with an initial learning rate set to 0.015 for a total of 1,200 epochs. The entire network is trained using the SGD optimizer. Before training, data augmentation is performed on the images, and the training images are randomly cropped to a size of 640×640. In addition, random horizontal flipping and random rotation are performed within the range of (-10°, 10°). The model proposed in the present invention is implemented based on the Python language, and the experimental environment is shown in Table 1:
[0112] Table 1: Experimental Environment
[0113]
[0114] 3. Experimental Results
[0115] The method proposed in the present invention was compared with previous scene text detection methods on four benchmark datasets. Experiments were conducted on multi-directional text datasets (MSRA-TD500 and ICDAR2015), arbitrary-shaped text datasets (Total-Text), and multi-language curved text datasets (CTW1500). The performance of the present invention was evaluated from four metrics: precision, recall, h-mean, and frames per second (FPS). Ablation experiments were also conducted on the MSRA-TD500 dataset and the CTW1500 dataset to illustrate the roles of the various components of the present invention.
[0116] Tables 2, 3, 4, and 5 show the quantitative evaluation results of each method on the text datasets. Among them, the best performance in each metric is shown in bold. As can be seen from Tables 2 and 3, the method proposed in the present invention can well balance detection accuracy and speed in the multi-directional text detection dataset. Especially on the MSRA-TD500 dataset, the three metrics of the method of the present invention all reached the highest performance, fully demonstrating the performance of the method of the present invention. Tables 4 and 5 show the detection results of the method of the present invention on the arbitrary-shaped text dataset and the multi-language curved text dataset. The method of the present invention achieved the highest recall rate on these two curved text datasets. The superiority of the technology of the present invention benefits from the fact that the method can effectively enhance the text contour features, distinguish the text region from the text-like background, and at the same time, can also compensate for the text edge information to obtain a more accurate text contour.
[0117] Table 2: Experimental results on the MSRA-TD500 dataset, with the best performance marked in bold
[0118]
[0119] Table 3: Experimental results on the ICDAR2015 dataset, with the best performance marked in bold
[0120]
[0121] Table 1: Experimental results on the Total-Text dataset, with the best performance marked in bold
[0122]
[0123] Table 2: Experimental results on the CTW1500 dataset, with the best performance marked in bold
[0124]
[0125]
[0126] To more intuitively prove the effectiveness of the method proposed by the present invention, the detection results of the method proposed by the present invention are visually compared with those of the classical text detection methods DBNet and DBNet++. The ground truth and the visualization results of each detection method are as Figure 4 shown. It can be clearly seen from Figure 4 that the method proposed by the present invention can compensate for the boundary information of the text region and ensure the accurate positioning of the text region. In addition, the method of the present invention can also enhance the text contour features and prevent misdetecting complex backgrounds in scene images as text instances. As Figure 4 (b) shows, the method proposed by the present invention can effectively detect multi-scale texts and texts of arbitrary shapes.
[0127] Table 6 conducts ablation experiments on the main modules of the method of the present invention on the CTW1500 dataset. It should be noted that when verifying the effectiveness of the orthogonal text attention module, the initial learning rate is set to 0.007. It can be seen from Table 6 that compared with the baseline network, the average precision of the model with the feature compensation module (FCM) added has increased by 2.05%. This shows that the FCM can enrich the detailed features, enhance the network's perception of the text region, thereby compensating for the loss of contour information in the feature fusion process and more accurately locating the text region. When only the orthogonal text attention module (OTAM) is added to the baseline network, the average precision of the model increases by 1.81%. This shows that the OTAM proposed by the present invention effectively captures text edge information from two orthogonal directions, enhances the contrast between the text region and the complex background, helps to accurately distinguish texts and similar elements, and enables accurate detection of text instances. When both SFCM and OTAM are added to the baseline network, the precision, recall rate, and average precision of the model increase by 5.3%, 1.77%, and 3.52% respectively. The experimental results fully prove the effectiveness of the feature compensation module (FCM) and the orthogonal text attention module (OTAM) proposed by the present invention.
[0128] Table 6: Ablation experiment results of the main modules, with the best performance marked in bold
[0129]
[0130] Tables 7, 8, and 9 conduct ablation experiments on each component of the main module. The feature compensation module (FCM) mainly adds boundary embedding and attention modules on the basis of the basic Swin Transformer block. Table 7 conducts ablation experiments on the CTW1500 dataset, illustrating the role of each component in the FCM for text contour feature compensation. Table 8 illustrates the influence of different-level feature maps on the FCM feature compensation effect when fusing features from top to bottom, where F 5 represents the high-level feature, F2 Denote low-level features. It can be seen from the experimental results that on the CTW1500 dataset, compared with using high-level features, the average precision of the model using low-level features has increased by 1.01%, which fully proves the effectiveness of FCM. Table 9 shows the ablation experiments on the MSRA-TD500 dataset to illustrate the role of the slicing operation and the division ratio when dividing along the channels in the orthogonal text attention module (OTAM).
[0131] Table 7: Ablation experiment results of the feature compensation module, with the best performance marked in bold
[0132]
[0133] Table 8: Ablation experiment results of FCM for processing different-level feature maps, with the best performance marked in bold
[0134]
[0135] Table 9: Ablation experiment results of the slicing operation and the division ratio in OTAM, with the best performance marked in bold, w / o means without the slicing operation
[0136]
[0137] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for detecting text in arbitrary shape scenes based on contour feature enhancement, characterized in that: The method is implemented by an arbitrary shape scene text detection system based on contour feature enhancement, the system is composed of a backbone network, a multi-scale feature fusion module and an output module, wherein the multi-scale fusion module includes a feature compensation module and an orthogonal text attention module; The method comprises the following steps: S1, use the backbone network to extract pyramid features and input the output of each stage into the multi-scale feature fusion module; S2, adopt a top-down strategy to fuse features from different levels, input the lowest level features into the feature compensation module, and use a lightweight swin transformer and convolutional neural network to enhance the model's perception of text areas with low computational complexity and compensate for the lack of text edge information; S3, using the attention module to adjust the information to enhance the interaction of feature map information in the channel dimension; S4, perform top-down feature fusion on the newly obtained multi-layer pyramid features, and then upsample them through 3×3 convolution and bilinear interpolation of different proportions, adjust the resolution of multi-scale feature maps to the same size, and adjust the number of channels to 64; then splice the multi-level feature maps along the channel dimension to obtain the feature map F; S5, using the two parallel branches of the orthogonal text attention module to learn the attention weights of the orthogonal directions of the fusion features respectively, and then using residual connections to avoid the loss of detail information; S6: Based on the S5 operation, a pair of orthogonal edge texture feature maps are obtained, and they are spliced along the channel dimension to obtain a new feature map F. C , then a nonlinear activation module is used to help gradient propagation to obtain the feature F*; S7, for the orthogonal feature map obtained by S6, first divide it into two parts along the channel: α times the channel and (1-α) times the channel, and then process it through the activation module to generate attention weights; use convolution with a kernel size of 1×1 to compress the number of channels to improve computational efficiency; finally, multiply the attention weights with the input feature map and perform residual connection to obtain a feature map with enhanced edge information; S8, input the feature map obtained in S7 to the output module for post-processing, wherein the output module is composed of two parallel branches with the same structure, each branch is composed of two deconvolution functions, and is used to predict the text probability map f P and threshold map f T , and then a binary mapping is obtained through a differentiable binarization function, and the predicted bounding box is obtained from it.
2. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 1, characterized in that: The S1 specifically includes the following contents: ResNet-50 is used as the backbone network to extract pyramid features. The outputs {F2, F3, F4, F5} of stage 2, stage 3, stage 4 and stage 5 are input into the multi-scale feature fusion module. The resolution of the multi-scale features is {1 / 4, 1 / 8, 1 / 16, 1 / 32} of the input image, and the number of channels is 256.
3. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 2, characterized in that: The S2 specifically includes the following contents: The input high-resolution features{F2}first pass through two serial swin transformer blocks, which use the window attention mechanism and multi-head self-attention mechanism to enhance the expressiveness of the model; In the second block, based on the shift window method, boundary embedding is further introduced to encode the initial pixels of the feature map before segmentation, and the encoded information is embedded into the model to make it easier for the model to perceive text information; the process of processing features by a single swintransformer block is as follows: F M =F E +f MSA (f LN (F E )) F S =F M +f LN (f MLP (F M )) Among them, f MSA represents long self-attention; f LN Representation layer normalization; F E represents the input features after layer normalization and multi-head self-attention module; F M represents the output features after layer normalization and multi-head self-attention module; f MLP Represents a multi-layer perceptron module; F S Represents the output of the Swin transformer block.
4. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 3, characterized in that: The information is adjusted using the attention module as described in S3, and its function is expressed as follows: F′2=F S ×σ(Conv k×k (f GAP (F S ))) Among them, σ represents the Sigmoid activation function; f GAP Represents global average pooling; Conv k×k (·) represents a two-dimensional convolution with a kernel size of k×k, and "×" represents element-by-element multiplication.
5. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 4, characterized in that: The newly obtained multi-layer pyramid features described in S4 are recorded as {F2′, F3, F4, F5}; the fused features are recorded as {P2, P3, P4, P5}; the functions of feature fusion and feature concatenation are expressed as follows: F=[f up×8 (P5),f up×4 (P4),f up×2 (P3),Conv 3×3 (P5)] Among them, f up×N (·) represents the bilinear interpolation upsampling operation with a scaling factor of N; Conv 3×3 (·) indicates a 3×3 convolutional layer; [·] indicates special concatenation along the channel dimension.
6. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 5, characterized in that: The S5 specifically includes the following contents: In the left branch of the orthogonal text attention module, global average pooling is used for feature aggregation in the horizontal direction, focusing only on the texture features along the horizontal direction, and then the feature map is expanded to the original size; in the right branch of the orthogonal text attention module, the same operation is performed in the vertical direction to process the feature map; the function of the above feature processing process is expressed as follows: F h =f ex (f GAP_h (F)) F v =f ex (f GAP_v (F)) Among them, f ex (·) indicates expansion operation; f GAP_h (·) and f GAP_v (·) represents global average pooling in horizontal and vertical directions respectively; F h and F v Represents the output features in orthogonal directions.
7. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 6, characterized in that: The function of S6 is expressed as: F C =[F h ,F v ] F * =ReLU(f BN (Conv. 1×1 (F C ))) Among them, [·] indicates feature concatenation along the channel dimension; Conv 1×1 (·) indicates a convolution with a kernel size of 1×1; f BN (·) represents the batch normalization operation; ReLU(·) represents the ReLU activation function.
8. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 7, characterized in that: The function of S7 is expressed as: F1 * ,F2 * =slice(F * ) F S1 =σ(Conv 1×1 (F1 * )) F S2 =σ(Conv 1×1 (F2 * )) F O =F+F×F S1 ×F S2 Where slice(·) represents the slicing operation; σ(·) represents the sigmoid activation function; Conv 1×1 (·) indicates convolution with a kernel size of 1×1; "×" indicates element-wise multiplication.
9. The method for detecting text in an arbitrary shape scene based on contour feature enhancement according to claim 8, characterized in that: The differentiable binarization function described in S8 is specifically: Among them, (i, j) represents the coordinate point of the pixel in the feature map.
Citation Information
Patent Citations
Scene text detection method, system and equipment based on deep convolutional neural network
CN114724155A
Sentence-level relation extraction method and device based on deep feature fusion
CN115640374A