Text detection method and system based on cross-level feature enhancement and auxiliary feature guidance

By adopting cross-level feature enhancement and auxiliary feature guidance techniques in text detection methods, the problem of large scale differences and extreme aspect ratio text detection in natural scene images is solved, and higher detection accuracy and performance are achieved.

CN120126110AActive Publication Date: 2025-06-10TIANJIN UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510175820.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Existing text detection methods are difficult to accurately detect text instances with large scale differences and extreme aspect ratios in natural scene images.

Method used

Using a method based on cross-level feature enhancement and auxiliary feature guidance, the receptive field of low-level features is dynamically adjusted, the edge high-frequency information of high-level features is enhanced, and the auxiliary feature guidance network is used to differentiate different text instances.

Benefits of technology

The positioning accuracy of texts of different scales is improved, and the problem of extreme aspect ratio text being wrongly divided into multiple text areas is avoided, and the detection performance of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126110A_ABST
    Figure CN120126110A_ABST
Patent Text Reader

Abstract

The invention discloses a text detection method and system based on cross-level feature enhancement and auxiliary feature guidance, and relates to the technical field of computer vision. The method comprises the following steps: extracting feature maps of different levels of an image to be detected; for the low-level features, a receptive field is dynamically adjusted by adjusting the size of a convolution kernel to obtain multi-scale information, and enhanced features of the low-level features are generated; for the high-level features, enhancing text edge high-frequency information by adopting differential convolution, and generating enhanced features of the high-level features; fusing features from different levels in a cross-level feature enhancement mode to generate fused features; semantic information of the high-level features is extracted, and auxiliary features are generated; for the generated auxiliary features and fusion features, discriminating the attribution of text kernel pixels by adopting an attention guidance mode to obtain features after attention guidance; and performing text-non-text prediction on the pixels to generate a text bounding box. According to the scheme, the network can more correctly segment the text region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a text detection method and system based on cross-level feature enhancement and auxiliary feature guidance. Background Art

[0002] As one of the important carriers for human communication and knowledge dissemination, text plays an important role in life, production, and learning. Natural scene images contain rich text information, which plays an important role in helping people understand the surrounding environment. In recent years, with the gradual maturity of artificial intelligence theory and the continuous development of computer vision devices, computer vision technology has attracted more and more attention. Accurately extracting text information from pictures helps to understand scene images. As the first step in obtaining text information in images, accurately and quickly locating text regions contributes to the efficient progress of subsequent tasks such as text recognition, image layout analysis, and machine translation.

[0003] Currently, text detection technology has been widely applied in various aspects of life. In the field of autonomous driving: Devices such as cameras and sensors installed on autonomous vehicles can capture surrounding scene images, and then text information such as indicator signs, billboards, and license plates in the images can be obtained through text detection and text recognition technologies, which helps the vehicle judge the surrounding environment and thus improve the safety of autonomous driving. In the field of real-time photo translation: With the continuous improvement of people's living standards, traveling abroad has become a popular way of traveling. Many translation software provide photo translation services. After reliably detecting and recognizing text from smartphone images, these software can translate text information in real time, which can provide great convenience for people's travel. Therefore, as one of the basic tasks in the field of computer vision technology, natural scene text detection technology has important research value and broad application prospects.

[0004] The main task of text detection is to detect text regions in natural scene images and label them with rectangular or polygonal bounding boxes. Early traditional text detection methods mainly studied the feature differences between text and background. With the continuous development of deep learning technology, text detection methods based on deep learning have gradually received extensive attention. These methods can be divided into regression-based natural scene text detection methods, contour fitting-based text detection methods, and segmentation-based text detection methods. Regression-based text detection methods generate a prediction score map through a fully convolutional network, and then directly regress the bounding boxes of text regions according to the score map. In 2017, Zhou et al. proposed an Efficient and Accurate Scene Text detector (EAST), which directly predicts the geometric features of text regions through a fully convolutional network without going through a complex anchor mechanism, improving the detection speed while ensuring the detection accuracy. In 2021, He et al. proposed the Multi-Oriented Scene Text Detector (MOST), which first generates rough detection results and then dynamically adjusts the receptive field through deformable convolution to generate more refined results. Although regression-based text detection methods have good performance in multi-directional text detection, due to the limitation of the shape of the rectangular box, it is difficult for these methods to achieve accurate results when detecting curved text.

[0005] Contour fitting-based text detection methods are similar to regression-based and segmentation-based text detection methods in terms of the model. The unique feature of these methods is that they directly fit the text contour in a novel way. In 2018, Long et al. proposed TextSnake (a flexible representation for detecting text of arbitrary shapes), which innovatively covers the text region with a series of overlapping disks, and forms the text boundary curve according to the tangents of all disks, converting the problem of detecting text regions into predicting the center coordinates, radius, and direction of each disk. The Fourier Contour Embedding for Arbitrary-Shaped Text Detection (FCENet) proposed in 2021 predicts the Fourier feature vectors of text instances and then reconstructs the text contour in the image spatial domain through inverse Fourier transform. Compared with segmentation-based text detection methods, these methods need to predict more parameters in the test stage, which affects the performance of the model.

[0006] The text detection method based on segmentation transforms the text detection task into a pixel-level classification problem, providing high-precision text localization. The Progressive Scale Expansion Network (PSENet) proposed in 2019 innovatively adopts a scale expansion method. This method generates a text kernel for each text instance and, starting from the pixels of the text kernel, iteratively merges adjacent text pixels to obtain the final text region. However, the post-processing process of this method is complex and the detection efficiency of the model is low. In 2020, the Differentiable Binarization Network (DBNet) proposed by Liao et al. effectively simplifies the post-processing part, greatly improves the detection speed, and also demonstrates excellent performance on lightweight backbone networks. This method mainly introduces a differentiable binarization function, which optimizes the binarization process together with the model during the training phase, and the differentiable binarization part can be removed during the test phase without affecting the detection speed of the model.

[0007] Although the above text detection methods have achieved excellent results, there are inherent difficulties in detecting text instances in natural scene images: large text scale differences and extreme aspect ratios. To address these difficulties, existing methods usually directly adopt a Feature Pyramid Network for feature fusion. However, when directly fusing multi-scale features through upsampling and pixel-level addition, small-scale text and large-scale text share feature maps with the same resolution, ignoring the uniqueness of features under different receptive fields, which may lead to missed detection of small-scale text. In addition, when directly fusing multi-level features from top to bottom, the feature differences between different levels are ignored, resulting in inaccurate text edge segmentation, and causing the same text instance with an extreme aspect ratio to be detected as multiple text regions, affecting the detection performance of the model.

[0008] To alleviate these problems, the present invention proposes a text detection method and system based on cross-level feature enhancement and auxiliary feature guidance. The text detection method proposed by the present invention is used to detect arbitrary-shaped scene text, solving the problem of difficult detection of text with large scale differences and extreme aspect ratios in natural scene images. This method adopts an unequal processing method for different-level features, dynamically adjusts the receptive field of low-level features to obtain multi-scale information, and introduces differential convolution to enhance the edge high-frequency information of high-level features, so as to more accurately locate text of different scales. In addition, this method also uses high-level features to generate auxiliary features to help the network distinguish different text instances, avoiding the problem of a text instance with an extreme aspect ratio being detected as multiple text instances. This method provides a research solution for deep learning methods for natural scene text detection and promotes the development of visual tasks such as image understanding to a certain extent. Summary of the Invention

[0009] The object of the present invention is to provide a text detection method and system based on cross-level feature enhancement and auxiliary feature guidance, so as to solve the problem in the above-mentioned background technology that it is difficult to detect texts with large scale differences and extreme aspect ratios in natural scene images.

[0010] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0011] In the first aspect, the present invention proposes a text detection method based on cross-level feature enhancement and auxiliary feature guidance, including the following steps:

[0012] S1. Extract feature maps of different levels from the image to be detected;

[0013] S2. For low-level features, dynamically adjust the receptive field by adjusting the size of the convolutional kernel to obtain multi-scale information, and generate enhanced features of the low-level features;

[0014] S3. For high-level features, use differential convolution to enhance the high-frequency information of text edges, and generate enhanced features of the high-level features;

[0015] S4. Adopt the method of cross-level feature enhancement to fuse features from different levels. The enhanced features of low-level features and the enhanced features of high-level features are fused one by one, and gradually fused until all features of different levels are fused to generate fused features;

[0016] S5. Extract the semantic information of high-level features to generate auxiliary features;

[0017] S6. For the generated auxiliary features and fused features, use the attention guidance method to determine the attribution of text kernel pixels to obtain the features after attention guidance;

[0018] S7. Based on the features after attention guidance, perform text-non-text prediction on pixels to generate text bounding boxes.

[0019] This text detection method is used to detect texts in scenes of arbitrary shapes, and solves the problem that it is difficult to detect texts with large scale differences and extreme aspect ratios in natural scene images. This method adopts an unequal processing method for features of different levels, dynamically adjusts the receptive field for low-level features to obtain multi-scale information, and introduces differential convolution for high-level features to enhance edge high-frequency information, so as to more accurately locate texts of different scales. In addition, this method also uses high-level features to generate auxiliary features to help the network distinguish different text instances, and avoid the problem that an extremely long and narrow text is detected as multiple text instances.

[0020] Preferably, the specific content of S1 is as follows:

[0021] The image to be detected passes through the backbone network to obtain feature maps at four different levels from stage 2 to stage 5; the number of channels is adjusted to 256 through 1×1 convolutional layers, and the multi-level feature representations after adjustment are denoted as {X 2 ,X 3 ,X 4 ,X 5}.

[0022] Preferably, S2 is specifically as follows:

[0023] The low-level features are first dynamically adjusted in receptive field through convolutional operations with different convolutional kernel sizes to obtain multi-scale information. Then, the feature maps with different receptive field sizes are concatenated along the channel dimension, and the number of channels is adjusted through 1×1 convolution to obtain multi-scale features. Subsequently, the channel importance of different receptive fields is determined by learning attention scores. Finally, the weights of the corresponding channels are multiplied by the multi-scale features to obtain the enhanced features of the low-level features; the process of generating the enhanced features is expressed as follows:

[0024] X' L =Conv 1×1 ([Conv 1×1 (X i ),Conv 3×3 (X i ),Conv 5×5 (X i )])

[0025] X L =σ(f FC (ReLU(f FC (f GAP (X' L )))))×X' L

[0026] Among them, X i represents the low-level features, X L represents the enhanced features of the low-level features, [·] represents concatenating features along the channel dimension, Conv k×k (·) represents a convolutional operation with a convolutional kernel size of k×k, f GAP (·) represents global average pooling, f FC (·) represents a fully connected layer, ReLU(·) represents the ReLU activation function, and σ(·) represents the Sigmoid activation function.

[0027] Preferably, S3 is specifically as follows:

[0028] Enhance the high-level features with differential convolution-based information. First, perform sub-pixel convolution to expand the resolution of the input features, and then pass through three convolutional layers to enhance the features. Among them, local features are aggregated through ordinary convolution with a kernel size of 3×3, and central differential convolution and corner differential convolution are used to enhance edge high-frequency information. Then, perform pixel-wise addition on the enhanced features of the three branches and generate the enhanced features of the high-level features through the ReLU activation function. The function representation of the process of generating the enhanced features is as follows:

[0029] X' H = f sub×N (Conv 3×3 (X j ))

[0030] X H = ReLU(Conv 3×3 (X' H ) + Conv CD (X' H ) + Conv AD (X' H )) + X' H

[0031] Among them, X j represents the high-level features, X H represents the enhanced features of the high-level features, Conv 3×3 (·) represents the convolution operation with a kernel size of 3×3, f sub×N (·) represents the sub-pixel convolution operation with an upsampling rate of N, Conv CD (·) represents the central differential convolution, Conv AD (·) represents the corner differential convolution, and ReLU(·) represents the ReLU activation function.

[0032] Preferably, the S4 is specifically as follows:

[0033] Obtain several groups of preliminary cross-level features by combining features of different levels in a one-to-one correspondence between low-level features and high-level features. First, generate respective enhanced features for each group of preliminary cross-level features for fusion. After fusion, the features are combined again in a one-to-one correspondence to obtain several groups of two-level cross-level features. Then, generate respective enhanced features for each group of two-level cross-level features for fusion until all features are fused;

[0034] First, perform the first-step feature enhancement and fusion on the cross-level features X 2 and X 4 , X 3 and X 5 respectively to generate X 2 ' and X 3 '. Subsequently, the different-level features obtained in the first step, X2 ′ and X 3 ′ perform the second - step feature enhancement fusion, and fuse the extracted multi - level features in a step - by - step manner to generate the fused feature X C ; among them, the feature enhancement fusion is to add the enhanced features of the low - level features and the enhanced features of the high - level features to obtain the fused feature after fusion, which is expressed as follows:

[0035] X′ i = X H +X L

[0036] Among them, X L represents the enhanced feature of the low - level feature, X H represents the enhanced feature of the high - level feature, and X i ′ represents the fused feature after the enhanced feature fusion.

[0037] Preferably, the S5 is specifically as follows:

[0038] Enrich the semantic information of the feature X 5 first through an information extraction block, which consists of a 1×1 convolution, a Relu activation function, and a batch normalization operation, then magnify the resolution through a sub - pixel convolution, and then perform a pixel - level subtraction operation with the feature X 4 to remove redundant information, obtaining the feature The function is expressed as follows:

[0039] X 5 ' = f BN (Conv 1×1 (ReLU(f BN (Conv 1×1 (X 5 )))))+X 5

[0040]

[0041] Among them, Conv 1×1 (·) represents a 1×1 convolution operation, f BN (·) represents batch normalization, ReLU(·) represents the ReLU activation function, and f sub×2 (·) represents a sub - pixel convolution operation with an up - sampling rate of 2;

[0042] Further perform feature enhancement on the feature after noise suppression ; first, retain the texture information through parallel global average pooling and global max pooling, then perform a pixel - level addition operation on the two pooled features and then generate weights to obtain the attention map X W , and finally use the attention map and the feature after noise suppression Perform a multiplication operation and a residual connection to obtain the auxiliary feature X A ; It is expressed as follows:

[0043]

[0044] Among them, f GAP (·) represents global average pooling, f GMP (·) represents global max pooling, f FC (·) represents a fully connected layer, ReLU(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, represents the deformation operation to reshape the vector into an attention map.

[0045] Preferably, the S6 is specifically as follows:

[0046] For the auxiliary feature and the fusion feature, first perform upsampling by bilinear interpolation to expand the resolution of the auxiliary feature by 4 times, then concatenate it with the fusion feature along the channel dimension, and then generate the attention weight X through a convolution with a kernel size of 3×3 and an activation function S , and finally multiply the attention weight by the fusion feature and perform a residual connection to obtain the guided feature X G ; It is expressed as follows:

[0047] X Cat =[X C , f up×4 (X A )]

[0048] X S =σ(Conv 3×3 (ReLU(Conv 3×3 (X Cat ))))

[0049] X G =X C ×X S +X C

[0050] Among them, f up×4 (·) represents the bilinear interpolation upsampling operation with a scaling factor of 4, [·] represents concatenating features along the channel dimension, Conv 3×3 (·) represents a 3×3 convolutional layer, ReLU(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function.

[0051] Preferably, the S7 is specifically as follows:

[0052] Input the feature X G into two parallel branches for processing, and respectively obtain a single-channel threshold map X Tand the single-channel text probability map X P , the parallel branches include an upsampling operation and text-non-text prediction for pixels; subsequently, a binary map is obtained through a post-processing process, and a text bounding box is generated therefrom. The prediction process and the function for generating the binary map are expressed as:

[0053] X O = σ(f deconv (f deconv (Conv 3×3 (X G ))))

[0054]

[0055] where Conv 3×3 (·) represents a 3×3 convolutional layer, f deconv (·) represents upsampling through a deconvolution operation to double the feature size, σ(·) represents the Sigmoid activation function, X B represents the generated binary map, and (i, j) represents the pixel coordinates in the image.

[0056] In the second aspect, the present invention proposes a text detection system based on cross-level feature enhancement and auxiliary feature guidance, including:

[0057] A backbone network, which is constructed by replacing the 3×3 convolutional layers in stages 3, 4, and 5 of ResNet-50 with deformable convolutions, and is used to obtain four different hierarchical feature maps {X 2 , X 3 , X 4 , X 5} of stages 2 - 5, and inputs the output features of each stage into the cross-level feature enhancement module;

[0058] A cross-level feature enhancement module, which is used to dynamically adjust the receptive field of low-level features to obtain multi-scale information, introduce differential convolutions to enhance edge high-frequency information for high-level features, and add the enhanced low-level features and high-level features to output the cross-level enhanced fusion feature X C ; adopt a cross-level feature enhancement method to fuse features from different levels, first send the cross-level features into the cross-level feature enhancement module, and then send the output two-layer features into the cross-level feature enhancement module to generate the fusion feature X C ;

[0059] An auxiliary weight generation module, which is used to enhance the semantic information in high-level features and generate an auxiliary feature X A ; use the information extraction block of the auxiliary weight generation module to enrich the semantic information of the high-level feature X 5 , and then perform a pixel-level subtraction operation with the feature X 4 to remove useless noise in the background to obtain a feature The texture information is retained through global average pooling and global max pooling operations, and then an attention map X is obtained through pixel-level addition and weight generation W . Finally, the attention map and the feature after suppressing noise are used for multiplication operation and residual connection to obtain the auxiliary feature X A ;

[0060] The attention guidance module is used to correctly segment the text region under attention guidance to obtain the guided feature X G ;

[0061] The output module uses an upsampling operation and makes text-non-text predictions on pixels to generate text bounding boxes.

[0062] Compared with the prior art, the beneficial effects of the present invention are:

[0063] The text detection method based on cross-level feature enhancement and auxiliary feature guidance proposed in the present invention aims to solve the problem of difficult detection of texts with large scale differences and extreme aspect ratios in natural scene images. The present invention mainly includes a cross-level feature enhancement module CFEM, an auxiliary weight generation module AWGM, and the attention guidance part of the latter to the former. Among them, the cross-level feature enhancement module CFEM enhances features by unequal processing of features at different levels to obtain a fused feature containing more effective information; and the semantic information in the high-level features is enhanced through the auxiliary weight generation module AWGM, enabling the network to correctly segment the text region. A large number of experiments conducted on four text detection benchmark datasets prove the effectiveness of the method proposed in the present invention. The method in the present invention provides a research solution for the deep learning method of natural scene text detection, and promotes the development of visual tasks such as image understanding to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is the overall network structure diagram of the text detection system based on cross-level feature enhancement and auxiliary feature guidance in the present invention;

[0065] Figure 2 is the specific structure diagram of the cross-level feature enhancement module in the present invention;

[0066] Figure 3 is the specific structure diagram of the auxiliary weight generation module in the present invention;

[0067] Figure 4 is the specific structure diagram of the attention guidance module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0069] Embodiment 1:

[0070] The text detection method based on cross-level feature enhancement and auxiliary feature guidance in the present invention proposes a text detection network based on cross-level feature enhancement and auxiliary feature guidance. The overall network structure of the text detection network is as Figure 1 shown. The text detection network consists of a feature extraction network, a cross-level feature enhancement module (Cross-level Feature Enhancement Module, CFEM), an auxiliary weight generation module (Auxiliary Weight Generation Module, AWGM), an attention guidance and output module. The specific text detection method in the present invention is as follows:

[0071] Step 1: Construct a backbone network for feature extraction.

[0072] Use ResNet-50 as the backbone network for feature extraction. To enhance the feature extraction ability of the backbone network, in the present invention, the 3×3 convolutional layers in the 3rd, 4th, and 5th stages of ResNet-50 are replaced with deformable convolutions. The input image passes through the backbone network to obtain feature maps of four different levels from stage 2 to stage 5.

[0073] First, the number of channels is adjusted to 256 through 1×1 convolutional layers respectively. The adjusted multi-level feature representations are denoted as {X 2 , X 3 , X 4 , X 5}, and the resolution sizes are {1 / 4, 1 / 8, 1 / 16, 1 / 32} of the input size respectively.

[0074] Step 2: Construct a cross-level feature enhancement module to enhance the features by an unequal processing method for different-level features, and obtain fused features containing more effective information.

[0075] The text scale in natural scene images varies greatly. However, simply using a series of upsampling operations and pixel addition to obtain the fused features will neglect the differences in the features extracted by different-depth convolutional networks, thus affecting the model's recognition ability for texts of different scales and reducing the detection accuracy of the model. Therefore, in the present invention, different processing methods are adopted for features of different levels through the Cross-level Feature Enhancement Module (CFEM). The specific structure of the cross-level feature enhancement module CFEM is as shown in Figure 2 shown. In the present invention, the feature maps of four levels are first sent into the cross-level feature enhancement module CFEM respectively, and then the enhanced features are passed through the cross-level feature enhancement module CFEM to obtain a more informative fused feature X C .

[0076] (1) Dynamically adjust the receptive field of low-level features to obtain multi-scale information;

[0077] In the present invention, the receptive field of the input low-level features is dynamically adjusted by adjusting the size of the convolutional kernel to obtain multi-scale information. After that, the feature maps with different receptive field sizes are concatenated together along the channel dimension and the number of channels is adjusted through 1×1 convolution to ensure the correspondence of the low-level features and the high-level features in the channel dimension. After obtaining the multi-scale features, the importance of the channels of different receptive fields is determined by learning the attention scores. For the multi-scale features, first, the spatial information is integrated through global average pooling, then the feature vectors are predicted through two fully connected layers and the ReLU activation function, and finally, the weights of the corresponding channels are obtained through the Sigmoid activation function and multiplied by the multi-scale features to obtain the enhanced features of the low-level features. The specific implementation process is as follows:

[0078] X' L =Conv 1×1 ([Conv 1×1 (X i ),Conv 3×3 (X i ),Conv 5×5 (X i )])

[0079] X L =σ(f FC (ReLU(f FC (f GAP (X' L )))))×X' L

[0080] where [·] represents concatenating features along the channel dimension, Conv k×k (·) represents a convolution operation with a convolutional kernel size of k×k, f GAP (·) represents global average pooling, fFC (·) represents a fully connected layer, ReLU(·) represents the ReLU activation function, and σ(·) represents the Sigmoid activation function.

[0081] (2) Introduce differential convolution to enhance edge high-frequency information for high-level features;

[0082] In natural scene images, the edge high-frequency information of text is of great significance for distinguishing text regions from background regions. As the feature extraction network deepens, the edge information becomes smoother, and the high-frequency information in high-level features gradually disappears, which seriously affects the model's ability to accurately locate text boundaries. Especially when detecting small-scale text, the loss of high-frequency information is more serious. Therefore, the present invention introduces differential convolution to enhance information for the input high-level features. For the input high-level features, first, sub-pixel convolution is used to expand the resolution of the input features. The sub-pixel convolution includes a 3×3 convolutional layer and channel shuffling. Subsequently, information enhancement based on differential convolution is performed on the generated high-resolution features.

[0083] Specifically, first, three convolutional layers are used to enhance features. Among them, local features are aggregated through ordinary convolution with a convolution kernel size of 3×3, and then edge high-frequency information is enhanced through central differential convolution and corner differential convolution respectively. Then, the enhanced features of the three branches are added pixel by pixel and passed through the ReLU activation function to generate the enhanced features of high-level features. Finally, the features obtained by enhancing the low-level features and high-level features respectively are added to output cross-level enhanced features, and the ability of the model to detect multi-scale text is enhanced by processing the cross-level features in an unequal manner. The specific implementation process can be expressed as follows:

[0084] X' H = f sub×N (Conv 3×3 (X j ))

[0085] X H = ReLU(Conv 3×3 (X' H ) + Conv CD (X' H ) + Conv AD (X' H )) + X' H

[0086] X′ i = X H + X L

[0087] Among them, Conv 3×3 (·) represents a convolution operation with a convolution kernel size of 3×3, and f sub×N (·) represents a sub-pixel convolution operation with an upsampling rate of N, such asFigure 1 As shown, in the first-step feature enhancement, N is set to 4, and in the second-step feature enhancement, N is set to 2. Conv CD (·) represents central difference convolution, and Conv AD (·) represents angular difference convolution, and ReLU(·) represents the ReLU activation function.

[0088] Step 3: Construct an auxiliary weight generation module to generate auxiliary features.

[0089] Convolutional neural networks can extract rich feature information from natural scene images. However, a large amount of redundant noise is contained in the low-level feature information, which seriously interferes with the accurate positioning of text instances by the model, causing the model to wrongly aggregate background pixels into text kernels or wrongly segment text regions. The semantic information contained in high-level features usually involves the entire image or text instance, which is convenient for understanding the context information of the image and helps the model correctly segment text regions. Therefore, to address the above problems, the present invention generates auxiliary features through an auxiliary weight generation module AWGM (Auxiliary Weight Generation Module, AWGM), and then accurately judges text instances through attention guidance with the fused features.

[0090] The specific structure of the auxiliary weight generation module AWGM is as Figure 3 shown. The input feature of this module is the high-level feature X extracted by the backbone network 4 and X 5 . The input feature X 5 first enriches the semantic information therein through an information extraction block, then doubles the resolution of the feature map through sub-pixel convolution, and then performs a pixel-level subtraction operation with the feature X 4 to remove the useless noise in the background. The information extraction block is mainly composed of 1×1 convolution, Relu activation function, and batch normalization operations to generate the feature The process is as follows:

[0091] X 5 ' = f BN (Conv 1×1 (ReLU(f BN (Conv 1×1 (X 5 )))) + X 5

[0092]

[0093] Among them, Conv 1×1 (·) represents a 1×1 convolution operation, f BN (·) represents batch normalization, ReLU(·) represents the ReLU activation function, f sub×2(·) represents a sub-pixel convolution operation with an upsampling rate of 2.

[0094] Subsequently, in order to generate auxiliary features more accurately for attention guidance of the fusion features, the present invention further enhances the features after suppressing noise. First, texture information is retained through parallel global average pooling and global max pooling. Subsequently, pixel-wise addition operation is performed on the two pooled features, and then weight generation is carried out to obtain the attention map X W , and finally, the attention map and the features after suppressing noise are used for multiplication operation and residual connection to obtain auxiliary features, which helps the model effectively distinguish text instances. The process of generating the auxiliary feature X A is as follows:

[0095]

[0096] where f GAP (·) represents global average pooling, f GMP (·) represents global max pooling, f FC (·) represents a fully connected layer, ReLU(·) represents the ReLU activation function, and σ(·) represents the Sigmoid activation function. represents a deformation operation to reshape the vector into the shape of the attention map for subsequent operations.

[0097] Step Four: Conduct attention guidance to correctly segment the text region.

[0098] For the generated auxiliary features and fusion features, the present invention helps the model determine the attribution of text kernel pixels through attention guidance. The process of attention guidance is as Figure 4 shown. The implementation process of attention guidance is as follows:

[0099] X Cat = [X C , f up×4 (X A )]

[0100] X S = σ(Conv 3×3 (ReLU(Conv 3×3 (X Cat ))))

[0101] X G = X C × X S + X C

[0102] where f up×4 (·) represents a bilinear interpolation upsampling operation with a scaling factor of 4, [·] represents concatenating features along the channel dimension, and Conv3×3 (·) represents a 3×3 convolutional layer, ReLU(·) represents the ReLU activation function, and σ(·) represents the Sigmoid activation function.

[0103] Step Five: Generate text bounding boxes.

[0104] Input X G into two parallel output modules for post - processing. Through prediction, a single - channel threshold map X T and a single - channel text probability map X P are obtained respectively. The output modules mainly include up - sampling operations and text - non - text predictions on pixels. Subsequently, through a post - processing process, a binary map is obtained, and text bounding boxes are generated therefrom. The implementation process is shown by the following formula:

[0105] X O = σ(f deconv (f deconv (Conv 3×3 (X G ))))

[0106]

[0107] where Conv 3×3 (·) represents a 3×3 convolutional layer, f deconv (·) represents up - sampling through de - convolution operations to double the feature size, σ(·) represents the Sigmoid activation function, X B represents the generated binary map, and (i, j) represents the pixel coordinates in the image.

[0108] Experimental verification:

[0109] (1) Dataset:

[0110] To verify the effectiveness of the method proposed in the present invention, experiments are carried out on four text detection benchmark datasets. The datasets used specifically include:

[0111] MSRA - TD500 contains a total of 500 natural images, including 300 training images and 200 test images. This dataset includes two languages, Chinese and English. Considering the small number of training images in the dataset, 400 additional images from the HUST - TR400 dataset are added during training.

[0112] The ICDAR2015 dataset is taken in natural street scenes, and the text in the images has phenomena such as blurring and deformation. The training set in this dataset contains 1000 images and the test set contains 500 images.

[0113] Total-Text is a classic arbitrary-shaped text detection dataset, which contains a total of 1,555 images, of which 1,255 are used for training and the remaining 300 are used for testing.

[0114] The CTW1500 dataset contains horizontal, tilted, and curved texts. This dataset contains 1,000 training images and 500 test images, covering multiple languages mainly in Chinese and English.

[0115] SynthText is a large synthetic dataset, containing more than 850,000 synthetic text images. The method proposed in the present invention is only pre-trained on this dataset.

[0116] (2) Experimental settings:

[0117] The method proposed in the present invention is implemented based on the Python language, and the experimental environment is shown in Table 1. The entire network is trained using the SGD optimizer. When initializing the parameters, the backbone network loads the weights pre-trained on ImageNet, and other network layers are initialized using the Kaiming initialization method. The pre-training stage of the model is carried out on the synthetic dataset SynthText, with the learning rate fixed at 0.007, the batch size set to 8, and a total of 2 epochs are trained. The fine-tuning stage is carried out on four benchmark datasets respectively, with the initial learning rate set to 0.007, and a total of 1200 epochs are trained. Before training, data augmentation is performed on the images, randomly cropping the training images to a size of 640×640, and using methods such as horizontal flipping and random rotation to enhance the generalization ability of the model.

[0118] Table 1 Experimental environment

[0119]

[0120]

[0121] (3) Experimental results:

[0122] The text detection method based on cross-level feature enhancement and auxiliary feature guidance proposed in the present invention is evaluated on four common text detection benchmark datasets and compared with existing methods. The method proposed in the present invention is evaluated from four indicators: precision, recall, mean average precision (Hmean), and frames per second (FPS). In addition, ablation experiments are carried out on the MSRA-TD500 dataset to verify the effectiveness of each part of the present invention.

[0123] The present invention conducts a performance comparison with other scene text detection methods on the benchmark dataset, and the evaluation results are shown in Tables 2 and 3. The best performance of each index in the table is marked in bold.

[0124] Table 2 Experimental results on the benchmark datasets MSRA-TD500 and ICDAR 2015

[0125]

[0126] Table 3 Experimental results on the benchmark datasets Total-Text and CTW1500

[0127]

[0128]

[0129] As can be seen from Table 2, the method proposed in the present invention achieves the highest accuracy on both the MSRA-TD500 and ICDAR2015 datasets. Especially on the MSRA-TD500 dataset, the three indexes of the present invention all reach the optimum, fully demonstrating the excellent performance of the method of the present invention in detecting multi-directional scene text. Table 3 shows the experimental results of each text detection method on the curved text dataset. As can be seen from the table, the method proposed in the present invention achieves the highest recall rate on both of these natural scene text datasets, and can well balance the detection accuracy and speed.

[0130] The excellent performance on the four basic datasets benefits from the method proposed in the present invention, which adopts an unequal processing method for high-level features and low-level features, and performs feature fusion in a cross-level manner, effectively solving the problem of large text scale differences and improving the detection effect. In addition, by generating auxiliary features to guide the fusion features, the method helps the model distinguish different text instances, avoiding the model mis-segmenting a text instance into multiple text regions or detecting multiple text instances as one text, and improving the detection performance of the model.

[0131] To further prove the effectiveness of the method of the present invention, ablation experiments on the functions of each component in the method of the present invention are carried out on the MSRA-TD500 dataset in Tables 4, 5 and 6.

[0132] Table 4 Ablation experiment results of the main module

[0133] CFEM AWGM Precision Recall Hmean / / 88.97 79.04 83.71 √ / 91.46 82.82 86.93 / √ 90.75 82.65 86.51 √ √ 92.52 85.05 88.63

[0134] Table 5 Influence of the two-part feature processing in AWGM on the experimental results

[0135]

[0136] Table 6 Influence of Different Hierarchical Feature Processing Methods in CFEM on Experimental Results

[0137] Low-level features High-level features Precision Recall Hmean / / 88.97 79.04 83.71 LTE / 92.28 80.07 85.74 LTE LTE 93.27 80.93 86.66 / HTE 91.90 79.90 85.48 HTE HTE 90.38 82.30 86.15 LTE HTE 91.46 82.82 86.93

[0138] Table 4 shows the ablation experiment on the main module of the method of the present invention. It can be seen from the table that, compared with the baseline network, the average precision of the model with the cross-level feature enhancement module (CFEM) added has increased by 3.21%. This shows that the CFEM proposed by the present invention can effectively dynamically adjust the receptive field and improve the model's ability to locate texts of different scales. When only the auxiliary weight generation module (AWGM) is added to the baseline model, the average precision of the model has increased by 2.8%. This shows that generating auxiliary features through high-level features can effectively guide the fusion features for text detection and more accurately assign pixels to different text regions. When both CFEM and AWGM are added to the baseline model, the precision, recall rate, and average precision of the model have increased by 3.55%, 6.01%, and 4.92% respectively. The experimental results fully prove the effectiveness of the cross-level feature enhancement module (CFEM) and the auxiliary weight generation module (AWGM) proposed by the present invention.

[0139] Table 5 illustrates the importance of the first-step information extraction and further feature enhancement in AWGM for differentiating text instances. To prove the effectiveness of the unequal processing method for different hierarchical features in the CFEM proposed by the present invention, different reinforcement methods are used for different hierarchical features. The experimental results are shown in Table 6. In Table 6, LTE represents the low-level feature enhancement method, that is, Figure 2 the processing method for the input low-level features in Figure 2 and HTE represents the high-level feature enhancement method, that is,

[0140] the feature enhancement method for the input high-level features in The experimental results fully illustrate the superior performance of each part of the CFEM proposed by the present invention.

[0140] As described above, it is only used to help understand the method of the present invention and its core essence. However, the protection scope of the present invention is not limited thereto. For those of ordinary skill in the art in the technical field of the present invention, any equivalent replacement or change made within the technical scope disclosed by the present invention according to the technical solution and inventive concept of the present invention should be covered within the protection scope of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A text detection method based on cross-level feature enhancement and auxiliary feature guidance, characterized in that: The following steps are involved: S1, extracting feature maps of different levels from the image to be detected; S2. For low-level features, the receptive field is dynamically adjusted by adjusting the size of the convolution kernel to obtain multi-scale information and generate enhanced features of low-level features; S3. For high-level features, differential convolution is used to enhance the high-frequency information of text edges and generate enhanced features of high-level features; S4. Use cross-level feature enhancement to fuse features from different levels. The enhanced features of low-level features and high-level features are fused one by one, and gradually fused until all features of different levels are fused to generate fused features. S5, extracting semantic information of high-level features and generating auxiliary features; S6. For the generated auxiliary features and fusion features, the attribution of the text core pixels is determined by using the attention guidance method to obtain the features after attention guidance; S7. Based on the attention-guided features, perform text-non-text prediction on pixels and generate text bounding boxes.

2. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 1 is characterized in that: The S1 is specifically as follows: The image to be detected passes through the backbone network to obtain feature maps of four different levels from stage 2 to stage 5; the number of channels is adjusted to 256 through 1×1 convolutional layers respectively, and the adjusted multi-level features are represented as {X2,X3,X4,X5}.

3. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 2 is characterized in that: The S2 is specifically as follows: The low-level features are first dynamically adjusted through convolution operations with different convolution kernel sizes to obtain multi-scale information. Then, the feature maps with different receptive field sizes are spliced ​​along the channel dimension, and the number of channels is adjusted through 1×1 convolution to obtain multi-scale features. Then, the importance of channels with different receptive fields is determined by learning attention scores. Finally, the weights of the corresponding channels are multiplied by the multi-scale features to obtain the enhanced features of the low-level features. The process of generating enhanced features is expressed as follows: X' L =Conv 1×1 ([Conv 1×1 (X i ),Conv 3×3 (X i ),Conv 5×5 (X i )]) X L =σ(f FC (ReLU(f FC (f GAP (X' L )))))×X' L Among them, X i represents low-level features, X L represents the enhanced features of low-level features, [·] represents the concatenation of features along the channel dimension, Conv k×k (·) represents the convolution operation with a kernel size of k×k, f GAP (·) represents global average pooling, f FC (·) represents a fully connected layer, ReLU(·) represents the ReLU activation function, and σ(·) represents the Sigmoid activation function.

4. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 3 is characterized in that: The S3 is as follows: The high-level features are enhanced based on differential convolution. The resolution of the input features is first expanded by sub-pixel convolution, and then three convolution layers are used to enhance the features. The local features are aggregated by ordinary convolution with a convolution kernel size of 3×3, and the edge high-frequency information is enhanced by center differential convolution and angle differential convolution. Then, the enhanced features of the three branches are added pixel by pixel and the enhanced features of the high-level features are generated by the ReLU activation function. The function of the process of generating enhanced features is expressed as follows: X' H =f sub×N (Conv 3×3 (X j )) X H =ReLU(Conv 3×3 (X' H )+Conv CD (X' H )+Conv AD (X' H ))+X' H Among them, X j represents high-level features, X H Represents the enhanced features of high-level features, Conv 3×3 (·) represents the convolution operation with a kernel size of 3×3, f sub×N (·) represents the sub-pixel convolution operation with upsampling rate N, Conv CD (·) represents the central difference convolution, Conv AD (·) represents angular differential convolution, and ReLU(·) represents the ReLU activation function.

5. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 4 is characterized in that: The S4 is specifically as follows: First, the cross-level features X2 and X4, X3 and X5 are fused by the first step feature enhancement to generate X′2 and X′3 respectively. Then, the different-level features X′2 and X′3 obtained in the first step are fused by the second step feature enhancement. The extracted multi-level features are fused by step-by-step fusion to generate the fused feature X C ; Among them, feature enhancement fusion is to add the enhanced features of low-level features and the enhanced features of high-level features to obtain the fused features after fusion, which is expressed as follows: X i =X H +X L Among them, X L represents the enhanced features of low-level features, X H Represents the enhanced features of high-level features, X′ i Represents the fused features after enhanced feature fusion.

6. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 5 is characterized in that: The S5 is specifically as follows: Feature X5 is first enriched with semantic information through the information extraction block, which consists of 1×1 convolution, Relu activation function and batch normalization operation, and then the resolution is enlarged through sub-pixel convolution. After that, it is subtracted from feature X4 at the pixel level to remove redundant information and obtain feature X5. Its function is expressed as follows: <h2 style=";text-align:left;direction:ltr">X5'=f<h2 style=";text-align:left;direction:ltr"> BN <h2 style=";text-align:left;direction:ltr"> (Conv<h2 style=";text-align:left;direction:ltr"> 1×1 <h2 style=";text-align:left;direction:ltr"> (ReLU(f<h2 style=";text-align:left;direction:ltr"> BN <h2 style=";text-align:left;direction:ltr"> (Conv<h2 style=";text-align:left;direction:ltr"> 1×1 <h2 style=";text-align:left;direction:ltr"> (X5)))))+X5 Among them, Conv 1×1 (·) represents a 1×1 convolution operation, f BN (·) represents batch normalization, ReLU(·) represents the ReLU activation function, and f sub×2 (·) indicates a sub-pixel convolution operation with an upsampling rate of 2; The characteristics after noise suppression Further feature enhancement is performed; first, the texture information is retained by parallel global average pooling and global maximum pooling, and then the two pooling features are added at the pixel level and weighted to obtain the attention map X W , and finally use the attention map and the features after noise suppression Perform multiplication and residual connection to obtain auxiliary feature X A ; means as follows: Among them, f GAP (·) represents global average pooling, f GMP (·) represents the global maximum pooling, f FC (·) represents the fully connected layer, ReLU(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, The representative deformation operation reshapes the vector into an attention map.

7. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 6 is characterized in that: The S6 is specifically as follows: For auxiliary features and fusion features, we first upsample the auxiliary features by 4 times through bilinear interpolation, then concatenate them with the fusion features along the channel dimension, and then generate the attention weight X through convolution with a kernel size of 3×3 and activation function. S Finally, the attention weight is multiplied by the fusion feature and a residual connection is performed to obtain the guided feature X G ; means as follows: X Cat =[X C ,f up×4 (X A )] X S =σ(Conv 3×3 (ReLU(Conv 3×3 (X Cat )))) X G =X C ×X S +X C Among them, f up×4 (·) represents the bilinear interpolation upsampling operation with a scaling factor of 4, [·] represents the concatenation of features along the channel dimension, Conv 3×3 (·) represents a 3×3 convolutional layer, ReLU(·) represents the ReLU activation function, and σ(·) represents the Sigmoid activation function.

8. The text detection method based on cross-level feature enhancement and auxiliary feature guidance according to claim 7, characterized in that: The S7 is specifically as follows: The feature X G After being input into two parallel branches for processing, the single channel threshold map X is obtained. T And the single-channel text probability map X P ,,The parallel branch includes upsampling operation and text-non-text prediction of pixels; then the binary image is obtained through post-processing, and the text bounding box is generated from it. The function of the prediction process and the generation of the binary image is expressed as: X O =σ(f deconv (f deconv (Conv 3×3 (X G )))) Among them, Conv 3×3 (·) represents a 3×3 convolutional layer, f deconv (·) indicates that the feature size is doubled by upsampling through deconvolution operation, σ(·) indicates Sigmoid activation function, X B Represents the generated binary image, and (i, j) represents the pixel coordinates in the image.

9. A text detection system based on cross-level feature enhancement and auxiliary feature guidance applied to the method according to any one of claims 1 to 8, characterized in that: include: The backbone network is constructed by replacing the 3×3 convolutional layers in stages 3, 4, and 5 of ResNet-50 with deformable convolutions to obtain feature maps {X2, X3, X4, X5} at four different levels from stage 2 to stage 5; The cross-level feature enhancement module is used to dynamically adjust the receptive field of low-level features to obtain multi-scale information, introduce differential convolution to high-level features to enhance edge high-frequency information, and add the enhanced low-level features to the high-level features to output the cross-level enhanced fusion feature X. C ; Auxiliary weight generation module, used to enhance the semantic information in high-level features and generate auxiliary features X A ; The attention guidance module is used to guide the correct segmentation of the text area and obtain the guided feature X G ; The output module performs upsampling and text-non-text prediction on pixels to generate text bounding boxes.

Citation Information

Patent Citations

  • Bill image text detection and recognition method

    CN110033000A

  • Cross-layer multi-model feature fusion and convolutional decoding-based image description method

    CN111859005A

  • Method and device for rapidly and accurately detecting camouflage object from thick to thin

    CN116363467A

  • Text detection method for guiding attention based on feature correction and difference

    CN117809294A

  • Non-reference frame image target detection method based on feature enhancement and multiple scales

    CN117893868A