Natural scene text detection method based on feature fusion and attention enhancement

By using ResNet50 feature extraction and feature fusion modules in natural scene text detection, combined with the three-branch query generation strategy of the attention enhancement module, the challenges of text detection in natural scenes are addressed, and the detection performance of irregular and dense small texts is improved.

CN121033869APending Publication Date: 2025-11-28DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511423595.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

The arbitrary distribution of text instances in natural scenes and the interference of lighting and background make it difficult for existing technologies to accurately detect and recognize text information, especially for irregular and dense small text.

Method used

Feature extraction is performed using a ResNet50-based backbone network. The feature fusion module and attention enhancement module are combined to enhance multi-scale feature representation. A three-branch query generation module is introduced into the encoder to model the interaction between channels and spatial dimensions, thereby improving the detection capability of irregular text and dense small text.

Benefits of technology

It improves the accuracy and recall of text detection in natural scenes, enhances the detection performance of irregular text and dense small text, and achieves more efficient text information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033869A_ABST
    Figure CN121033869A_ABST
Patent Text Reader

Abstract

The invention discloses a natural scene text detection method based on feature fusion and attention enhancement. The method comprises the following steps: acquiring a to-be-detected image; constructing a text detection model in a natural scene, wherein the text detection model is used for extracting text information in the image; training the text detection model in the natural scene to obtain a trained text detection model in the natural scene; and inputting a to-be-detected image into the trained text detection model in the natural scene to realize extraction of text information contained in the image. Through an attention enhancement module, the model is more sensitive to text instance information, and detection of irregular text instances is enhanced; and finally, the natural scene text detection performance is improved. Expression results of various indexes such as subjective impression, accuracy, recall rate and the like on a mainstream data set show that the method has important significance in the aspect of natural scene text detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of image processing and deep learning, and relates to a natural scene text detection method based on feature fusion and attention enhancement. BACKGROUND

[0002] Natural scene text detection is an important branch of computer vision, which involves identifying and understanding text information in natural environments, such as street signs, billboards, traffic signs, etc. In map construction, autonomous driving, image retrieval and many other practical applications, this task is of great significance. In recent years, although significant progress has been made in this field, natural scene text detection is still a challenging problem that needs further exploration. The key challenges include that the distribution of natural scene text instances is often arbitrary direction, which makes it difficult to accurately frame the text instances in the rectangular detection box commonly used in object detection tasks; at the same time, affected by background factors such as light in natural scenes, the color and texture of text instances may be similar to the background, bringing difficulties to detection and recognition tasks.

[0003] Currently, natural scene text detection techniques are mainly divided into classical methods and deep learning methods. The classical methods include connected component analysis-based methods and sliding detection window-based methods. The connected component analysis-based methods use digital image processing techniques to preprocess the input image, obtain text candidate regions, and then realize the connection and positioning of characters and texts through different connected component analysis methods. According to the differences in region generation and feature representation, this type of method can be divided into edge-based methods, stroke width transform-based methods, and maximum stable extremal region-based methods. The sliding detection window-based methods scan the image from top to bottom by designing a sliding detection window, and regard the region covered by each window as a text candidate region, and then extract features for the classifier to predict and verify. The deep learning methods include region proposal-based methods, segmentation-based methods, hybrid methods, and end-to-end text recognition methods. The region proposal-based methods draw on the general object detection process (such as Faster R-CNN, SSD), generate text candidate boxes through the region proposal network (RPN), and complete the positioning by combining the classification and regression branches. In view of the slender shape, multi-directional distribution and scale difference of text, the anchor box design, feature fusion and loss function are optimized accordingly. The segmentation-based methods obtain the pixel-level mask of the text region through semantic segmentation or instance segmentation, and generate the text bounding box through post-processing (such as connected component analysis, polygon fitting), which is especially suitable for detecting text of any shape (such as curved lines, irregular arrangement). The hybrid method combines the advantages of region proposal-based methods and segmentation-based methods, such as combining detection and segmentation ideas, obtaining corner points through the regression method of the detection process, sampling and recombining to obtain candidate boxes, and then predicting the score of the rotation position sensitive segmentation map to assist in judging the quality of the candidate box. The end-to-end text recognition method integrates text detection and text recognition into the same framework, shares the underlying features, and optimizes the detection using the recognition loss. SUMMARY

[0004] To solve the above problems, the technical scheme adopted by the present application is a natural scene text detection method based on feature fusion and attention enhancement, characterized by comprising the following steps:

[0005] obtaining an image to be detected;

[0006] constructing a text detection model in a natural scene for extracting text information in the image;

[0007] training the text detection model in the natural scene to obtain a trained text detection model in the natural scene;

[0008] inputting the image to be detected into the trained text detection model in the natural scene to extract the text information contained in the image.

[0009] Further, the text information includes English letters and English punctuation.

[0010] Further, the text detection model in the natural scene includes:

[0011] The backbone network: using ResNet50 for feature extraction;

[0012] The feature fusion module: fuses the four layers of feature maps extracted by the backbone network to enhance the multi-scale feature representation,

[0013] The encoder: for converting image features into feature maps rich in text semantic and location information, strengthening text features such as character edges and text structures, and suppressing useless background noise based on the fused image transmitted by the feature fusion module;

[0014] The decoder: for restoring the abstract features output by the encoder into fine-grained feature maps that are directly used for text positioning and aligning with the original image size.

[0015] Further, the feature fusion module: fuses the four layers of feature maps extracted by the backbone network to enhance the multi-scale feature representation process as follows:

[0016] The last three layers of feature maps extracted by the backbone network are respectively passed through 3x3 convolution layers, and spatial dimension sampling is performed through step 2 to extract key spatial information. Then, 1x1 convolution is used to compress and integrate channel information through linear transformation without changing the spatial size of the feature map, highlighting the key feature channel. After that, the whole is connected with the residual connection of the upper layer feature map. The highest layer feature map is only applied with a simple self-attention mechanism. Finally, the four layers of feature maps are flattened and spliced.

[0017] Further, the encoder includes an attention enhancement module: for modeling the interaction between channel and height, channel and width, and channel dimension based on the feature vector output by the feature fusion module, to realize the detection of irregular text and dense small text.

[0018] Further, the specific process of the above-mentioned process of passing the last three layers of feature maps through 3x3 convolution layers, and performing spatial dimension sampling through step 2 to extract key spatial information, and then using 1x1 convolution to compress and integrate channel information through linear transformation without changing the spatial size of the feature map, highlighting the key feature channel, and then connecting the whole with the residual connection of the upper layer feature map, and only applying the highest layer feature map with a simple self-attention mechanism, and finally flattening and splicing the four layers of feature maps is as follows:

[0019] Step one: given the input feature map First, get

[0020]

[0021] is the value of the input feature map at position (i * S + k - 1, j * S + 1 - 1) and channel c, S represents the step size of the convolution operation, is the feature map after spatial sampling,

[0022] Then, the feature map is mixed in the channel dimension by applying a 1x1 convolution kernel, which is represented as:

[0023]

[0024] wherein, is a 1x1 convolution kernel, is the feature map after feature sampling;

[0025] Step three: the feature map after feature sampling is fused layer by layer, which is represented as:

[0026]

[0027] wherein, y' represents the current feature map, is the lower layer feature map of y' after feature sampling, and y is the fused feature map;

[0028] Step four: given the last layer feature map X ∈ R B×C×H×W First, a 1-size convolution kernel Conv 1×1 Adjust the number of channels:

[0029] y = Conv 1×1 (x)

[0030] Then, it is divided into two parts according to the channel dimension, which are stored in a and b, respectively, which is represented as:

[0031]

[0032] Step five: based on the feature vector b, the query, key and value are generated according to the traditional method. Then, the attention score matrix is calculated by means of the softmax function, and the attention score matrix is applied to the value to complete the weighted summation operation;

[0033]

[0034] wherein: key_dim represents the dimension of the feature vector key, query t represents the transpose of the query vector;

[0035] ​​Step 6: Pass the value through a convolutional layer of size 1 to make its shape consistent with b', and then add it to b' to fuse information. Finally, input the summed result into a fully connected layer for linear projection.

[0036] b1' = Liner proj (value×b'+Conv 1×1 (value))

[0037] Step 7: After simply adding b and b', input the result into a feedforward neural network for processing, and then concatenate it with tensor a in the channel dimension to obtain the enhanced feature representation.

[0038] x=Concat(a,b+FFN(b+b2'),dim=1)

[0039] Where: Concat represents the concatenation operation, and FFN is a simple feedforward neural network;

[0040] Step 8: Concatenate and flatten the four feature maps, and use them as input to the encoder. The final feature output of this method is represented as follows:

[0041]

[0042] Furthermore, the process of modeling the channel and the height dimension is as follows:

[0043] Given an input vector X1∈R C×H×W First, adjust the vector shape to (W×H×C):

[0044] X′ 1-cw =P(X1,(C,W))

[0045] X′ 1-cw After average pooling and max pooling operations, and compression of the width dimension, the shape is (1×H×C):

[0046]

[0047] in: ψ represents average pooling operation, and ψ represents max pooling operation.

[0048] X′ 1-maxpool , X′ 1-avgpool The concatenation is performed, and weights are generated using the sigmoid function and applied to X1 to obtain:

[0049] X1'=σ[Cconcat(X′ 1-maxpool ,X′ 1-avgpool )]×X1

[0050] Where σ represents the sigmoid function.

[0051] Furthermore, the process of modeling the width dimension and channel dimension is as follows:

[0052] Step 1: Given an input vector X2∈R C×H×W First, adjust the vector shape to (H×C×W);

[0053] X′ 2-ch =P(X2,(C,H))

[0054] Step 2: X′ 2-ch After average pooling and max pooling operations, and compression of the width dimension, the shape is (1×C×W):

[0055]

[0056] in: ψ represents average pooling operation, and ψ represents max pooling operation.

[0057] Step 3: Set X′ 2-maxpool , X′ 2-avgpool The concatenation is performed, and weights are generated using the sigmoid function, then applied to X2 to obtain:

[0058] X2'=σ[Concat(X′ 2-maxpool ,X′ 2-avgpool )]×X2

[0059] Where σ represents the sigmoid function.

[0060] Furthermore, the process of modeling the channel dimensions is as follows:

[0061] Step 1: Given the input tensor X3∈R C×H×W The number of channels is compressed to 1 through pooling operations;

[0062] X′ 3-pool =τ(X3)

[0063] Where: τ represents channel pooling operation;

[0064] Step 2: X′ 3-pool The convolutional features are normalized using a standard convolutional layer and a batch normalization layer, resulting in the intermediate feature X′. 3-tmp ;

[0065] X′ 3-tmp =BatchNorm(Conv(X′) 3-pool ))

[0066] Step 3: Generate weights using the sigmoid function and apply them to X3, as follows:

[0067] X′3=X3×σ(X′ 3-tmp )

[0068] Finally, the feature vectors of the three branches are simply added together and averaged to complete the information integration of different branches. That is, the final query generated given the feature vector is:

[0069] Query=Avg(X′1+X′2+X′3)×W Q

[0070] Where: Avg represents the averaging operation, W Q This represents the weight matrix.

[0071] This invention provides a natural scene text detection method and model based on feature fusion and attention enhancement. The model utilizes ResNet50 as the backbone network to extract features, and fuses the four-layer feature maps through a feature fusion module to enhance multi-scale feature representation. A three-branch query generation module is introduced into the encoder to model the interaction between channels and spatial dimensions, improving the detection capability for irregular and dense small text. The effectiveness of this invention is demonstrated by subjective perception and performance metrics.

[0072] This invention employs two modules: a feature fusion module, which fuses information from feature maps at different scales, allowing high-level feature maps to retain more detailed texture information from the original image. This balances the advantages of high-level feature maps having rich semantics with the advantages of low-level feature maps having more texture details.

[0073] The attention enhancement module, used in the encoder's query generation stage, takes the feature map after feature fusion as input and models the relationships between channels and height, channels and width, and channel dimensions. This makes the model more sensitive to text instance information and enhances the detection of irregular text instances, ultimately improving the performance of text detection in natural scenes. Performance results on mainstream datasets, including subjective impressions, accuracy, and recall, demonstrate the significant implications of this invention for text detection in natural scenes. Attached Figure Description

[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0075] Figure 1 This is a flowchart of the method;

[0076] Figure 2 This is a structural diagram of a text detection model in natural scenes;

[0077] Figure 3 This is a structural diagram of the feature fusion module;

[0078] Figure 4 This is a structural diagram of the attention enhancement module. Detailed Implementation

[0079] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0080] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0081] Figure 1 This is a flowchart of the method;

[0082] A natural scene text detection method based on feature fusion and attention enhancement includes the following steps:

[0083] S1: Acquire the image to be detected;

[0084] S2: Construct a text detection model for natural scenes to extract text information from images;

[0085] S3: Train the text detection model in the natural scene to obtain a trained text detection model in the natural scene.

[0086] S4: Input the image to be detected into a trained text detection model in a natural scene to extract the text information contained in the image.

[0087] Steps S1 / S2 / S3 / S4 are executed sequentially;

[0088] The text information includes English letters and English punctuation.

[0089] Figure 2 This is a structural diagram of a text detection model in natural scenes;

[0090] The text detection model for natural scenes includes:

[0091] Backbone network: Feature extraction is performed using ResNet50;

[0092] Feature fusion module: fuses the feature maps extracted from the four layers of the backbone network to enhance multi-scale feature representation.

[0093] Encoder: Used to transform the image features into feature maps rich in text semantics and location information based on the fused image transmitted by the feature fusion module, enhance text features such as character edges and text structure, and suppress useless background noise;

[0094] Decoder: Used to restore the abstract features output by the encoder into a refined feature map that is aligned with the original image size and can be directly used for text localization.

[0095] The feature fusion module fuses the feature maps extracted from the last four layers of the backbone network to enhance multi-scale feature representation. The process is as follows:

[0096] The feature maps of the last three layers extracted from the backbone network are passed through 3×3 convolutional layers. Spatial dimension sampling is performed with a stride of 2 to extract key spatial information. Then, 1×1 convolution is used to compress and integrate channel information and highlight key feature channels through linear transformation without changing the spatial size of the feature maps. After that, the whole is residually connected with the upper layer feature maps. The highest layer feature map is only applied with a simple self-attention mechanism. Finally, the four layers of feature maps are flattened and stitched together. Figure 3 This is a structural diagram of the feature fusion module;

[0097] The encoder includes:

[0098] Attention Enhancement Module: Based on the feature vectors output by the feature fusion module, it models the interaction relationships between channels and height, channels and width, and channel dimensions to achieve the detection of irregular text and dense small text. Figure 4 This is a structural diagram of the attention enhancement module.

[0099] Furthermore, the lower three feature maps are processed by 3×3 convolutional layers with a stride of 2 for spatial dimension sampling to extract key spatial information. Then, 1×1 convolutions are used to compress and integrate channel information and highlight key feature channels through linear transformation without changing the spatial size of the feature maps. After that, the entire feature map is residually connected to the upper feature map. The highest feature map is only subject to a simple self-attention mechanism. Finally, the four feature maps are flattened and stitched together. The specific process is as follows:

[0100] Step 1: Given the input feature map First, it is obtained through spatial dimension sampling.

[0101]

[0102] Let be the value of the convolution kernel at position (k,l) and channel c, and S represent the stride of the convolution operation. It is the value of the input feature map at position (i*S+k-1, j*S+l-1) and channel c. It is a feature map after spatial sampling.

[0103] After that Channel blending is performed using a 1x1 convolutional kernel, as shown below:

[0104]

[0105] in, It is a 1x1 convolution kernel. Feature map after feature sampling;

[0106] Step 3: Perform feature fusion layer by layer on the feature maps after feature sampling, which is represented as:

[0107]

[0108] Where y' represents the current feature map, y' is the lower-level feature map after feature sampling, and y is the fused feature map;

[0109] Step 4: Given the last layer feature map X∈R B×C×H×W First, a convolution kernel of size 1 is used... 1×1 Adjust the number of channels:

[0110] y = Conv 1×1 (x)

[0111] Then, it is divided into two parts according to the channel dimension, and stored in a and b respectively, which are represented as follows:

[0112]

[0113] Step 5: Based on feature vector b, generate query, key, and value using traditional methods. Then, calculate the attention score matrix using the softmax function, and apply this attention score matrix to the value to complete the weighted summation operation.

[0114]

[0115] Where: key_dim represents the dimension of the feature vector key, query t Represents the transpose of the query vector;

[0116] Step 6: Pass the value through a convolutional layer of size 1 to make its shape consistent with b', and then add it to b' to fuse information. Finally, input the summed result into a fully connected layer for linear projection.

[0117] b′1=Liner proj (value×b'+Conv 1×1 (value))

[0118] Step 7: After simply adding b and b', input the result into a feedforward neural network for processing, and then concatenate it with tensor a in the channel dimension to obtain the enhanced feature representation.

[0119] x=Concat(a,b+FFN(b+b'2),dim=1)

[0120] Where: Concat represents the concatenation operation, and FFN is a simple feedforward neural network;

[0121] Step 8: Concatenate and flatten the four feature maps, and use them as input to the encoder. The final feature output of this method is represented as follows:

[0122]

[0123] The feature maps extracted from the last four layers of the backbone network are input into the feature fusion module. The resulting feature vector after multi-scale feature fusion is calculated as follows:

[0124] X final =F fusion (X)

[0125] Where X is the original feature vector, F fusion As a feature fusion processor, it samples the feature maps of the last four layers extracted from the backbone network layer by layer and fuses them with the feature maps of the upper layers. final This is the fused feature vector.

[0126] After the feature fusion module, the feature vector contains more detailed information, but it is still difficult to detect irregularly shaped text in the image. The attention enhancement module improves the detection performance of irregularly shaped text by introducing a three-branch query generation strategy. The feature vector is fed into the encoder, and then the interaction relationship between channels and height, channels and width, and channel dimensions is modeled and calculated as follows:

[0127] The process of modeling the channel and height dimension is as follows:

[0128] Given an input vector X1∈R C×H×W First, adjust the vector shape to (W×H×C):

[0129] X′ 1-cw =P(X1,(C,W))

[0130] X′ 1-cw After average pooling and max pooling operations, and compression of the width dimension, the shape is (1×H×C):

[0131]

[0132] in: ψ represents average pooling operation, and ψ represents max pooling operation.

[0133] X′ 1-maxpool , X′ 1-avgpool The concatenation is performed, and weights are generated using the sigmoid function and applied to X1 to obtain:

[0134] X1'=σ[Concat(X′ 1-maxpool ,X′ 1-avgpool )]×X1

[0135] Where σ represents the sigmoid function.

[0136] The process of modeling the width dimension and channel dimension is as follows:

[0137] Step 1: Given an input vector X2∈R C×H×W First, adjust the vector shape to (H×C×W);

[0138] X′ 2-ch =P(X2,(C,H))

[0139] Step 2: X′ 2-ch After average pooling and max pooling operations, and compression of the width dimension, the shape is (1×C×W):

[0140]

[0141] in: ψ represents average pooling operation, and ψ represents max pooling operation.

[0142] Step 3: Set X′ 2-maxpool , X′ 2-avgpool The concatenation is performed, and weights are generated using the sigmoid function, then applied to X2 to obtain:

[0143] X2'=σ[Concat(X′ 2-maxpool ,X′ 2-avgpool )]×X2

[0144] Where σ represents the sigmoid function.

[0145] The process of modeling the channel dimensions is as follows:

[0146] Step 1: Given the input tensor X3∈R C×H×W The number of channels is compressed to 1 through pooling operations;

[0147] X′ 3-pool =τ(X3)

[0148] Where: τ represents channel pooling operation;

[0149] Step 2: X′ 3-pool The convolutional features are normalized using a standard convolutional layer and a batch normalization layer, resulting in the intermediate feature X′. 3-tmp ;

[0150] X′ 3-tmp =BatchNorm(Conv(X′) 3-pool ))

[0151] Step 3: Generate weights using the sigmoid function and apply them to X3, as follows:

[0152] X′3=X3×σ(X′ 3-tmp )

[0153] Finally, the feature vectors of the three branches are simply added together and averaged to complete the information integration of different branches. That is, the final query generated given the feature vector is:

[0154] Query=Avg(X′1+X′2+X′3)×W Q

[0155] Where: Avg represents the averaging operation, W Q This represents the weight matrix.

[0156] Model the relationships between channels and height, channels and width, and channel dimensions of the feature vectors obtained by the feature fusion module.

[0157] The modeling and calculations are as follows:

[0158] X1'=F c-h ×X1

[0159] X2'=F c-w ×X1

[0160] X3'=F c-c ×X3

[0161] Wherein: F c-hF represents the feature interaction between the channel and the height dimension. c-w F represents the feature interaction between the channel and width dimensions. c-c , representing the feature interaction of the channel dimension.

[0162] The initial calculation of the query vector is as follows:

[0163] Query=Avg(X′1+X′2+X′3)×W Q

[0164] Where Avg represents the average operation W Q The weight matrix to be learned is represented by Query, which is used to generate the model's self-attention score and is only used in the first three layers of the encoder. The generation of the key and value vectors does not involve the above calculations.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A natural scene text detection method based on feature fusion and attention enhancement, characterized in that, Includes the following steps: Acquire the image to be detected; Construct a text detection model for natural scenes to extract text information from images; The text detection model in the natural scene is trained to obtain a trained text detection model in the natural scene. The image to be detected is input into a pre-trained text detection model in a natural scene to extract the text information contained in the image.

2. The natural scene text detection method based on feature fusion and attention enhancement according to claim 1, characterized in that, The text information includes English letters and English punctuation.

3. The natural scene text detection method based on feature fusion and attention enhancement according to claim 1, characterized in that, The text detection model for natural scenes includes: Backbone network: Feature extraction is performed using ResNet50; Feature fusion module: fuses the feature maps extracted from the four layers of the backbone network to enhance multi-scale feature representation. Encoder: Used to transform the image features into feature maps rich in text semantics and location information based on the fused image transmitted by the feature fusion module, enhance text features such as character edges and text structure, and suppress useless background noise; Decoder: Used to restore the abstract features output by the encoder into a refined feature map that is aligned with the original image size and can be directly used for text localization.

4. The natural scene text detection method based on feature fusion and attention enhancement according to claim 3, characterized in that, The feature fusion module fuses the feature maps extracted from the last four layers of the backbone network to enhance multi-scale feature representation. The process is as follows: The feature maps of the last three layers extracted from the backbone network are passed through 3×3 convolutional layers. Spatial dimension sampling is performed with a stride of 2 to extract key spatial information. Then, 1×1 convolution is used to compress and integrate channel information and highlight key feature channels through linear transformation without changing the spatial size of the feature maps. After that, the whole is residually connected with the upper layer feature maps. The highest layer feature map is only applied with a simple self-attention mechanism. Finally, the four layers of feature maps are flattened and stitched together.

5. The natural scene text detection method based on feature fusion and attention enhancement according to claim 3, characterized in that, The encoder includes: Attention Enhancement Module: Based on the feature vectors output by the feature fusion module, it models the interaction relationships between channels and height, channels and width, and channel dimensions to achieve the detection of irregular text and dense small text.

6. The natural scene text detection method based on feature fusion and attention enhancement according to claim 4, characterized in that, The process of applying 3×3 convolutional layers to the lower three feature maps, sampling spatial dimensions with a stride of 2 to extract key spatial information, and then using 1×1 convolutions to compress and integrate channel information and highlight key feature channels through linear transformation without changing the spatial size of the feature maps, is then performed. The entire dataset is then residually connected to the upper feature maps. The highest-level feature map is only applied with a simple self-attention mechanism. Finally, the four feature maps are flattened and stitched together. The specific process is as follows: Step 1: Given the input feature map First, it is obtained through spatial dimension sampling. Let be the value of the convolution kernel at position (k,l) and channel c, and S represent the stride of the convolution operation. It is the value of the input feature map at position (i*S+k-1, j*S+l-1) and channel c. It is a feature map after spatial sampling. After that Channel blending is performed using a 1x1 convolutional kernel, as shown below: in, It is a 1x1 convolution kernel. Feature map after feature sampling; Step 3: Perform feature fusion layer by layer on the feature maps after feature sampling, which is represented as: Where y' represents the current feature map, y' is the lower-level feature map after feature sampling, and y is the fused feature map; Step 4: Given the last layer feature map X∈R B×C×H×W First, a convolution kernel of size 1 is used... 1×1 Adjust the number of channels: y=Conv 1×1 (x) Then, it is divided into two parts according to the channel dimension, and stored in a and b respectively, which are represented as follows: Step 5: Based on feature vector b, generate query, key, and value using traditional methods. Then, calculate the attention score matrix using the softmax function, and apply this attention score matrix to the value to complete the weighted summation operation. Where: key_dim represents the dimension of the feature vector key, query t Represents the transpose of the query vector; Step 6: Pass the value through a convolutional layer of size 1 to make its shape consistent with b', and then add it to b' to fuse information. Finally, input the summed result into a fully connected layer for linear projection. b1'=Liner proj (value×b'+Conv 1×1 (value)) Step 7: After simply adding b and b', input the result into a feedforward neural network for processing, and then concatenate it with tensor a in the channel dimension to obtain the enhanced feature representation. x=Concat(a,b+FFN(b+b'2),dim=1) Where: Concat represents the concatenation operation, and FFN is a simple feedforward neural network; Step 8: Concatenate and flatten the four feature maps, and use them as input to the encoder. The final feature output of this method is represented as follows:

7. The natural scene text detection method based on feature fusion and attention enhancement according to claim 5, characterized in that, The process of modeling channels and height dimensions is as follows: Given an input vector X1∈R C×H×W First, adjust the vector shape to (W×H×) C): X′ 1-cw <P(X1,(C,W)) X′ 1-cw After average pooling and max pooling operations, and compression of the width dimension, the shape is (1×H×C): in: ψ represents average pooling operation, and ψ represents max pooling operation. X′ 1-maxpool , X′ 1-avgpool The concatenation is performed, and weights are generated using the sigmoid function and applied to X1 to obtain: X1'=σ[Concat(X′ 1-maxpool ,X′ 1-avgpool )]×X1 Where σ represents the sigmoid function.

8. A natural scene text detection method based on feature fusion and attention enhancement according to claim 5, characterized in that, The process of modeling the width dimension and the channel dimension is as follows: Step 1: Given an input vector X2∈R C×H×W First, adjust the vector shape to (H×C×W); X′ 2-ch =P(X2,(C,H)) Step 2: X′ 2-ch After average pooling and max pooling operations, and compression of the width dimension, the shape is (1×C×W): in: ψ represents average pooling operation, and ψ represents max pooling operation. Step 3: Set X′ 2-maxpool , X′ 2-avgpool The concatenation is performed, and weights are generated using the sigmoid function, then applied to X2 to obtain: X2‘=σ[Concat(X′ 2-maxpool ,X′ 2-avgpool )]×X2 Where σ represents the sigmoid function.

9. A natural scene text detection method based on feature fusion and attention enhancement according to claim 5, characterized in that, The process of modeling the channel dimensions is as follows: Step 1: Given the input tensor X3∈R C×H×W The number of channels is compressed to 1 through pooling operations; X′ 3-pool =τ(X3) Where: τ represents channel pooling operation; Step 2: X′ 3-pool The convolutional features are normalized using a standard convolutional layer and a batch normalization layer, resulting in the intermediate feature X′. 3-tmp ; X′ 3-tmp =BatchNorm(Conv(X′ 3-pool )) Step 3: Generate weights using the sigmoid function and apply them to X3, as follows: X′3=X3×σ(X′ 3-tmp ) Finally, the feature vectors of the three branches are simply added together and averaged to complete the information integration of different branches. That is, the final query generated given the feature vector is: Query=Avg(X′1+X′2+X′3)×W Q Where: Avg represents the averaging operation, W Q This represents the weight matrix.