Natural scene text detection method based on feature pyramid

By adaptively adjusting feature weights through a feature pyramid enhancement module and an SE attention mechanism, combined with an FDB model, the complex background and multi-scale problems in Chinese text detection in natural scenes are solved, achieving high-precision and robust text detection.

CN120953973APending Publication Date: 2025-11-14WUHAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510956373.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies for Chinese text detection in natural scenes suffer from limited detection accuracy, difficulty in handling complex backgrounds and multi-scale text, and insufficient model generalization ability. In particular, traditional methods and deep learning models struggle to effectively integrate feature information at different scales in the context of multi-form and complex Chinese text.

Method used

We adopt a natural scene text detection method based on feature pyramids. The feature pyramid enhancement module processes multi-scale feature maps, and combines upsampling and downsampling operations with SE attention mechanism to adaptively adjust the weights of features at each scale and construct a global feature representation. The FDB model is then used for analysis.

Benefits of technology

It significantly improves the accuracy and robustness of Chinese text detection in natural scenes, reduces false detection and false negative rates, optimizes the feature fusion process, reduces computational load, improves model inference speed, and meets the real-time requirements of practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953973A_ABST
    Figure CN120953973A_ABST
Patent Text Reader

Abstract

The invention discloses a natural scene text detection method based on a feature pyramid, and the method comprises the steps: extracting a multi-scale feature map based on an input natural scene image; the multi-scale feature map is processed through a feature pyramid enhancement module, the feature pyramid enhancement module comprises up-sampling and down-sampling operation, an SE attention mechanism is fused, and the weight of each scale feature is adjusted in a self-adaptive mode; performing feature fusion on the processed feature map, and constructing global feature representation; and analyzing the global feature representation through an FDB model, and outputting a text detection result. The invention further discloses a natural scene text detection device based on the feature pyramid, corresponding equipment and a storage medium. According to the method, the precision and robustness of Chinese text detection in a natural scene can be remarkably improved, and the problem of text detection under a complex background is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and pattern recognition technology, and more specifically, to a method for natural scene text detection based on feature pyramids. Background Technology

[0002] With the rapid development of artificial intelligence technology, natural scene text detection and recognition technology plays an important role in many practical scenarios, such as cargo label recognition in logistics and warehousing management, information reading from traffic signs in smart cities, and text monitoring in intelligent security surveillance. Text information in natural scenes is characterized by multi-scale, multi-form, and complex backgrounds, especially Chinese text, whose unique stroke structure and arrangement (such as vertical, curved, and slanted strokes) increase the complexity of detection. Accurate and efficient detection of Chinese text in natural scenes is of great significance for information extraction and subsequent processing.

[0003] Current technologies have several shortcomings in Chinese text detection in natural scenes. Traditional methods, such as those based on edge detection and connected component analysis, have limited accuracy in complex backgrounds and with significant differences in text scale, making them unsuitable for real-world applications. While deep learning-based text detection models (such as DBNet) have improved in recent years, their unidirectional top-down feature fusion approach struggles to effectively integrate feature information from different scales when processing cross-scale Chinese text. Furthermore, existing public datasets lack sufficient support for Chinese text, particularly in realistic scenarios with complex backgrounds, multi-angle shots, and text occlusion. This results in poor generalization ability of these models in practical applications, failing to meet the demands for high accuracy and robustness. Summary of the Invention

[0004] In view of at least one defect or improvement need of the prior art, this application provides a natural scene text detection method based on feature pyramid, which can solve at least one of the problems existing in the above background art.

[0005] To achieve the above objectives, according to the first aspect of this application, a natural scene text detection method based on feature pyramids is provided, the method comprising: Extract multi-scale feature maps from the input natural scene image; The multi-scale feature map is processed by a feature pyramid enhancement module, which includes upsampling and downsampling operations, integrates the SE attention mechanism, and adaptively adjusts the weights of features at each scale. The processed feature maps are fused to construct a global feature representation; The global feature representation is analyzed using the FDB model, and the text detection results are output.

[0006] Furthermore, in the above-mentioned natural scene text detection method based on feature pyramids, the extraction of feature maps of the image to be detected includes extracting feature maps of four different scales, which are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original image to be detected.

[0007] Furthermore, the aforementioned natural scene text detection method based on feature pyramids, before extracting multi-scale feature maps, also includes constructing a Chinese text image dataset and training the FDB model, specifically including: Collect natural scene images containing Chinese text, including various scenes; Add polygonal vertex annotations to the text region to create an annotation file; Enhanced samples are generated through random perspective transformation, illumination perturbation, and background synthesis. The FDB model is trained based on the labeled file and the augmented samples.

[0008] Furthermore, in the above-mentioned natural scene text detection method based on feature pyramids, the upsampling operation adopts a 2x linear upsampling, and the downsampling operation adopts a convolution operation with a stride of 2.

[0009] Furthermore, in the aforementioned natural scene text detection method based on feature pyramids, the SE attention mechanism obtains global information for each channel through global average pooling, and uses two fully connected layers to learn the dependencies between channels to generate corresponding weight coefficients.

[0010] Furthermore, in the above-mentioned natural scene text detection method based on feature pyramids, the feature fusion step includes performing a summation operation on feature maps of the same size, upsampling all feature maps, and cascading them to a uniform size.

[0011] Furthermore, in the aforementioned natural scene text detection method based on feature pyramids, the FDB model uses ResNet50 as the backbone network and is trained using the Adam optimizer.

[0012] According to a second aspect of this application, a natural scene text detection device based on feature pyramids is also provided, comprising: The feature map extraction module is used to extract multi-scale feature maps based on the input natural scene image; The sampling module is used to process the multi-scale feature map through the feature pyramid enhancement module. The feature pyramid enhancement module includes upsampling and downsampling operations, integrates the SE attention mechanism, and adaptively adjusts the weights of features at each scale. The feature fusion module is used to fuse features in the processed feature map and construct a global feature representation. The results output module is used to analyze the global feature representation through the FDB model and output the text detection results.

[0013] According to a third aspect of this application, a natural scene text detection device based on feature pyramids is also provided, which includes at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program that, when executed by the processing unit, causes the processing unit to perform the steps of any of the methods described above.

[0014] According to a fourth aspect of this application, a storage medium is also provided, which stores a computer program executable by a feature pyramid-based natural scene text detection device, which, when run on the feature pyramid-based natural scene text detection device, causes the feature pyramid-based natural scene text detection device to perform the steps of any of the methods described above.

[0015] In summary, compared with the prior art, the above-described technical solutions conceived in this application can achieve the following beneficial effects: The natural scene text detection method based on feature pyramid provided in this application processes multi-scale feature maps by using a feature pyramid enhancement module. Combined with upsampling, downsampling operations and SE attention mechanism, it achieves adaptive adjustment of feature weights at each scale, which can significantly improve the accuracy and robustness of Chinese text detection in natural scenes. It effectively solves the problem of text detection in complex backgrounds, can accurately identify text information of different scales and forms, reduce false detection and false negative rates, and optimize the feature fusion process to reduce data redundancy and computation, improve model inference speed, meet the real-time requirements of practical applications, and ensure high detection performance under adverse conditions such as insufficient lighting, complex backgrounds or text occlusion. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating the natural scene text detection method based on feature pyramids provided in this application embodiment; Figure 2 This is a schematic diagram of the FPEMs network structure provided in the embodiments of this application; Figure 3 This is a schematic diagram of the FDB model structure provided in an embodiment of this application; Figure 4This is a schematic diagram of the feature fusion network (FFM) structure provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Furthermore, the technical features involved in the various embodiments described below can be combined with each other as long as they do not conflict with each other.

[0019] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0020] Figure 1 A flowchart illustrating the natural scene text detection method based on feature pyramids provided in this application is shown below. Figure 1 As shown in the embodiment of the application, the natural scene text detection method based on feature pyramids includes: Extract multi-scale feature maps from the input natural scene image; The multi-scale feature map is processed by a feature pyramid enhancement module, which includes upsampling and downsampling operations, integrates the SE attention mechanism, and adaptively adjusts the weights of features at each scale. The processed feature maps are fused to construct a global feature representation; The global feature representation is analyzed using the FDB model, and the text detection results are output.

[0021] Specifically, the input natural scene image is fed into the backbone network to extract multi-scale feature maps. These feature maps correspond to different resolutions of the original image, such as 1 / 4, 1 / 8, 1 / 16, and 1 / 32, thereby capturing multi-dimensional information from fine details to high-level semantics. For example, features at different scales, including buildings and road signs, are extracted from street view images.

[0022] like Figure 2As shown, the FDB model in this application employs Feature Pyramid Enhancement Modules (FPEMs), which have a U-shaped structure and include two stages: scale-up enhancement and scale-down enhancement. In the scale-up enhancement stage, to address the difficulty of detecting small target text, an iterative 2x linear upsampling operation is used to gradually increase the size of the feature maps, restoring a resolution close to the original image and capturing fine spatial details. Adjacent feature maps are matched in size through 2x linear upsampling, and element-wise addition is used as a fusion strategy to preserve original information and fuse complementary information between different scales. In the scale-down enhancement stage, a convolution operation with a stride of 2 is used to reduce the size of the feature maps, extracting more abstract and representative features.

[0023] In both the upsampling and downsampling stages of FPEMs, the SE attention mechanism is integrated. After dimensionality reduction and cross-channel information integration of the feature map using 1×1 convolutions, the SE mechanism obtains global information for each channel through global average pooling and learns the dependencies between channels using two fully connected layers to generate weight coefficients. The scale operation redistributes the weights to the original feature map, realizing the reweighting of features, enabling the model to focus on key text regions and improve feature perception capabilities.

[0024] The feature maps processed by the feature pyramid enhancement module are then fused using a Feature Fusion Network (FFM) strategy. First, a summation operation is performed on feature maps of the same size to merge redundant feature information and reduce data redundancy. Then, all feature maps are upsampled and concatenated to a uniform size to construct a global feature representation. This effectively avoids the problem of a sharp increase in dimensionality in traditional fusion methods, preserving the advantages of features at each scale while reducing computational cost.

[0025] The constructed global feature representation is input into a well-trained FDB model, which is trained on a CTD dataset containing samples of complex scenes and can effectively recognize Chinese text in natural scenes. Figure 3 As shown, the FDB model analyzes global feature representations to output text detection results, accurately locating text regions in images and providing accurate data support for subsequent text recognition and other tasks. In practical applications, such as the recognition of Chinese information on cargo labels in automated sorting systems in logistics and warehousing management, or the reading of Chinese content on traffic signs in smart cities, it can operate efficiently, improving text detection accuracy and robustness.

[0026] The natural scene text detection method based on feature pyramid provided in this application uses a feature pyramid enhancement module to process multi-scale feature maps. By combining upsampling, downsampling operations and SE attention mechanism, it achieves adaptive adjustment of feature weights at each scale, which can significantly improve the accuracy and robustness of Chinese text detection in natural scenes. It effectively solves the problem of text detection in complex backgrounds, can accurately identify text information of different scales and forms, reduce false detection and false negative rates, and optimize the feature fusion process to reduce data redundancy and computational load, improve model inference speed, meet the real-time requirements of practical applications, and ensure high detection performance under adverse conditions such as insufficient lighting, complex backgrounds or text occlusion.

[0027] Optionally, the natural scene text detection method based on feature pyramid provided in this application embodiment includes extracting feature maps of four different scales, namely 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original image to be detected.

[0028] Specifically, it addresses the common scenario in Chinese text where small-sized text (such as product labels and license plate numbers) coexists with large-sized text (such as billboards and building signs). It extracts feature maps of four different scales as input, corresponding to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image, respectively, covering multi-dimensional information from fine details to high-level semantics.

[0029] Optionally, the natural scene text detection method based on feature pyramids provided in this application embodiment further includes constructing a Chinese text image dataset and training the FDB model before extracting multi-scale feature maps, specifically including: Collect natural scene images containing Chinese text, including various scenes; Add polygonal vertex annotations to the text region to create an annotation file; Enhanced samples are generated through random perspective transformation, illumination perturbation, and background synthesis. The FDB model is trained based on the labeled file and the augmented samples.

[0030] Specifically, current publicly available datasets lack image data for Chinese scenes, especially samples containing realistic interference factors such as complex background textures, multi-angle shots, text occlusion, and bending deformation. This makes it difficult for trained models to adapt to real-world scenarios, often resulting in missed detections and false detections. To address these issues, this application constructs the Chinese Text Image Dataset (CTD). Starting from practical engineering needs, it adopts a synthesis and selection integration approach to create a dataset closely resembling real-world applications. During the construction process, representative Chinese text images were first selected from existing public datasets such as CTW1500, MSRA-TD500, and PaddleOCR. Simultaneously, new images were synthesized at a 1:4 ratio, ultimately forming a training set containing 5000 images.

[0031] In the image synthesis stage, to ensure the broad applicability of the dataset and the robustness of the algorithm, diverse background images were selected from public background datasets. Chinese characters of different colors were embedded based on the background brightness to simulate the blending of text and background in real-world scenes. Figure 4 As shown.

[0032] The CTD dataset focuses on covering complex text geometry and interference factors in real-world scenarios. For text rotation, scaling, projection distortion, bending, and elongated layouts, images with these features are integrated. For issues like reflections, perspective transformations, and distorted fonts, images from real-world scenes affected by lighting and shooting angles are collected, such as reflective billboards and signs photographed at an angle. These images contain multiple layers of background texture and exhibit complex changes under different lighting conditions, greatly reflecting the challenges faced in Chinese text detection in practical applications.

[0033] By comparing with mainstream text detection models on multiple public datasets, the FDB model demonstrates significant advantages in accuracy and robustness for Chinese text detection in complex scenarios, effectively meeting the requirements of practical engineering applications for high precision and strong adaptability in Chinese text detection algorithms. This strategy of using pure Chinese text as the training basis effectively avoids interference from multilingual mixing on model learning, thereby improving the detection model's feature learning ability and robustness towards Chinese text.

[0034] Optionally, in the natural scene text detection method based on feature pyramid provided in this application embodiment, the upsampling operation adopts 2x linear upsampling, and the downsampling operation adopts a convolution operation with a stride of 2.

[0035] Specifically, size matching is achieved between adjacent feature maps through a 2x linear upsampling, which smoothly increases the feature map resolution and avoids introducing excessive noise. Simultaneously, a simple element-wise addition strategy is used, adding the upsampled feature map element-wise to the feature map at the corresponding scale. This preserves the original information in the feature maps while effectively fusing complementary information between different scales, maintaining the integrity and continuity of the information.

[0036] Assume the first Layer input features are First, the feature map of the current layer is upsampled and then added to the previous layer. :

[0037] Then, after passing through DWConv and a 1×1 convolution, the convolutional enhanced features are obtained. :

[0038] In the scale-down enhancement stage, FPEM reduces the size of the feature map after feature fusion by using a convolution operation with a stride of 2. This process helps to extract more abstract and representative features. It is the downsampled feature map of the current layer. These are enhanced feature maps from the upsampling stage:

[0039] Then, after passing through DWConv and 1×1 convolution, convolutional enhancement features are obtained. : .

[0040] In text detection tasks in natural scenes, models need to process multi-scale and multi-form text information. Traditional feature fusion methods are prone to data redundancy and dimensionality expansion, affecting model training efficiency and inference speed. To solve this problem, this application adopts the Feature Fusion Network (FFM) strategy, which effectively improves model performance and efficiency through an innovative two-step fusion mechanism. In practical applications, such as processing street view images, text may exist in regions of different sizes, corresponding to feature maps of different scales output by the model. The FFM strategy first performs a summation operation on feature maps of the same size, significantly reducing data redundancy by merging redundantly expressed feature information. Taking cargo label detection in logistics and warehousing scenarios as an example, summing the feature maps output by multiple detection branches at the same scale can avoid repeatedly calculating similar text edges, structures, and other features, thus reducing the computational burden.

[0041] After summing, the FFM strategy upsamples all feature maps and concatenates them to a uniform size to construct a global feature representation. This step effectively avoids the problem of a sharp increase in dimensionality caused by direct concatenation in traditional fusion methods. This fusion method not only preserves the advantages of features at each scale but also reduces computation, enabling the model to converge faster during training and respond quickly during actual inference, meeting the efficiency requirements of scenarios such as real-time road sign recognition in intelligent transportation.

[0042] The specific operations of the scaling down phase (SE) are the same as those of the scaling up phase (SE). Optionally, in the natural scene text detection method based on feature pyramid provided in this application embodiment, the SE attention mechanism obtains global information of each channel through global average pooling operation, and uses two fully connected layers to learn the dependency relationship between channels to generate corresponding weight coefficients.

[0043] Specifically, in text detection tasks in natural scenes, text features at different scales have varying impacts on the detection results. Traditional methods struggle to effectively balance the weights of features at different scales, leading to performance limitations when handling complex scenarios. To address this issue, this application places the SE mechanism after a 1×1 convolution and improves the model's ability to handle multi-scale text features through a unique weight adjustment strategy. In practical applications such as intelligent traffic road sign recognition, road signs may exhibit vastly different scales due to varying shooting distances, ranging from tiny warning signs hundreds of meters away to large directional signs at close range at intersections. 1×1 convolutions can reduce the dimensionality of feature maps and integrate cross-channel information, providing a simpler and more effective input for the SE (Search Engine) mechanism. Based on this, the SE mechanism obtains global information for each channel through global average pooling, then uses two fully connected layers to learn the dependencies between channels, generating corresponding weight coefficients based on the global importance of features at different scales.

[0044] Squeeze applies to each scale Feature map Perform global average pooling to compress the feature maps at each scale into a global description. :

[0045] The excitation operation yields the weights for each scale. :

[0046] Where δ represents the ReLU activation function and σ represents the Sigmoid function. and This is a learnable weight matrix.

[0047] The scale operation assigns weights to each scale. Reassigned to the original feature map This involves multiplying the features at each scale by their corresponding weight values ​​to reweight the features. The formula is as follows: .

[0048] Optionally, the feature fusion step in the natural scene text detection method based on feature pyramid provided in this application includes performing a summation operation on feature maps of the same size, upsampling all feature maps, and cascading them to a uniform size.

[0049] Specifically, by performing a summation operation on feature maps of the same size during the feature fusion step, redundant feature information can be merged, effectively reducing data redundancy. Subsequently, all feature maps are upsampled and concatenated to a uniform size, enabling the construction of a global feature representation while avoiding the problem of a sharp increase in dimensionality caused by direct concatenation in traditional fusion methods. In this way, the advantages of features at each scale are preserved while reducing computational load, allowing the model to converge faster during training, improving inference speed, and meeting the efficiency requirements of practical applications. This enables efficient and accurate feature fusion and detection in text detection tasks in natural scenes.

[0050] Optionally, in the natural scene text detection method based on feature pyramid provided in this application embodiment, the FDB model uses ResNet50 as the backbone network and is trained using the Adam optimizer.

[0051] Optionally, embodiments of this application also provide a natural scene text detection device based on feature pyramids, comprising: The feature map extraction module is used to extract multi-scale feature maps based on the input natural scene image; The sampling module is used to process the multi-scale feature map through the feature pyramid enhancement module. The feature pyramid enhancement module includes upsampling and downsampling operations, integrates the SE attention mechanism, and adaptively adjusts the weights of features at each scale. The feature fusion module is used to fuse features in the processed feature map and construct a global feature representation. The results output module is used to analyze the global feature representation through the FDB model and output the text detection results.

[0052] The following describes the natural scene text detection method based on feature pyramids provided in this application and its technical effects, using a specific embodiment as an example: (1) Selecting the dataset and experimental environment Experiments were conducted on the CTW1500 and MSRA-TD500 datasets. The FDB model used ResNet50 as the backbone network, with a batch size of 8, and was trained for 1200 epochs.

[0053] To address the complexity of background and text types in natural scene text images, this embodiment preserves the original image information during data preprocessing and employs two 2080Ti GPUs for parallel training to improve model training efficiency and generalization ability. Due to significant gradient noise, the initial learning rate is set to 0.01, and a decay strategy of halving the learning rate every three epochs is adopted to maintain the stability and convergence of the training process. The Adam optimizer is selected as the optimization algorithm.

[0054] (2) Selecting a comparison model CTPN, EAST, PixelLink, PAN, TCM-DBNet, and ODM-DBNet were selected as comparison models. Among them, CTPN and EAST use candidate box regression methods, PixelLink combines image segmentation and connectivity relationships, PAN improves detection performance through pixel aggregation, and TCM-DBNet and ODM-DBNet are improved models based on DBNet.

[0055] (3) Performance evaluation and analysis Precision, Recall, and F1 score were used as performance evaluation metrics.

[0056] The specific formula is as follows:

[0057]

[0058]

[0059] Where TP represents true positives (the number of correctly detected text boxes), FP represents false positives (the number of falsely detected text boxes), and FN represents false negatives (the number of text boxes that were missed).

[0060] (4) Implementation examples and results analysis The trained FDB model and the comparison model were tested on the test sets of the CTW1500 and CTD datasets. For each model, the test results were recorded, including the number of detected text boxes, the number of correctly detected text boxes, the number of missed text boxes, and the number of falsely detected text boxes, to provide data support for subsequent performance evaluation.

[0061] Experimental results show that the FDB model outperforms other models on both datasets, as shown in Tables 1 and 2. The superior performance of FDB in Chinese text detection is evident, especially in its ability to better distinguish morphologically changing Chinese characters from the background in complex contexts, avoiding misinterpretation of background information as text. Furthermore, it outperforms other models when text and background are difficult to distinguish.

[0062] On the CTW1500 dataset, the FDB model's F1 score not only surpasses all other comparison models but also shows a 0.4% improvement over the second-place ODM-DBNet model. Compared to the classic detection model CTPN, the FDB model's F1 score is improved by 23%. In terms of precision, FDB still offers a 0.2% improvement over the latest ODM-DBNet. Compared to the baseline model DBNet, the FDB model improves precision by 2.4%, recall by 1.3%, and F1 score by 1.3%, while its recall score is slightly lower than ODM-DBNet; the choice of detection threshold influences the model's recall capability to some extent. Compared to all comparison models, the FDB model is superior in detecting Chinese text in complex backgrounds. Table 1 Comparison of CTW1500 dataset results

[0063] On the MSRA-TD500 dataset, the FDB model shows improvements in both precision and F1 score compared to other models. Compared to the baseline model DBNet, the FDB model improves precision by 0.9%, recall by 2.6%, and F1 score by 1.5%. It also shows improvements in both precision and F1 score compared to the latest models TCM-DBNet and ODM-DBNet.

[0064] Table 2 Comparison of MSRA-TD500 Dataset Results

[0065] In summary, the comparative experimental results on the CTW1500 and MSRA-TD500 datasets fully demonstrate that the scene text detection model FDB based on attention feature pyramids proposed in this application has higher text detection accuracy, especially in Chinese text detection. In real-world scene image detection containing complex backgrounds and multi-scale text, the FDB model effectively solves the cross-scale feature fusion problem through the synergistic effect of the Feature Pyramid Enhancement Modules (FPEMs) and the SE attention mechanism, and can accurately identify various types of Chinese text of different sizes. Furthermore, targeted training based on the self-built CTD dataset makes the FDB model more adaptable to irregular arrangements of Chinese text and the influence of complex backgrounds. When dealing with extreme cases such as backlighting and tilting, the detection accuracy is significantly improved compared to traditional models. The above experimental data fully verify the technical innovation and practicality of the FDB model in the field of Chinese text detection, providing a more efficient and reliable solution for the development of text detection technology in natural scenes.

[0066] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, as well as magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0067] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0068] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0069] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0070] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0071] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0072] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0073] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0074] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

[0075] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0076] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A natural scene text detection method based on feature pyramids, characterized in that, include: Extract multi-scale feature maps from the input natural scene image; The multi-scale feature map is processed by a feature pyramid enhancement module, which includes upsampling and downsampling operations, integrates the SE attention mechanism, and adaptively adjusts the weights of features at each scale. The processed feature maps are fused to construct a global feature representation; The global feature representation is analyzed using the FDB model, and the text detection results are output.

2. The natural scene text detection method based on feature pyramids as described in claim 1, characterized in that, The extraction of feature maps from the image to be detected includes extracting feature maps at four different scales, which are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original image to be detected.

3. The natural scene text detection method based on feature pyramids as described in claim 1, characterized in that, Before extracting multi-scale feature maps, the process also includes constructing a Chinese text-image dataset and training the FDB model, specifically including: Collect natural scene images containing Chinese text, including various scenes; Add polygonal vertex annotations to the text region to create an annotation file; Enhanced samples are generated through random perspective transformation, illumination perturbation, and background synthesis. The FDB model is trained based on the labeled file and the augmented samples.

4. The natural scene text detection method based on feature pyramids as described in claim 1, characterized in that, The upsampling operation uses a 2x linear upsampling, and the downsampling operation uses a convolution operation with a stride of 2.

5. The natural scene text detection method based on feature pyramids as described in claim 1, characterized in that, The SE attention mechanism obtains global information for each channel through global average pooling, learns the dependencies between channels using two fully connected layers, and generates corresponding weight coefficients.

6. The natural scene text detection method based on feature pyramid as described in claim 1, characterized in that, The feature fusion step includes performing a summation operation on feature maps of the same size, upsampling all feature maps, and cascading them to a uniform size.

7. The natural scene text detection method based on feature pyramid as described in claim 3, characterized in that, The FDB model uses ResNet50 as the backbone network and is trained using the Adam optimizer.

8. A natural scene text detection device based on feature pyramids, characterized in that, include: The feature map extraction module is used to extract multi-scale feature maps based on the input natural scene image; The sampling module is used to process the multi-scale feature map through the feature pyramid enhancement module. The feature pyramid enhancement module includes upsampling and downsampling operations, integrates the SE attention mechanism, and adaptively adjusts the weights of features at each scale. The feature fusion module is used to fuse features in the processed feature map and construct a global feature representation. The results output module is used to analyze the global feature representation through the FDB model and output the text detection results.

9. A natural scene text detection device based on feature pyramids, characterized in that, The method includes at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program that, when executed by the processing unit, causes the processing unit to perform the steps of the method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, It stores a computer program executable by a feature pyramid-based natural scene text detection device, which, when run on the feature pyramid-based natural scene text detection device, causes the feature pyramid-based natural scene text detection device to perform the steps of the method according to any one of claims 1 to 7.