Scene text detection method and device

Through the combination of multi-scale backbone network, feature pyramid enhancement network and dual attention module, the existing scene text detection methods have poor detection effect and long inference time in complex environments, and efficient and accurate scene text detection is achieved.

CN119992528APending Publication Date: 2025-05-13BEIJING INST OF ENVIRONMENTAL FEATURES
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510058141.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing scene text detection methods have poor detection results in complex natural environments and harsh imaging conditions, and the model inference time is long, making it difficult to achieve the balance of detection performance and efficiency.

Method used

Multi-scale backbone network module is used to extract multi-scale feature, build the initial feature pyramid, and perform two-stage feature enhancement processing through the feature pyramid enhancement network module. At the same time, the dual attention module is used to fusion across scale features to improve the ability to characterize multi-scale features and the adequacy of fusion across scale features.

Benefits of technology

The performance and real-time performance of the scene text detection model are improved, the multi-scale feature representation ability and the adequacy of cross-scale feature fusion are enhanced, and the detection efficiency and accuracy are balanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992528A_ABST
    Figure CN119992528A_ABST
Patent Text Reader

Abstract

The invention discloses a scene text detection method and device, and belongs to the field of computers. The method comprises the following steps: performing multi-scale feature extraction on a to-be-detected image by using a multi-scale backbone network module to obtain a multi-scale feature map; constructing an initial feature pyramid based on the multi-scale feature map; using a feature pyramid enhancement network module to perform two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map, the two stages including a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage including a grouping shuffling convolution operation; performing cross-scale feature fusion on the enhanced feature map by using a double attention module to obtain a multi-scale fusion feature map; and performing scene text detection on the multi-scale fusion feature map by using a detection network module to obtain a scene text detection result. By means of the method, the scene text detection effect can be improved, and meanwhile the scene text detection real-time performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a scene text detection method and device. Background Art

[0002] Scene text detection is an important task in computer vision, aiming to locate text regions in natural images. Currently, scene text detection is widely used in office automation, intelligent navigation, image retrieval, and instant translation. However, scene text detection still faces many challenges in practical applications, such as complex natural environments, harsh imaging conditions, and variable text sizes and directions.

[0003] Traditional scene text detection methods can be divided into sliding window-based detection methods and connected domain-based detection methods. These detection methods mainly build detection models through manually designed text visual features and text prior information, and their detection accuracy and speed are difficult to meet the needs of actual applications.

[0004] Compared with traditional detection methods, scene text detection methods based on deep learning can extract deeper text instance features, so the detection performance is better. Text detection methods based on deep learning can be divided into detection methods based on region regression and detection methods based on pixel segmentation. The detection method based on region regression mainly uses a general target detection framework to improve the traditional sliding window method according to the characteristics of text. The idea of ​​this type of method is to first extract a large number of text candidate boxes in the image, and then obtain the text detection results through bounding box regression of the text area. The detection method based on pixel segmentation mainly improves the traditional connected domain method based on the general semantic segmentation framework, combining pixel-level semantic segmentation and corresponding post-processing algorithms to obtain the text bounding box.

[0005] In related technologies, scene text detection methods still have problems such as weak multi-scale feature representation capabilities and insufficient cross-scale feature fusion, which leads to poor small text detection effects. In addition, the model inference time is long, making it difficult to achieve a balance between detection performance and efficiency. Summary of the invention

[0006] The present invention provides a scene text detection method and device, which can solve the problem of poor scene text detection effect and time-consuming in the related art. The technical solution is as follows:

[0007] In a first aspect, a scene text detection method is provided, comprising: using a multi-scale backbone network module to perform multi-scale feature extraction on an image to be detected to obtain a multi-scale feature map; constructing an initial feature pyramid based on the multi-scale feature map; using a feature pyramid enhancement network module to perform two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map, wherein the two stages include a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation; using a dual attention module to perform cross-scale feature fusion on the enhanced feature map to obtain a multi-scale fused feature map, wherein the dual attention module includes a channel attention module and a spatial attention module; using a detection network module to perform scene text detection on the multi-scale fused feature map to obtain a scene text detection result.

[0008] In some embodiments, constructing an initial feature pyramid based on the multi-scale feature map includes: performing a scale-invariant adaptive average pooling operation on the feature map of the smallest scale in the multi-scale feature map to obtain context feature maps of multiple scales; performing adaptive spatial fusion on the context feature maps of the multiple scales to obtain a fused context feature map; and constructing the initial feature pyramid based on the fused context feature map and the multi-scale feature map.

[0009] In some embodiments, the adaptive spatial fusion of the context feature maps of the multiple scales to obtain a fused context feature map includes: performing 1×1 convolution and upsampling processing on the context feature maps of the multiple scales to obtain multiple context feature maps of the same scale; performing channel cascading operations on the multiple context feature maps of the same scale to obtain a cascaded feature map; and using a spatial attention module to perform weighted summation of spatial features in the cascaded feature map to obtain the fused context feature map.

[0010] In some embodiments, constructing the initial feature pyramid based on the fused context feature map and the multi-scale feature map includes: using the fused context feature map as the top-level feature map of the initial feature pyramid; and using the multi-scale feature map as the feature map structure located below the top-level feature map in the initial feature pyramid.

[0011] In some embodiments, the use of the feature pyramid enhancement network module to perform two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map includes: in the top-down feature enhancement stage, upsampling each layer of the feature map in the feature pyramid, and then adding and summing the next layer of feature maps adjacent to the feature map of this layer to obtain a fused feature map; performing grouped shuffled convolution, batch normalization, and activation processing on the fused feature map to obtain the first-stage enhanced feature pyramid; in the bottom-up feature enhancement stage, adding the upsampling results of each layer of the feature map in the first-stage enhanced feature pyramid and the previous layer of feature map adjacent to the feature map of this layer to obtain a fused feature map; performing grouped shuffled convolution, batch normalization, and activation processing on the fused features to obtain the second-stage enhanced feature pyramid; and performing element-by-element addition on the second-stage enhanced feature pyramid and the initial feature pyramid to obtain the enhanced feature map.

[0012] In some embodiments, the multi-scale backbone network module is a deep residual network module.

[0013] In a second aspect, a scene text detection device is provided, comprising: an extraction module, used to use a multi-scale backbone network module to perform multi-scale feature extraction on a to-be-detected image to obtain a multi-scale feature map; a construction module, used to construct an initial feature pyramid based on the multi-scale feature map; an enhancement module, used to use a feature pyramid enhancement network module to perform two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map, wherein the two stages include a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation; a fusion module, used to use a dual attention module to perform cross-scale feature fusion on the enhanced feature map to obtain a multi-scale fused feature map, wherein the dual attention module includes a channel attention module and a spatial attention module; a detection module, used to use a detection network module to perform scene text detection on the multi-scale fused feature map to obtain a scene text detection result.

[0014] In a third aspect, a computer device is provided, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of the scene text detection method as described above.

[0015] In a fourth aspect, a computer-readable storage medium is provided, characterized in that a computer program is stored in the storage medium, and when the computer program is executed by a processor, the steps of the scene text detection method as described above are implemented.

[0016] In a fifth aspect, a computer program product is provided, characterized in that it includes a computer program, and when the computer program is executed by a processor, the steps of the scene text detection method as described above are implemented.

[0017] The technical solution provided by the present invention can at least bring the following beneficial effects: on the one hand, by combining the technical means of performing multi-scale feature extraction on the image to be detected, performing two-stage feature enhancement processing on the initial feature pyramid using the feature pyramid enhancement network module, and performing cross-scale feature fusion on the enhanced feature map using the dual attention module, the multi-scale feature representation capability can be improved, the adequacy of cross-scale feature fusion can be enhanced, and the performance of the scene text detection model can be improved; on the other hand, by using lightweight group shuffle convolution in the feature pyramid enhancement network module, the computational consumption is reduced and the real-time performance of scene text detection is improved. The combination of the above two aspects can take into account both the efficiency and detection accuracy of scene text detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 is a flow chart of a scene text detection method provided by an embodiment of the present invention;

[0020] Figure 2 It is a schematic flow chart of the steps of constructing an initial pyramid provided by an embodiment of the present invention;

[0021] Figure 3 It is a flow chart of a feature enhancement step provided by an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of a scene text detection model provided by an embodiment of the present invention;

[0023] Figure 5 is a schematic diagram of a feature pyramid enhancement network module provided by an embodiment of the present invention;

[0024] Figure 6 is a schematic diagram of a dual attention module provided by an embodiment of the present invention;

[0025] Figure 7 is a schematic diagram of a channel attention module provided by an embodiment of the present invention;

[0026] Figure 8is a schematic diagram of a spatial attention module provided by an embodiment of the present invention;

[0027] Fig. 9 is a structural schematic diagram of a scene text detection device provided by an embodiment of the present invention;

[0028] Fig.10 It is a hardware architecture diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0030] In order to solve the problems existing in the related art, the present invention provides a scene text detection method and device, which can improve the accuracy and real-time performance of scene text detection.

[0031] The following is a detailed description with reference to specific embodiments.

[0032] Please refer to Figure 1 , the scene text detection method provided by the embodiment of the present invention includes steps S11 to S15.

[0033] Step S11, using a multi-scale backbone network module, performs multi-scale feature extraction on the image to be detected to obtain a multi-scale feature map.

[0034] The multi-scale feature map includes feature maps of multiple different scales. For example, the multi-scale feature map includes feature maps of four different scales.

[0035] In some examples, the multi-scale backbone network module is a deep residual network module. For example, the multi-scale backbone network module is a Res2Net network. By using the Res2Net network for multi-scale feature extraction, the model's ability to represent multi-scale features can be enhanced, and the model's ability to detect text boxes of different scales can be improved.

[0036] Step S12: constructing an initial feature pyramid based on the multi-scale feature map.

[0037] In some examples, in step S12, based on Figure 2 The process shown below constructs the initial feature pyramid. Figure 2As shown, step S12 includes: step S121, performing a scale-invariant adaptive average pooling operation on the feature map of the smallest scale in the multi-scale feature map (i.e., the top-level feature map) to obtain context feature maps of multiple scales; step S122, performing adaptive spatial fusion on the context feature maps of multiple scales to obtain a fused context feature map; step S123, constructing an initial feature pyramid based on the fused context feature map and the multi-scale feature map.

[0038] For example, in step S121, the feature map with the smallest scale in the multi-scale feature map obtained in step S11 is used as the top-level feature map of the pyramid; a scale-invariant adaptive average pooling operation is performed on the top-level feature map of the pyramid to obtain multiple context feature maps of different scales. For example, four context feature maps of different scales are obtained.

[0039] For example, in step S122, 1*1 convolution and upsampling are performed on context feature maps of multiple scales to unify these context feature maps to the same scale; then, channel cascading operation is performed on the context feature maps unified to the same scale to obtain a cascaded feature map; next, the spatial features in the cascaded feature map are weighted summed using the spatial attention module to obtain a fused context feature map.

[0040] Specifically, in step S122, when the cascaded feature map is processed using the spatial attention module, the cascaded feature map can first be subjected to spatial attention processing and dimensional expansion to obtain weight features, and then the weight features are multiplied and summed with the feature points of the corresponding scale to obtain a fused context feature map.

[0041] For example, in step S123, the fused context features are used as the top-level feature map of the initial feature pyramid; and the multi-scale feature map obtained in step S11 is used as the feature map structure located below the top-level feature map in the initial feature pyramid.

[0042] In the disclosed embodiment, by performing adaptive pooling operations and adaptive spatial fusion operations on the top-level feature map of the pyramid to obtain a fused contextual feature map, and constructing an initial feature pyramid based on this, it is possible to compensate for the semantic information loss of the top-level features of the feature pyramid, enhance the representation capability of deep contextual features, and further enhance the model's representation capability for multi-scale features, thereby helping to improve the accuracy of scene text detection.

[0043] Step S13, using the feature pyramid enhancement network module, performs two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map.

[0044] The two-stage feature enhancement process includes a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation.

[0045] In some examples, in step S13, based on Figure 3 The process shown in the figure performs feature enhancement processing. Figure 3 As shown, step S13 includes: S131, a top-down feature enhancement stage; S132, a bottom-up feature enhancement stage; S133, fusing the feature pyramid enhanced in the second stage and the initial feature pyramid to obtain an enhanced feature map.

[0046] For example, in step S131, each layer of feature maps in the initial feature pyramid obtained by step S12 is upsampled, and then the next layer of feature maps adjacent to the feature map of this layer are added and summed to obtain a fused feature map; the fused feature map is subjected to group shuffle convolution, batch normalization, and activation processing to obtain a feature pyramid after the first stage enhancement.

[0047] For example, in step S132, the upsampling results of each layer of feature graphs in the first-stage enhanced feature pyramid and the upper layer of feature graphs adjacent to the layer of feature graphs are added to obtain a fused feature graph; the fused features are subjected to group shuffle convolution, batch normalization, and activation processing to obtain a second-stage enhanced feature pyramid. For example, in step S133, the second-stage enhanced feature pyramid and the initial feature pyramid are element-by-element added to obtain an enhanced feature graph.

[0048] In the embodiments disclosed herein, on the one hand, by introducing not only a top-down feature enhancement stage but also a bottom-up feature enhancement stage, the positioning information of the high-level features of the pyramid is enriched, the information transmission path from the low-level to the high-level is shortened, and the low-level information is easier to be transmitted to the high-level of the pyramid, which helps to improve the accuracy of scene text detection; on the other hand, by adopting low-computational group shuffle convolution in the two-stage feature enhancement stage, the speed of scene text detection can be improved while improving the accuracy of scene text detection.

[0049] Step S14, using the dual attention module, performs cross-scale feature fusion on the enhanced feature map to obtain a multi-scale fused feature map.

[0050] Among them, the dual attention module includes a channel attention module and a spatial attention module.

[0051] For example, in step S14, the enhanced feature pyramid is taken as input, and the input is first processed based on the channel attention mechanism to obtain a channel weighted feature map; then the channel weighted feature map is processed based on the spatial attention mechanism to obtain a global feature map; then the global feature map is subjected to 1×1 convolution processing and Sigmoid activation operation to obtain weight matrices corresponding to feature maps of different scales; then each weight matrix is ​​subjected to a dot multiplication operation with the feature map of the corresponding scale in the enhanced feature pyramid to obtain a multi-scale fusion feature map.

[0052] In the disclosed embodiment, by learning the weights of the input feature map on different channels based on the channel attention module, the channels containing text instance-related features are highlighted, and the channels containing a lot of noise and background information are suppressed; further, by learning the weights of the feature map at different positions based on the spatial attention module, the feature information related to the target is emphasized, and the useless background and noise features are suppressed. Therefore, through the dual attention module, a weight matrix containing a lot of effective weight information can be learned, and a more robust multi-scale fusion feature map can be obtained, thereby improving the accuracy of multi-scale text detection.

[0053] Step S15, using the detection network module, performing scene text detection on the multi-scale fusion feature map to obtain a scene text detection result.

[0054] In some examples, in step S15, the multi-scale fusion feature map is processed using a detection network module to obtain a probability map and a threshold map for scene text detection; then an approximate binary map is obtained based on the probability map and the text map; next, a scene text detection result is obtained based on the binary map.

[0055] Specifically, after obtaining the probability map, a binarization operation is required to divide the text core area pixels and background pixels in the image. The standard binarization operation is to set a fixed threshold and divide each pixel according to the fixed threshold. The standard binarization calculation process is shown in formula (1).

[0056]

[0057] Among them, (i, j) represents the pixel position in the image, t is the set threshold, P i,j Indicates the probability that a pixel is text.

[0058] It can be seen from formula (1) that the standard binarization calculation is not differentiable and cannot be jointly trained and optimized with other modules of the model. In order to add the threshold as a learnable parameter to the optimization iteration process of the model, the differentiable binarization operation can be implemented by approximating the step function so that each pixel obtains an adaptive threshold. The differentiable binarization calculation is shown in formula (2):

[0059]

[0060] in, P i,j , T i,j The values ​​of the approximate binary map, probability map, and threshold map of the pixel at the position (i, j) are represented in sequence, and k is the magnification factor of the approximate step function. For example, k is 50.

[0061] In the disclosed embodiment, the above process can improve the scene text detection effect while improving the real-time performance of the scene text detection.

[0062] The following combination Figures 4 to 8 The scene text detection method is further explained. Figure 4 As shown, the scene text detection model includes a backbone network module, a feature pyramid enhancement network module, a dual attention module and a detection network module.

[0063] Among them, the backbone network module can specifically adopt a deep residual network, such as a Res2Net-50 network. Each residual block of the deep residual network sets a residual connection of the layer grouping to perform multi-scale decoupling of the original single 3×3 convolution operation. Specifically, feature extraction of the image to be detected based on the Res2Net-50 network includes: based on the Res2Net-50 network, the image to be detected is processed in multiple convolution stages to generate four feature maps of different scales. Among them, each convolution stage can use a 3×3 convolution operation. In specific implementation, the input of the Res2Net-50 network can be grouped to obtain 4 groups of input sub-feature maps. Among them, the first group of input sub-feature maps is subjected to a 3×3 convolution to obtain an output sub-feature map, and each subsequent group of input sub-feature maps is added to the previous group of output feature maps and then subjected to a 3×3 convolution operation to obtain the output sub-feature map of the group. It can be seen that each group of output sub-feature maps has a larger receptive field than the previous group of output sub-feature maps. Therefore, by processing the image to be detected through the Res2Net network, features containing receptive field combinations of various sizes and numbers are obtained, rich global and local features are extracted, and the ability to represent multi-scale features is enhanced, which helps to improve the accuracy of subsequent text detection at different scales.

[0064] After obtaining multi-scale feature maps based on the Res2Net-50 network (for example, feature maps C2, C3, C4, and C5, whose scales are 160*160, 80*80, 40*40, and 20*20, respectively), we can Figure 5 The modules shown perform feature enhancement processing. Figure 5As shown in the figure, after obtaining the feature maps C2 to C5, the feature map C5 of the smallest scale in the multi-scale feature map can be first subjected to a scale-invariant adaptive average pooling operation to obtain context feature maps of four scales; the context feature maps of the four scales are adaptively spatially fused to obtain a fused context feature map T6; T6 and the multi-scale feature maps C2 to C5 are used as the initial feature pyramid.

[0065] Next, use Figure 5 The feature enhancement pyramid network module shown performs two-stage feature enhancement processing on the initial feature pyramid to obtain enhanced feature maps P2 to P5. Specifically, in the downward enhancement stage, each deep feature map in the initial feature pyramid is first upsampled, and then added and summed with the feature map of the next layer. The feature map obtained by the sum is processed by the grouped shuffled convolution layer, batch normalization layer, and ReLU activation layer to obtain an enhanced shallow feature map; in the upward enhancement stage, each shallow feature map in the feature map obtained in the downward enhancement stage is added and summed with the upsampling result of the feature map of the previous layer. The feature map obtained by the sum is processed by the grouped shuffled convolution, batch normalization, and activation operation to obtain an enhanced deep feature map; finally, the pyramid feature map output in the upward enhancement stage is added element by element with the pyramid feature map before bidirectional enhancement to obtain the final enhanced feature maps P2, P3, P4, and P5.

[0066] After obtaining the enhanced feature maps P2, P3, P4 and P5, based on Figure 6 The dual attention module shown in FIG. performs cross-scale feature fusion on the enhanced feature map to obtain a multi-scale fused feature map F. Among them, the channel attention module in the dual attention module can be used Figure 7 The structure shown in the figure, the spatial attention module in the dual attention module can be adopted Figure 8 The structure shown.

[0067] After obtaining the multi-scale fusion feature map F, use Figure 4 The detection network module shown performs text detection on the multi-scale fusion feature map F to obtain a probability map and a threshold map; obtains an approximate binary map based on the probability map and the threshold map; and then obtains the scene text detection result of the image to be detected based on the approximate binary map.

[0068] In the present invention, the following technical effects can be achieved through the above scene text detection method: 1) by introducing hierarchical residual connections in each residual block in the deep residual network during multi-scale feature extraction, the network receptive field is increased, and the network's feature extraction capability for texts of various scales is enhanced; 2) by compensating for the semantic information loss of the top-level features of the pyramid through a scale-invariant adaptive average pooling operation when constructing the initial feature pyramid, by introducing two-stage feature enhancement during feature enhancement, especially introducing a bottom-up feature enhancement stage, it is helpful to enhance the positioning information of the high-level features of the pyramid, and the above two aspects are combined to help improve the accuracy of subsequent scene text detection; and, by introducing lightweight grouped shuffled convolutions in the two-stage feature enhancement stage, the computational consumption is reduced, so that while improving the accuracy of scene text detection, real-time detection of complex scene texts is achieved, and the detection efficiency and detection performance of the model are balanced; 3) by processing using a channel and space dual attention mechanism in the feature fusion stage, a more robust cross-scale fusion feature map can be obtained, which helps to further improve the accuracy of subsequent scene text detection.

[0069] In the present invention, a training method for a scene text detection model is also provided. The scene text detection model includes a multi-scale backbone network module, a feature pyramid enhancement network module, a dual attention module, and a detection network module. The training method of the scene text detection model includes: using the multi-scale backbone network module to perform multi-scale feature extraction on the training image to obtain a multi-scale feature map; based on the multi-scale feature map, constructing an initial feature pyramid; using the feature pyramid enhancement network module, performing two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map, the two stages include a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation; using the dual attention module, cross-scale feature fusion is performed on the enhanced feature map to obtain a multi-scale fused feature map; using the detection network module, the multi-scale fused feature map is subjected to scene text prediction to obtain a scene text prediction result; based on the scene text prediction result, the scene text detection model is trained to obtain a trained scene text detection model. Through the above model training method, the performance of the scene text detection model can be improved while the convergence speed of the scene text detection model can be improved.

[0070] like Fig. 9 As shown, the present invention also provides a scene text detection device for executing the scene text detection method as described above, and the device specifically includes an extraction module 91, a construction module 92, an enhancement module 93, a fusion module 94, and a detection module 95.

[0071] The extraction module 91 is used to use the multi-scale backbone network module to perform multi-scale feature extraction on the image to be detected to obtain a multi-scale feature map.

[0072] The construction module 92 is used to construct an initial feature pyramid based on the multi-scale feature map.

[0073] The enhancement module 93 is used to utilize the feature pyramid enhancement network module to perform two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map, wherein the two stages include a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation.

[0074] The fusion module 94 is used to perform cross-scale feature fusion on the enhanced feature map using a dual attention module to obtain a multi-scale fused feature map. The dual attention module includes a channel attention module and a spatial attention module.

[0075] The detection module 95 is used to perform scene text detection on the multi-scale fusion feature map using the detection network module to obtain a scene text detection result.

[0076] In an embodiment of the present invention, the above device can improve the scene text detection effect while improving the real-time performance of the scene text detection. It should be noted that the scene text detection device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the scene text detection device provided in the above embodiment belongs to the same concept as the scene text detection method embodiment. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0077] The embodiment of the present application also provides a computer device, please refer to Fig.10 The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the scene text detection method provided by the above-mentioned method embodiments.

[0078] An embodiment of the present application also provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the scene text detection method provided by the above-mentioned method embodiments.

[0079] An embodiment of the present application further provides a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes any of the scene text detection methods described in the above embodiments.

[0080] For the convenience of description, the above system or device is described by dividing it into various modules or units according to its functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0081] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0082] Finally, it should be noted that, in this article, relational terms such as first, second, third and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0083] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A scene text detection method, characterized in that: The method comprises: Using the multi-scale backbone network module, multi-scale feature extraction is performed on the image to be detected to obtain a multi-scale feature map; Based on the multi-scale feature map, construct an initial feature pyramid; Using a feature pyramid enhancement network module, the initial feature pyramid is subjected to a two-stage feature enhancement process to obtain an enhanced feature map, wherein the two stages include a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation; Using a dual attention module, cross-scale feature fusion is performed on the enhanced feature map to obtain a multi-scale fused feature map, wherein the dual attention module includes a channel attention module and a spatial attention module; The detection network module is used to perform scene text detection on the multi-scale fusion feature map to obtain a scene text detection result.

2. The scene text detection method according to claim 1, characterized in that: The constructing an initial feature pyramid based on the multi-scale feature map includes: Performing a scale-invariant adaptive average pooling operation on the feature map of the smallest scale in the multi-scale feature map to obtain context feature maps of multiple scales; Adaptively spatially fusing the context feature maps of the multiple scales to obtain a fused context feature map; The initial feature pyramid is constructed according to the fused context feature map and the multi-scale feature map.

3. The scene text detection method according to claim 2, characterized in that: The adaptively spatially fusing the context feature maps of the multiple scales to obtain a fused context feature map comprises: Performing 1×1 convolution processing and upsampling processing on the context feature maps of the multiple scales to obtain multiple context feature maps of the same scale; Performing a channel cascade operation on the multiple context feature maps of the same scale to obtain a cascaded feature map; The spatial features in the cascaded feature map are weightedly summed using a spatial attention module to obtain the fused context feature map.

4. The scene text detection method according to claim 2, characterized in that: The constructing the initial feature pyramid according to the fused context feature map and the multi-scale feature map comprises: The fused context feature map is used as the top-level feature map of the initial feature pyramid; and the multi-scale feature map is used as the feature map structure located below the top-level feature map in the initial feature pyramid.

5. The scene text detection method according to claim 1, characterized in that: The method of using the feature pyramid enhancement network module to perform two-stage feature enhancement processing on the initial feature pyramid to obtain an enhanced feature map includes: In the top-down feature enhancement stage, each layer of feature maps in the feature pyramid is upsampled, and then the next layer of feature maps adjacent to the layer of feature maps are added and summed to obtain a fused feature map; the fused feature map is subjected to group shuffle convolution, batch normalization, and activation processing to obtain a feature pyramid enhanced in the first stage; In the bottom-up feature enhancement stage, the upsampling results of each layer of feature maps in the feature pyramid enhanced in the first stage and the feature maps of the previous layer adjacent to the feature maps of the layer are added to obtain a fused feature map; the fused features are subjected to group shuffle convolution, batch normalization, and activation processing to obtain a feature pyramid enhanced in the second stage; The enhanced feature pyramid in the second stage and the initial feature pyramid are added element by element to obtain the enhanced feature map.

6. The scene text detection method according to any one of claims 1 to 5, characterized in that: The multi-scale backbone network module is a deep residual network module.

7. A scene text detection device, characterized in that: The device comprises: An extraction module is used to extract multi-scale features from the image to be detected using a multi-scale backbone network module to obtain a multi-scale feature map; A construction module, used to construct an initial feature pyramid based on the multi-scale feature map; An enhancement module, used for performing two-stage feature enhancement processing on the initial feature pyramid using a feature pyramid enhancement network module to obtain an enhanced feature map, wherein the two stages include a top-down feature enhancement stage and a bottom-up feature enhancement stage, and each stage includes a group shuffle convolution operation; A fusion module, used to perform cross-scale feature fusion on the enhanced feature map using a dual attention module to obtain a multi-scale fused feature map, wherein the dual attention module includes a channel attention module and a spatial attention module; The detection module is used to perform scene text detection on the multi-scale fusion feature map using the detection network module to obtain a scene text detection result.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of the scene text detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the scene text detection method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The method comprises a computer program, which, when executed by a processor, implements the steps of the scene text detection method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Coal ash fusibility determination method based on residual attention and feature pyramid

    CN120182734A

  • Electric power communication cross-medium data identification method and device and storage medium

    CN122200675A

  • Power communication cross-medium data identification method and device and storage medium

    CN122200675B