Text content position prediction method, device, electronic device and storage medium
By extracting the semantic features of the text content area in the image and performing bar pooling and feature weighting processing, the problem of incomplete detection of curved and long text content in the prior art is solved, and the accuracy of text content position prediction is improved.
Patent Information
- Application Number
- CN202111621546.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Existing text content position prediction methods cannot effectively detect curved text content, and the detection of long text content is not complete enough.
By acquiring the image to be processed, the first semantic features of the text content area are extracted, and the second semantic features are extracted from the first semantic features and the image. Then, the second semantic feature is bar pooled and feature weighted to obtain the location information of the text content area.
Improve the accuracy of text content position prediction and enable more efficient detection of the position of curved and long text content.
Smart Images

Figure CN114417108B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technology, and in particular to a method, device, electronic device and computer-readable storage medium for predicting the location of text content. Background Art
[0002] The position prediction of text content in an image is of great significance for understanding the image and is the algorithmic cornerstone of image analysis. Therefore, building a more robust method for predicting the position of text content is very important for analyzing images.
[0003] At present, text content location prediction methods are mainly divided into two categories, namely regression-based methods and segmentation-based methods.
[0004] (1) Regression-based method. The model outputs a rotatable text region candidate box, and finds the tilt angle of the text region candidate box to be regressed during the bounding box regression calculation process. For example, the text region candidate box output by the Rotation Region Proposal Network (RRPN) model is represented as a 5-tuple, namely x, y, h, w and θ. Among them, x represents the horizontal coordinate of a vertex of the text region candidate box, y represents the vertical coordinate of the vertex of the text region candidate box, h represents the height of the text region candidate box, w represents the width of the text region candidate box, and θ represents the tilt angle of the text region candidate box. Figure 1 , Figure 1 Figure 1 shows a schematic diagram of the output results of the RRPN model. Figure 1 There are two text region candidate boxes, one of which contains the text content "CENTRAL PARK" and the other contains the text content "WEST". However, the regression-based method cannot detect curved text content and is not easy to fully detect long text content.
[0005] (2) Semantic segmentation-based methods. Semantic segmentation-based methods use the text content location prediction as an image segmentation task and use a segmentation-based framework for detection. The specific approach is to treat each text content as a mask area, and then predict the location of the mask area, and then obtain the final text content location through post-processing. Figure 2a and Figure 2b , Figure 2a A schematic diagram showing the original image, Figure 2bThe schematic diagram of the mask area is shown. The advantage of the semantic segmentation method is that it can fit text content of any shape, such as detecting curved text content, but the segmentation result will affect the final mask area. However, the position of the detected text content is greatly affected by the segmentation result based on the semantic segmentation method. Summary of the invention
[0006] In view of the above problems, embodiments of the present invention are proposed to provide a method, device, electronic device and computer-readable storage medium for predicting the position of text content that overcome the above problems or at least partially solve the above problems.
[0007] In order to solve the above problem, according to a first aspect of an embodiment of the present invention, a method for predicting the position of text content is disclosed, the method comprising: obtaining an image to be processed, the image containing at least one strip-shaped text content area; extracting a first semantic feature of at least one of the text content areas from the image, and extracting a second semantic feature of at least one of the text content areas from the image based on the first semantic feature and the image; performing strip pooling processing and feature weighting processing on the second semantic feature respectively to obtain the position information of at least one of the text content areas.
[0008] Optionally, extracting a first semantic feature of at least one text content area from the image includes: performing multiple convolution processes from the image in a bottom-up direction to obtain the first semantic feature; wherein the input item of each convolution process is the output item of the previous convolution process.
[0009] Optionally, extracting a second semantic feature of at least one text content area from the image based on the first semantic feature and the image includes: performing convolution processing on the first semantic feature to obtain a top-level convolution processing result; performing multiple deconvolution processing in a top-down direction starting from the top-level convolution processing result to obtain the second semantic feature; wherein the input item of each deconvolution processing is the output item of the previous deconvolution processing and the corresponding output item of each convolution processing.
[0010] Optionally, performing strip pooling processing on the second semantic feature includes: performing horizontal pooling processing and vertical pooling processing on the second semantic feature respectively to obtain a pooling processing result.
[0011] Optionally, the performing feature weighted processing on the second semantic feature includes: performing spatial attention feature weighted processing on the second semantic feature to obtain a weighted processing result.
[0012] Optionally, performing strip pooling processing and feature weighting processing on the second semantic features respectively to obtain position information of at least one text content area includes: fusing the pooling processing result and the weighting processing result to obtain the position information.
[0013] According to a second aspect of an embodiment of the present invention, a device for predicting the position of text content is also disclosed, the device comprising: an image acquisition module, used to acquire an image to be processed, the image containing at least one strip-shaped text content area; a recursive module, used to extract a first semantic feature of at least one text content area from the image, and extract a second semantic feature of at least one text content area from the image based on the first semantic feature and the image; an attention module, used to perform strip pooling processing and feature weighting processing on the second semantic feature, respectively, to obtain the position information of at least one text content area.
[0014] Optionally, the recursive module includes: a convolution processing module, used to perform multiple convolution processes from the image in a bottom-up direction to obtain the first semantic feature; wherein the input item of each convolution process is the output item of the previous convolution process.
[0015] Optionally, the convolution processing module is also used to convolve the first semantic feature to obtain a top-level convolution processing result; the recursive module also includes: a deconvolution processing module, used to perform multiple deconvolution processes in a top-down direction starting from the top-level convolution processing result to obtain the second semantic feature; wherein the input item of each deconvolution processing is the output item of the previous deconvolution processing and the corresponding output item of each convolution processing.
[0016] Optionally, the attention module includes: a pooling processing module, used to perform horizontal pooling processing and vertical pooling processing on the second semantic features to obtain pooling processing results.
[0017] Optionally, the attention module further includes: a weighted processing module, configured to perform spatial attention feature weighted processing on the second semantic feature to obtain a weighted processing result.
[0018] Optionally, the attention module further includes: a result fusion module, used to fuse the pooling processing result and the weighted processing result to obtain the position information.
[0019] According to a third aspect of an embodiment of the present invention, an electronic device is also disclosed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for predicting the position of text content described in the first aspect when executing the computer program.
[0020] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is further disclosed, on which a computer program is stored, and when the program is executed by a processor, the method for predicting the position of text content described in the first aspect is implemented.
[0021] Compared with the prior art, the technical solution provided by the embodiment of the present invention has the following advantages:
[0022] The embodiment of the present invention provides a text content location prediction scheme, which obtains an image to be processed, and the image to be processed may contain at least one strip-shaped text content area. A first semantic feature of at least one text content area is extracted from the image, and a second semantic feature of at least one text content area is extracted from the image based on the first semantic feature and the image. Then, the second semantic feature is subjected to strip pooling processing and feature weighting processing respectively to obtain the location information of at least one text content area. After extracting the first semantic feature, the embodiment of the present invention performs a secondary extraction on the first semantic feature to obtain the second semantic feature, which can obtain more refined edge information, improve the semantic segmentation result, and thus improve the accuracy of the location prediction of the text content. Moreover, strip pooling processing is performed on the second semantic feature to capture the long dependency relationship of the spatial region and prevent interference from irrelevant regions. Feature weighting processing is performed on the second semantic feature to enhance the features of the text content, which can simultaneously aggregate global and local context information and enhance the location prediction capability of long text content. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the output results of the RRPN model;
[0024] Figure 2a is a schematic diagram of the original image;
[0025] Figure 2b It is a schematic diagram of the mask area;
[0026] Figure 3 is a flowchart of the steps of a method for predicting the position of text content according to an embodiment of the present invention;
[0027] Figure 4 is a schematic diagram of a process of extracting a first semantic feature according to an embodiment of the present invention;
[0028] Figure 5 is a schematic diagram of a process of extracting a second semantic feature according to an embodiment of the present invention;
[0029] Figure 6 is a flow chart of a strip pooling process according to an embodiment of the present invention;
[0030] Figure 7 is a flow chart of a feature weighting process according to an embodiment of the present invention;
[0031] Figure 8 is a structural block diagram of a device for predicting the position of text content according to an embodiment of the present invention;
[0032] Fig. 9 It is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] Reference Figure 3 , shows a flowchart of a method for predicting the position of text content according to an embodiment of the present invention. The method for predicting the position of text content may specifically include the following steps:
[0035] Step 301: Obtain an image to be processed.
[0036] In an embodiment of the present invention, an image may include at least one strip-shaped text content area. In practical applications, the strip-shaped text content area may be a rectangular text content area, a trapezoidal text content area, or a fan-shaped text content area, etc., and the embodiment of the present invention does not specifically limit the specific form of the strip. Moreover, the text content may be words, letters, symbols, etc., and the embodiment of the present invention does not specifically limit the type, quantity, color, etc. of the text content.
[0037] Step 302: extract a first semantic feature of at least one text content region from the image, and extract a second semantic feature of at least one text content region from the image based on the first semantic feature and the image.
[0038] In an embodiment of the present invention, a secondary feature extraction is performed on the text content area in the image. Specifically, a secondary feature extraction can be performed on the text content area in the image based on a Feature Pyramid Networks (FPN) model. The first feature extraction obtains the first semantic feature, and the second feature extraction obtains the second semantic feature.
[0039] Step 303 , performing strip pooling processing and feature weighting processing on the second semantic feature respectively to obtain position information of at least one text content area.
[0040] In an embodiment of the present invention, strip pooling processing and feature weighting processing are performed on the second semantic features based on the long-short range attention mechanism (LSRA) to obtain the position information of the text content area.
[0041] The embodiment of the present invention provides a text content location prediction scheme, which obtains an image to be processed, and the image to be processed may contain at least one strip-shaped text content area. A first semantic feature of at least one text content area is extracted from the image, and a second semantic feature of at least one text content area is extracted from the image based on the first semantic feature and the image. Then, the second semantic feature is subjected to strip pooling processing and feature weighting processing respectively to obtain the location information of at least one text content area. After extracting the first semantic feature, the embodiment of the present invention performs a secondary extraction on the first semantic feature to obtain the second semantic feature, which can obtain more refined edge information, improve the semantic segmentation result, and thus improve the accuracy of the location prediction of the text content. Moreover, strip pooling processing is performed on the second semantic feature to capture the long dependency relationship of the spatial region and prevent interference from irrelevant regions. Feature weighting processing is performed on the second semantic feature to enhance the features of the text content, which can simultaneously aggregate global and local context information and enhance the location prediction capability of long text content.
[0042] In a preferred embodiment of the present invention, an implementation method of extracting the first semantic feature of at least one text content area from an image is to perform multiple convolution processes from the image in a bottom-up direction to obtain the first semantic feature. The input item of each convolution process is the output item of the previous convolution process. Figure 4 , Figure 4 A schematic diagram of the process of extracting the first semantic feature is shown. Figure 4 In the figure, the bottom layer image is the image to be processed, and the arrow direction indicates the direction from bottom to top. Figure 4 It can be seen that three convolution processes are performed in the process of extracting the first semantic feature. That is, the image is subjected to the first convolution process to obtain the first convolution process result, the first convolution process result is subjected to the second convolution process to obtain the second convolution process result, and the second convolution process result is subjected to the third convolution process to obtain the third convolution process result. The third convolution process result is the first semantic feature. It should be noted that when the image contains multiple text content areas, the first semantic feature is the semantic feature of multiple text content areas.
[0043] In a preferred embodiment of the present invention, an implementation method of extracting a second semantic feature of at least one text content area from an image based on the first semantic feature and the image is to perform convolution processing on the first semantic feature to obtain a top-level convolution processing result. Starting from the top-level convolution processing result, multiple deconvolution processing is performed in a top-down direction to obtain the second semantic feature. The input item of each deconvolution processing is the output item of the previous deconvolution processing and the corresponding output item of each convolution processing. Figure 5 , Figure 5A schematic diagram of the process of extracting the second semantic feature is shown. Figure 5 In the above method, after extracting the first semantic feature, the first semantic feature is subjected to horizontal convolution processing to obtain the top convolution processing result. Then, the top convolution processing result and the second convolution processing result are subjected to a first deconvolution processing to obtain the first deconvolution processing result. Then, the first deconvolution processing result and the first convolution processing result are subjected to a second deconvolution processing to obtain the second deconvolution processing result. The second deconvolution processing result is the second semantic feature.
[0044] It should be noted that the number of convolution operations in the first semantic feature extraction process is one more than the number of deconvolution operations in the second semantic feature extraction process. That is, three convolution operations are performed in the process of extracting the first semantic feature, and two deconvolution operations are performed in the process of extracting the second semantic feature.
[0045] The first semantic feature contains rich semantic information, but due to its relatively low resolution, it is not easy to accurately save the location information of the text content. Although the second semantic feature has relatively less semantic information, it has a relatively high resolution and can save the location information of the text content more accurately. Therefore, the first semantic feature and the second semantic feature are merged to accurately locate the location of the text content.
[0046] In a preferred embodiment of the present invention, an implementation method of performing strip pooling processing on the second semantic feature is to perform horizontal pooling processing and vertical pooling processing on the second semantic feature to obtain a pooling processing result. Figure 6 , Figure 6 A flow chart of strip pooling processing is shown. Figure 6 In , the second semantic feature (Feature map) is processed by horizontal pooling (Horizontal pooling) and vertical pooling (Vertical pooling) to obtain the pooling result.
[0047] In a preferred embodiment of the present invention, an implementation method of performing feature weighting processing on the second semantic feature is to perform spatial attention feature weighting processing on the second semantic feature to obtain a weighted processing result. Figure 7 , Figure 7 A flow chart of feature weighting processing is shown. Figure 7 In the above, the second semantic feature (Feature map) is weighted by the spatial attention feature to obtain the weighted processing result.
[0048] In a preferred embodiment of the present invention, strip pooling processing and feature weighting processing are performed on the second semantic feature respectively, and an implementation method for obtaining the position information of at least one text content area is to fuse the pooling processing result and the weighted processing result to obtain the position information.
[0049] Based on the above description of an embodiment of a method for predicting the position of text content, a recursive text line detection scheme based on long-short dependency is introduced below. The detection scheme is based on a text line detection framework. The framework may include a recursive network and a long-short dependency attention mechanism. Among them, the basic structure of the recursive network is an FPN model, and the first semantic feature is extracted by using the FPN model, and then the recursive idea is adopted to extract the feature again using the FPN model to obtain the second semantic feature. In other words, the FPN model is used for secondary feature extraction to refine the semantic feature. When the framework includes multiple recursive networks, the long-short dependency attention mechanism can be embedded in the FPN model of any recursive network, which is a plug-and-play spatial attention mechanism. The long-short dependency attention mechanism may include strip pooling and feature weighting. Among them, the strip pooling has a long kernel shape in the spatial dimension, which can capture the dependency relationship of distant spatial regions. Due to the long strip kernel shape, while capturing the dependency relationship of the feature map, it can also prevent interference from irrelevant regions. The feature weighting process is to weight the feature map, enhance the features of the text area, so that the framework can aggregate global and local context information at the same time.
[0050] The text line detection framework proposed in the embodiment of the present invention enhances the detection capability of long text lines based on the long-short dependency attention mechanism, and adopts a recursive network to improve the semantic segmentation effect, thereby improving the ability of text line detection.
[0051] The text line detection framework proposed in the embodiment of the present invention has been tested on both public data sets and business data. The test results are as follows:
[0052]
[0053] Table 1
[0054] Table 1 lists the test results of each target detection module. Compared with the FPN model, each target detection module has an improvement in F value, which shows that the text line detection framework proposed in the embodiment of the present invention is very effective.
[0055] The text line detection framework proposed in the embodiment of the present invention can also be applied to the insurance document recognition business and the watermark text line recognition business. Both businesses only focus on the recall rate in terms of indicators. The insurance document recognition business data has a detection test set and an end-to-end test set; the watermark text line recognition business only has an end-to-end data set.
[0056]
[0057] Table 2
[0058] Table 2 lists the results on insurance document recognition business data. The data set is an insurance document recognition test set, which has a total of 1,000 test images. The results show that the text line detection framework proposed in the embodiment of the present invention can improve the recall rate.
[0059]
[0060]
[0061] Table 3
[0062] Table 3 lists the results of the watermark text recognition business. The data set is an end-to-end data set with a total of 105 pictures. The results show that the text line detection framework proposed in the embodiment of the present invention can improve the recall rate.
[0063] The text line detection framework proposed in the embodiment of the present invention extracts long-line features from the feature map based on the distribution characteristics of the text lines, which can capture the long-term dependencies of the spatial region, improve the detection effect of long-line text, and prevent breaks.
[0064] The text line detection framework proposed in the embodiment of the present invention performs secondary recursion on the input features to further refine the segmentation results, thereby improving the text line detection effect, making the text line detection more accurate, and improving recall and precision.
[0065] The text line detection framework proposed in the embodiment of the present invention can be applied to watermark text line detection and insurance document recognition services. It can also be applied to signboard detection, health code recognition, safety network map detection and other services to improve the text line detection effect.
[0066] The text line detection framework proposed in the embodiment of the present invention can be used as the basis of optical character recognition (OCR for short). All business projects requiring OCR can be upgraded using the text line detection framework proposed in the embodiment of the present invention to improve technology and business indicators, thereby increasing related business revenue.
[0067] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0068] Reference Figure 8 , shows a structural block diagram of a text content position prediction device according to an embodiment of the present invention, and the text content position prediction device may specifically include the following modules:
[0069] An image acquisition module 81 is used to acquire an image to be processed, wherein the image includes at least one strip-shaped text content area;
[0070] A recursive module 82, configured to extract a first semantic feature of at least one text content region from the image, and extract a second semantic feature of at least one text content region from the image based on the first semantic feature and the image;
[0071] The attention module 83 is used to perform strip pooling processing and feature weighting processing on the second semantic features respectively to obtain position information of at least one of the text content areas.
[0072] In a preferred embodiment of the present invention, the recursive module 82 comprises:
[0073] A convolution processing module, configured to perform multiple convolution processes from the image in a bottom-up direction to obtain the first semantic feature;
[0074] The input item of each convolution process is the output item of the previous convolution process.
[0075] In a preferred embodiment of the present invention, the convolution processing module is further used to perform convolution processing on the first semantic feature to obtain a top-level convolution processing result;
[0076] The recursive module 82 further includes:
[0077] A deconvolution processing module, configured to perform multiple deconvolution processes from the top-level convolution processing result in a top-down direction to obtain the second semantic feature;
[0078] The input item of each deconvolution process is the output item of the previous deconvolution process and the corresponding output item of each convolution process.
[0079] In a preferred embodiment of the present invention, the attention module 83 comprises:
[0080] The pooling processing module is used to perform horizontal pooling processing and vertical pooling processing on the second semantic feature to obtain a pooling processing result.
[0081] In a preferred embodiment of the present invention, the attention module 83 further includes:
[0082] A weighted processing module is used to perform spatial attention feature weighted processing on the second semantic feature to obtain a weighted processing result.
[0083] In a preferred embodiment of the present invention, the attention module 83 further includes:
[0084] A result fusion module is used to fuse the pooling processing result and the weighted processing result to obtain the position information.
[0085] The embodiment of the present invention further provides an electronic device, see Fig. 9 , including: a processor 901, a memory 902, and a computer program 9021 stored in the memory 902 and executable on the processor 901, wherein when the processor 901 executes the program 9021, the text content position prediction method of the aforementioned embodiment is implemented.
[0086] An embodiment of the present invention further provides a readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for predicting the position of text content of the aforementioned embodiment is implemented.
[0087] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0088] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0089] Those skilled in the art will appreciate that the embodiments of the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0090] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0091] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0092] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0093] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0094] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0095] The above is a detailed introduction to a method and device for predicting the position of text content provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for predicting the position of text content, It is characterized in that The method comprises: obtaining an image to be processed, the image comprising at least one strip-shaped text content region; extracting a first semantic feature of at least one text content region from the image, and extracting a second semantic feature of at least one text content region from the image according to the first semantic feature and the image; performing strip-shaped pooling processing and feature weighting processing on the second semantic feature based on a long-short dependency attention mechanism, respectively, to obtain position information of at least one text content region; Extracting the first semantic feature of at least one text content area from the image includes: performing multiple convolution processes from the image in a bottom-up direction to obtain the first semantic feature; wherein the input item of each convolution process is the output item of the previous convolution process; The extracting a second semantic feature of at least one text content area from the image based on the first semantic feature and the image includes: performing convolution processing on the first semantic feature to obtain a top-level convolution processing result; performing multiple deconvolution processing in a top-down direction starting from the top-level convolution processing result to obtain the second semantic feature; wherein the input item of each deconvolution processing is the output item of the previous deconvolution processing and the corresponding output item of each convolution processing.
2. The method according to claim 1, It is characterized in that The performing strip pooling processing on the second semantic feature includes: performing horizontal pooling processing and vertical pooling processing on the second semantic feature respectively to obtain a pooling processing result.
3. The method according to claim 1, It is characterized in that The performing feature weighted processing on the second semantic feature includes: performing spatial attention feature weighted processing on the second semantic feature to obtain a weighted processing result.
4. The method according to claim 3, It is characterized in that The performing strip pooling processing and feature weighting processing on the second semantic features respectively to obtain the position information of at least one of the text content areas includes: fusing the pooling processing result and the weighting processing result to obtain the position information.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method for predicting the position of text content described in any one of claims 1 to 4 is implemented.
6. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method for predicting the position of text content described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Image semantic segmentation method and device, electronic device and computer-readable medium
CN109447990A
Luggage X-ray contraband image semantic segmentation method combined with attention mechanism
CN110533045A
Character detection method, system and device and storage medium
CN111914843A