Text image processing method and device and electronic equipment

By performing object detection and deviation correction based on preset text types in text image processing, the accuracy and adaptability problems of traditional methods when identifying complex layout text objects are solved, and a more efficient text image processing effect is achieved.

CN120220154APending Publication Date: 2025-06-27UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213035.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional text image processing methods are difficult to effectively identify and locate target objects in complex layouts, and multiple hyperparameters need to be adjusted to adapt to different document types and layouts, increasing the complexity of the model and the difficulty of adapting to new document fields.

Method used

By obtaining the text image to be processed, object detection is performed based on the preset text type, multiple text detection boxes are obtained, and text detection boxes that do not conform to the dependencies between text types are corrected.

Benefits of technology

It improves the accuracy and comprehensiveness of the object detection results of text images, simplifies the model adjustment process, and enhances the processing ability of complex layout documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220154A_ABST
    Figure CN120220154A_ABST
Patent Text Reader

Abstract

The invention provides a text image processing method and apparatus, and an electronic device. The method comprises the steps of obtaining a to-be-processed text image; based on a preset text type, performing target detection on the text image to obtain a plurality of text detection boxes; and correcting the text detection box which does not conform to the dependency relationship between the text types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technology, and more particularly, to a method, apparatus, and electronic device for processing text images. Background Art

[0002] In the field of document processing, object detection technology is a key component, which is responsible for identifying and locating specific target objects from text images. Traditional methods usually focus on a single type of document element, such as tables, pictures, or mathematical formulas. These early solutions rely on a series of predefined heuristic rules for extracting and identifying target objects from a single type of text image.

[0003] However, this method has limitations because they usually require adjusting multiple hyperparameters to adapt to different document types and layouts, which increases the complexity of the model and the difficulty of adapting to new document domains, and it is difficult to solve the detection of complex layouts. Summary of the Invention

[0004] An object of embodiments of the present disclosure is to provide a method, apparatus, and electronic device for processing text images.

[0005] According to a first aspect of embodiments of the present disclosure, there is provided a method for processing a text image, including:

[0006] Obtain a text image to be processed;

[0007] Perform object detection on the text image based on a preset text type to obtain a plurality of text detection boxes;

[0008] Rectify the text detection boxes that do not conform to the dependency relationship between text types.

[0009] Optionally, the performing object detection on the text image based on a preset text type to obtain a plurality of text detection boxes includes:

[0010] Obtain a target feature map of the text image;

[0011] Perform object detection on the text image according to the target feature map to obtain the plurality of text detection boxes.

[0012] Optionally, the obtaining a target feature map of the text image includes:

[0013] Based on the set first resolution, second resolution, and third resolution respectively, perform feature extraction processing on the text image to obtain the first feature map of the first resolution, the second feature map of the second resolution, and the third feature map of the third resolution. The first resolution is N times the second resolution, and the second resolution is N times the third resolution, where N is a positive integer;

[0014] Perform upsampling processing on the third feature map to obtain a fourth feature map, and the resolution of the fourth feature map is the same as that of the second feature map;

[0015] Perform merging processing on the fourth feature map and the second feature map to obtain a fifth feature map;

[0016] Perform convolution processing on the fifth feature map and then perform upsampling processing to obtain a sixth feature map, and the resolution of the sixth feature map is the same as that of the first feature map;

[0017] Perform splicing processing on the first feature map and the sixth feature map to obtain the target feature map.

[0018] Optionally, the text detection box is represented by corresponding vertex coordinates.

[0019] Optionally, the rectification of the text detection box that does not conform to the dependency relationship between text types includes:

[0020] Determine the template semantic information representing the text type and position corresponding to the text detection box;

[0021] Segment the text image into multiple image regions and identify the image region corresponding to the text detection box;

[0022] Construct a graph network of the text image according to the image region corresponding to the text detection box and the position of the text detection box, where each text detection box serves as a node in the graph network, and the geometric relationship between text detection boxes serves as the edge of the graph network;

[0023] Rectify the text detection box that does not conform to the dependency relationship according to the template semantic information and the graph network.

[0024] Optionally, the rectification of the text detection box that does not conform to the dependency relationship between text types includes:

[0025] Change the text type corresponding to the text detection box that does not conform to the dependency relationship;

[0026] Remove the text detection box with a corresponding incorrect text type.

[0027] Optionally, the method further includes:

[0028] Deduplicate the text detection boxes corresponding to the same text instance.

[0029] Optionally, the deduplication of the text detection boxes corresponding to the same text instance includes:

[0030] Cluster the text detection boxes according to the positions of the text detection boxes to obtain multiple clusters of text detection boxes; the text detection boxes in the same cluster of text detection boxes correspond to the same text instance, and the text detection boxes in different clusters of text detection boxes correspond to different text instances;

[0031] Determine the cluster of text detection boxes containing multiple text detection boxes as the dense text detection box cluster;

[0032] Construct a sub-graph of instances corresponding to each dense text detection box cluster, where the nodes in the sub-graph of instances are the text detection boxes contained in the corresponding dense text detection box cluster, and the edges in the sub-graph of instances represent the set relationship between the text detection boxes in the corresponding dense text detection box cluster;

[0033] According to the sub-graph of instances, fuse the text detection boxes contained in the corresponding dense text detection box cluster into one text detection box.

[0034] According to a second aspect of the present disclosure, there is provided a processing device for text images, including:

[0035] An image acquisition module for acquiring a text image to be processed;

[0036] A target detection module for performing target detection on the text image based on a preset text type to obtain multiple text detection boxes;

[0037] A detection box rectification module for rectifying the text detection boxes that do not conform to the dependency relationship between text types.

[0038] According to a third aspect of the present disclosure, there is provided an electronic device, including a processor and a memory, the memory is used to store a computer program, and the processor is used to execute the method as described in the first aspect of the present disclosure under the control of the computer program.

[0039] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the method as described in the first aspect of the present disclosure when executed by a processor.

[0040] Through the embodiments of the present disclosure, based on a preset text type, target detection is performed on a text image to obtain a plurality of text detection frames, and then the text detection frames that do not conform to the dependency relationship between text types are corrected, which can improve the accuracy and recall rate of the target detection result of the text image.

[0041] Other features and advantages of the present invention will become clear from the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.

[0043] Figure 1 is a block diagram showing the hardware configuration of an electronic device that can implement the embodiments of the present disclosure;

[0044] Figure 2 is a flowchart of a method for processing a text image according to an embodiment of the present disclosure;

[0045] Figure 3 is a schematic diagram of a feature extraction network according to an embodiment of the present disclosure;

[0046] Figure 4 is a block diagram of a device for processing a text image according to an embodiment of the present disclosure;

[0047] Figure 5 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Now, various exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.

[0049] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way limits the present invention, its application, or its use.

[0050] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0051] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0052] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0053] <Hardware Configuration>

[0054] Figure 1 is a block diagram showing the hardware configuration of an electronic device 1000 that can implement the embodiments of the present disclosure.

[0055] The electronic device 1000 can be an electronic product with image processing functions such as a laptop computer, a desktop computer, a mobile phone, a tablet computer, etc. As Figure 1 shown, the electronic device 1000 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and so on. Among them, the processor 1100 can be a processor CPU, a microprocessor MCU, etc. The memory 1200 includes, for example, ROM (Read Only Memory), RAM (Random Access Memory), non-volatile memory such as a hard disk, etc. The interface device 1300 includes, for example, a USB interface, a headphone interface, etc. The communication device 1400 can perform wired or wireless communication, and specifically can include Wifi communication, Bluetooth communication, 2G / 3G / 4G / 5G communication, etc. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, etc. The input device 1600 can include, for example, a touch screen, a keyboard, a somatosensory input, etc. The user can input / output voice information through the speaker 1700 and the microphone 1800.

[0056] Figure 1 The electronic device shown is merely illustrative and in no way limits the present disclosure, its application, or its use. Applied to the embodiments of the present disclosure, the memory 1200 of the electronic device 1000 is used to store instructions for controlling the processor 1100 to operate to execute any one of the methods provided by the embodiments of the present disclosure. Those skilled in the art should understand that although multiple devices are shown for the electronic device 1000 in Figure 1 , the present disclosure may only relate to some of the devices. For example, the electronic device 1000 only relates to the processor 1100 and the memory 1200. Those skilled in the art can design instructions according to the solutions disclosed in the present disclosure. How the instructions control the processor to operate is well known in the art and will not be described in detail here.

[0057] <Method Embodiment>

[0058] The present disclosure provides a method for processing a text image. The method for processing the text image may be implemented by an electronic device. Specifically, the electronic device may be the electronic device 1000 described in the foregoing embodiment.

[0059] Figure 2 FIG. is a flowchart of a method for processing a text image according to an embodiment of the present disclosure.

[0060] As Figure 2 shown, the method includes steps S2100 to S2300 as follows:

[0061] Step S2100, obtain a text image to be processed.

[0062] The text image in this embodiment may be an image containing text elements. Among them, the text elements may include text, graphics, tables, formulas, etc.

[0063] In one example, the text image may be obtained by scanning text content. Specifically, the text image may be transmitted from another electronic device to the electronic device executing this embodiment, or may be downloaded by the electronic device executing this embodiment from the network, or may be obtained by the electronic device executing this embodiment by scanning the text content through a connected scanning device.

[0064] In one example, the text image to be processed may be obtained by scanning a test paper.

[0065] Step S2200, perform target detection on the text image based on a preset text type to obtain a plurality of text detection frames.

[0066] In this embodiment, the text type may be preset according to the content of the text image. For example, it may include multiple types such as printed text, printed formula, handwritten text, handwritten formula, table, picture, and question number.

[0067] Among the multiple text types, there may be an inclusion relationship between at least two text types. For example, printed text may include printed formulas, tables, question numbers, etc. For another example, handwritten text may include handwritten formulas.

[0068] In this embodiment, each text detection frame corresponds to a text instance, and a text instance may be a line of text, a formula, a question number, etc., representing specific text content.

[0069] Furthermore, each text detection frame also corresponds to a text type.

[0070] In one embodiment of the present disclosure, based on a preset text type, target detection is performed on the text image to obtain a plurality of text detection frames, including: obtaining a target feature map of the text image; performing target detection on the text image according to the target feature map to obtain a plurality of text detection frames.

[0071] Specifically, it may be based on a feature extraction network to perform feature extraction on the text image to obtain a target feature map of the text image.

[0072] In one embodiment, the feature extraction network may be a Feature Pyramid Networks (FPN) structure.

[0073] The feature pyramid upsamples the features of the bottom layer and fuses them with the bottom layer features to obtain high-resolution and strong-semantic features (i.e., enhancing feature extraction), which are used as the target feature map.

[0074] In another embodiment, the feature extraction network may be a structure as Figure 3 shown. In this embodiment, obtaining the target feature map of the text image includes steps S2211 to S221 as shown below:

[0075] Step S2211, respectively perform feature extraction processing on the text image based on a set first resolution, second resolution, and third resolution to obtain a first feature map of the first resolution, a second feature map of the second resolution, and a third feature map of the third resolution. The first resolution is N times the second resolution, and the second resolution is N times the third resolution, where N is a positive integer.

[0076] In this embodiment, the first resolution, second resolution, and third resolution may be preset according to the application scenario or specific requirements. For example, the first resolution is W*H / 8, the second resolution is W*H / 16, and the third resolution is W*H / 32, where W*H is the resolution of the text image.

[0077] In this embodiment, it may be to respectively perform convolution processing on the text image using dilated convolutions with dilation rates corresponding to the first resolution, second resolution, and third resolution to obtain the first feature map, the second feature map, and the third feature map.

[0078] Specifically, as Figure 3 shown, it may be to perform feature extraction processing on the text image based on convolution network 1 corresponding to the first resolution to obtain the first feature map of the first resolution; perform feature extraction processing on the text image based on convolution network 2 corresponding to the second resolution to obtain the second feature map of the second resolution; perform feature extraction processing on the text image based on convolution network 3 corresponding to the third resolution to obtain the third feature map of the third resolution.

[0079] If global average pooling is used or an overly large convolutional kernel is used to quickly expand the receptive field, it will result in the loss of certain size features and the overlap of image information. Therefore, in this embodiment, dilated convolutions with different dilation rates are used to increase the receptive field, which can effectively balance the feature maps with multiple resolutions.

[0080] Step S2212: Upsample the third feature map to obtain a fourth feature map, and the resolution of the fourth feature map is the same as that of the second feature map.

[0081] Upsampling refers to restoring a smaller-sized image or feature map to a larger size. Specifically, as Figure 3 shown, it can be based on the upsampling network 1 to upsample the third feature map to obtain the fourth feature map.

[0082] Bilinear interpolation is a commonly used upsampling method. It performs a weighted average of the target pixel based on the surrounding pixel values in the image, thereby smoothly magnifying the small image. Bilinear interpolation considers the 4 neighboring pixels of the target pixel for weighted averaging, making the transition of the upsampled image smoother.

[0083] Step S2213: Combine the fourth feature map and the second feature map to obtain a fifth feature map.

[0084] Upsample the output of 1 / 32 size through bilinear interpolation to expand its size, and then combine the result after upsampling with the output of 1 / 16 size in order to combine feature information of different scales and enhance the final feature map or output.

[0085] In this embodiment, as Figure 3 shown, it can be based on the combination network to combine the fourth feature map and the second feature map to obtain the fifth feature map.

[0086] Step S2214: After performing convolution processing on the fifth feature map and then performing upsampling processing, obtain a sixth feature map, and the resolution of the sixth feature map is the same as that of the first feature map.

[0087] In this embodiment, as Figure 3 shown, it can be based on the convolutional network 4 to perform convolution processing on the fifth feature map, and based on the upsampling network 2 to perform upsampling processing on the fifth feature map after convolution processing to obtain the sixth feature map.

[0088] Step S2215: Concatenate the first feature map and the sixth feature map to obtain the target feature map.

[0089] In this embodiment, as Figure 3As shown, it may be based on a splicing network to splice the first feature map and the sixth feature map to obtain a target feature map.

[0090] Through this embodiment, the three-level hierarchical module can be understood as a small decoder, which can be used to restore some detailed information, and can fuse feature maps of three resolutions at a high level with a small computational cost, thereby improving the target extraction speed of the document image.

[0091] In an embodiment of the present disclosure, it may be based on a neural network to perform target detection on a text image. Specifically, the target feature map may be input into the neural network to obtain a plurality of text detection boxes.

[0092] In this embodiment, the neural network may be a Deep Guided Progressive Network (DGPNet), which can, based on the target separation technology guided by depth information, in addition to the prediction of classification and localization during the text image detection process, also estimate the depth and use it to guide the feature map in a weighted manner to obtain a better spatial representation.

[0093] In another embodiment of the present disclosure, the text detection box may be represented by corresponding vertex coordinates, and the vertex coordinates are the pixel coordinates of the vertices of the corresponding text detection box in the text image, which can represent the position, shape, size, etc. of the text detection box in the text image.

[0094] In this embodiment, the detection of the bounding box is cancelled, and only vertex regression is retained. The smallest rectangle formed by the adjusted vertices can be regarded as a horizontal bounding box. In this way, more accurate and faster target detection can be achieved.

[0095] For handwritten text, the method of extracting vertex features and regressing to predict vertices and maximizing the minimum circumscribed rectangle can be adopted. Specifically, the detection head can be modified to improve the detection of multi-directional text boxes.

[0096] First, the relationship between the predicted vertices can be used to improve vertex prediction, that is, align the smallest rectangle formed by the vertices, and use the predicted detection box to calculate the CIoU (Complete Intersection over Union Loss) loss with the ground truth box.

[0097] The final loss L can be expressed by the following formula:

[0098]

[0099] where L cls is the classification loss, L vr is the vertex regression loss, L va (B, B gt ) is the vertex adjustment loss, B gtis the ground truth box, γ is the loss balance term, which is set to 16 by default. B is the smallest rectangular box formed by the predicted vertices.

[0100] This loss is used to control vertex regression, so as to achieve more accurate and faster object detection.

[0101] Step S2300: Correct the text detection boxes that do not conform to the dependency relationship between text types.

[0102] The dependency relationship in this embodiment can represent the relative position relationship between text instances of each text type. For example, it can represent that the text instance of the question number should be at the forefront of the text instance of the printed text, the text instance of the printed formula should be included in the text instance of the printed text, the text instance of the handwritten formula should be included in the text instance of the handwritten text, and so on.

[0103] Then, the text detection boxes that do not conform to the dependency relationship may include: the text detection box of the corresponding question number whose corresponding text instance is not at the forefront of the text instance representing the printed text, the text detection box of the corresponding printed formula whose corresponding text instance is included in the text instance of the handwritten text, the text detection box of the corresponding handwritten formula whose corresponding text instance is included in the text instance of the printed text, etc.

[0104] In an embodiment of the present disclosure, correcting the text detection boxes that do not conform to the dependency relationship between text types includes: changing the text type corresponding to the text detection box that does not conform to the dependency relationship; or, removing the text detection box with the corresponding text type error.

[0105] For example, it may be removing the text detection box of the corresponding question number whose corresponding text instance is not at the forefront of the text instance representing the printed text.

[0106] For another example, it may be changing the text type corresponding to the text detection box of the corresponding printed formula whose corresponding text instance is included in the text instance of the handwritten text from printed formula to handwritten formula; it may also be changing the text type corresponding to the text detection box of the corresponding handwritten formula whose corresponding text instance is included in the text instance of the printed text from handwritten formula to printed formula.

[0107] In an embodiment of the present disclosure, correcting the text detection boxes that do not conform to the dependency relationship between text types may include the following steps S2310 to S2340:

[0108] Step S2310: Determine the template semantic information representing the text type and position corresponding to the text detection box.

[0109] In this embodiment, the template semantic information may be determined according to the detection result obtained in step S2200.

[0110] Step S2320: Segment the text image into multiple image regions, and identify the image regions corresponding to the text detection boxes.

[0111] In this embodiment, one image region obtained by segmentation may correspond to a text instance, a question, etc.

[0112] Specifically, it may be to perform fine-grained classification on the pixels in each image region according to the positions of the text detection boxes obtained in step S2200, including text lines, tables, formulas, pictures, etc. Then, perform element classification on each image region to identify the text detection boxes corresponding to each image region.

[0113] Step S2330: Construct a graph network of the text image according to the image regions corresponding to the text detection boxes and the positions of the text detection boxes.

[0114] Among them, each text detection box serves as a node in the graph network, and the geometric relationships (such as relative positions, context dependencies, etc.) between the text detection boxes serve as the edges of the graph network.

[0115] In this embodiment, it may be to construct a graph network corresponding to the text image. The edges between the nodes in this graph network represent the positional relationships of the corresponding text detection boxes in space, which can specifically reflect inclusion relationships or adjacency relationships.

[0116] For example, if the text detection box corresponding to a text instance of a line of printed text contains the text detection box corresponding to a text instance of a printed formula or a question number, an inclusion relationship will be formed between these two text detection boxes.

[0117] Step S2340: Correct the text detection boxes that do not conform to the dependency relationship according to the template semantic information and the graph network.

[0118] In this embodiment, the template semantic information can be introduced as a guidance for constraining the graph network to achieve the fusion of the template semantic information and the constrained graph network.

[0119] Furthermore, it may be to process the graph network constructed in step S2330 through the fused constrained graph network to correct the text detection boxes that do not conform to the dependency relationship.

[0120] This embodiment uses the graph network to reason about the dependency relationships between the text detection boxes, thereby improving the spatial relationships and category consistency between the text detection boxes.

[0121] In addition, through the graph attention module, more relevant parts of each specific task are sampled to improve the relevance of the transmitted knowledge. This process also reduces the computational complexity by reducing the total number of nodes in the graph network. More specifically, the graph attention module first uses the Projector module to encode the samples and labels of the support set into task representations, and then samples task-specific potential subgraphs G l from G lat . Then the graph sampler samples G k from G lat according to the query.

[0122] The graph attention module can adaptively sample specific subgraphs from the graph network, optimize the transmission of task-related information, improve accuracy, and reduce computational complexity.

[0123] Embed the template semantics into the constraint graph network, and through the weakly associated knowledge embedding technology, while using a small amount of labeled data, make full use of the template semantics for training to improve the recognition accuracy of the model.

[0124] Through the embodiments of the present disclosure, based on a preset text type, target detection is performed on a text image to obtain multiple text detection frames, and then the text detection frames that do not conform to the dependency relationship between text types are corrected, which can improve the accuracy and recall rate of the target detection results of the text image.

[0125] In an embodiment of the present disclosure, the method further includes: removing duplicates from the text detection frames corresponding to the same text instance.

[0126] In this embodiment, the adaptive aggregation graph network technology can be used to turn multiple text detection frames corresponding to the same text instance into one text detection frame.

[0127] Specifically, it can be based on the new graph fusion network GFNe for multi-directional target detection. The new graph fusion network is scalable and can adaptively fuse multiple text detection frames corresponding to the same text instance to achieve more accurate and comprehensive target detection of the text image.

[0128] In one embodiment of the present disclosure, duplicate removal processing is performed on text detection frames corresponding to the same text instance, including: clustering the text detection frames according to the positions of the text detection frames to obtain multiple text detection frame clusters; the text detection frames in the same text detection frame cluster correspond to the same text instance, and the text detection frames in different text detection frame clusters correspond to different text instances; determining a text detection frame cluster containing multiple text detection frames as a dense text detection frame cluster; constructing a sub-graph of instances corresponding to each dense text detection frame cluster, where the nodes in the sub-graph of instances are the text detection frames included in the corresponding dense text detection frame cluster, and the edges in the sub-graph of instances represent the set relationship between the text detection frames in the corresponding dense text detection frame cluster; and fusing the text detection frames included in the corresponding dense text detection frame cluster into one text detection frame according to the sub-graph of instances.

[0129] In this embodiment, the basis for clustering is the geometric position.

[0130] In this embodiment, the set relationship (the degree of overlap between text detection frames) of each text detection frame can also be determined according to the position of the text detection frame, and clustering is performed according to this set relationship.

[0131] In this embodiment, it may be only the text detection frames included in the dense text detection frame cluster that are fused. For the text detection frames in a text detection frame cluster that contains only one text detection frame, duplicate removal processing may not be performed.

[0132] The constructed sub-graph of instances corresponding to any dense text detection frame cluster, the text detection frames in the dense text detection frame cluster are the nodes of the sub-graph, and the set relationship (such as the degree of overlap) between the text detection frames in the dense text detection frame cluster is the edge of the graph.

[0133] In this embodiment, a graph-based fusion network may be proposed through a graph convolutional network (GCN) to learn to reason and fuse text detection frames, so as to perform duplicate removal processing on text detection frames corresponding to the same text instance, and obtain the final text detection frame of the text image.

[0134] In this embodiment, each text detection frame in each dense text detection frame cluster may be weighted and fused respectively, so that the obtained one text detection frame may be the minimum rectangle frame of the corresponding text instance.

[0135] <Device Embodiment>

[0136] This embodiment provides a processing device for text images, such as Figure 4As shown, the processing device 4000 of the text image includes an image acquisition module 4100, an object detection module 4200, and a detection box rectification module 4300. The image acquisition module 4100 is used to acquire the text image to be processed; the object detection module 4200 is used to perform object detection on the text image based on a preset text type to obtain a plurality of text detection boxes; the detection box rectification module 4300 is used to rectify the text detection boxes that do not conform to the dependency relationship between text types.

[0137] In an embodiment of the present disclosure, the object detection module 4200 is used to:

[0138] Obtain the target feature map of the text image;

[0139] Perform object detection on the text image according to the target feature map to obtain the plurality of text detection boxes.

[0140] In an embodiment of the present disclosure, the obtaining of the target feature map of the text image includes:

[0141] Perform feature extraction processing on the text image respectively based on a set first resolution, second resolution, and third resolution to obtain a first feature map with the first resolution, a second feature map with the second resolution, and a third feature map with the third resolution. The first resolution is N times the second resolution, and the second resolution is N times the third resolution, where N is a positive integer;

[0142] Perform upsampling processing on the third feature map to obtain a fourth feature map, and the resolution of the fourth feature map is the same as that of the second feature map;

[0143] Perform merging processing on the fourth feature map and the second feature map to obtain a fifth feature map;

[0144] Perform convolution processing on the fifth feature map and then perform upsampling processing to obtain a sixth feature map, and the resolution of the sixth feature map is the same as that of the first feature map;

[0145] Perform splicing processing on the first feature map and the sixth feature map to obtain the target feature map.

[0146] In an embodiment of the present disclosure, the text detection box is represented by corresponding vertex coordinates.

[0147] In an embodiment of the present disclosure, the detection box rectification module 4300 is used to:

[0148] Determine the template semantic information representing the text type and position corresponding to the text detection box;

[0149] Segment the text image into multiple image regions, and identify the image regions corresponding to the text detection frames;

[0150] Construct a graph network of the text image according to the image regions corresponding to the text detection frames and the positions of the text detection frames, where each text detection frame serves as a node in the graph network, and the geometric relationship between the text detection frames serves as the edge of the graph network;

[0151] Rectify the text detection frames that do not conform to the dependency relationship according to the template semantic information and the graph network.

[0152] In one embodiment of the present disclosure, rectifying the text detection frames that do not conform to the dependency relationship between text types includes:

[0153] Change the text types corresponding to the text detection frames that do not conform to the dependency relationship;

[0154] Eliminate the text detection frames with corresponding incorrect text types.

[0155] In one embodiment of the present disclosure, the text image processing device 4000 further includes:

[0156] A duplicate removal processing module for performing duplicate removal processing on the text detection frames corresponding to the same text instance.

[0157] In one embodiment of the present disclosure, the duplicate removal processing module specifically is used for:

[0158] Cluster the text detection frames according to the positions of the text detection frames to obtain multiple text detection frame clusters; the text detection frames in the same text detection frame cluster correspond to the same text instance, and the text detection frames in different text detection frame clusters correspond to different text instances;

[0159] Determine the text detection frame clusters containing multiple text detection frames as dense text detection frame clusters;

[0160] Construct a sub-graph of instances corresponding to each dense text detection frame cluster, where the nodes in the sub-graph of instances are the text detection frames included in the corresponding dense text detection frame cluster, and the edges in the sub-graph of instances represent the set relationship between the text detection frames in the corresponding dense text detection frame cluster;

[0161] Fuse the text detection frames included in the corresponding dense text detection frame cluster into one text detection frame according to the sub-graph of instances.

[0162] <Embodiment of the electronic device>

[0163] This embodiment provides an electronic device. On the one hand, the electronic device may include the aforementioned text image processing device 4000.

[0164] In another aspect, as Figure 5 shown, the electronic device 5000 may include a processor 5100 and a memory 5200. The memory 5200 is used to store a computer program, and the processor 5100 is used to control the electronic device to execute the method of any embodiment of the present disclosure under the control of the computer program.

[0165] <Readable storage medium embodiment>

[0166] This embodiment provides a computer-readable storage medium storing a computer program, which when executed by a processor, executes the method described in any method embodiment of the present disclosure.

[0167] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0168] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0169] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0170] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0171] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0172] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, generate a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0173] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0174] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0175] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A text image processing method, characterized in that: include: Get the text image to be processed; Based on a preset text type, target detection is performed on the text image to obtain a plurality of text detection frames; Correct the text detection boxes that do not conform to the dependency relationship between text types.

2. The method according to claim 1, characterized in that The method of performing target detection on the text image based on a preset text type to obtain a plurality of text detection frames includes: Obtaining a target feature map of the text image; Performing target detection on the text image according to the target feature map to obtain the multiple text detection frames.

3. The method according to claim 2, characterized in that The step of obtaining a target feature map of the text image comprises: Based on the set first resolution, second resolution and third resolution, feature extraction processing is performed on the text image to obtain a first feature map of the first resolution, a second feature map of the second resolution and a third feature map of the third resolution, wherein the first resolution is N times the second resolution, the second resolution is N times the third resolution, and N is a positive integer; Performing upsampling processing on the third feature map to obtain a fourth feature map, where a resolution of the fourth feature map is the same as a resolution of the second feature map; Merging the fourth feature map and the second feature map to obtain a fifth feature map; Performing convolution processing on the fifth feature map and then upsampling processing to obtain a sixth feature map, where the resolution of the sixth feature map is the same as the resolution of the first feature map; The first feature map and the sixth feature map are concatenated to obtain the target feature map.

4. The method according to claim 2, characterized in that: The text detection box is represented by corresponding vertex coordinates.

5. The method according to claim 1, characterized in that The correction of the text detection frame that does not conform to the dependency relationship between text types includes: Determining template semantic information representing the text type and position corresponding to the text detection box; Segmenting the text image into a plurality of image regions, and identifying the image region corresponding to the text detection frame; According to the image area corresponding to the text detection box and the position of the text detection box, a graph network of the text image is constructed, wherein each text detection box serves as a node in the graph network, and the geometric relationship between the text detection boxes serves as an edge of the graph network; According to the template semantic information and the graph network, the text detection box that does not conform to the dependency relationship is corrected.

6. The method according to claim 1, characterized in that The correction of the text detection frame that does not conform to the dependency relationship between text types includes: Changing the text type corresponding to the text detection box that does not meet the dependency relationship; Eliminate text detection boxes with incorrect text types.

7. The method according to claim 1, characterized in that The method further comprises: De-duplicate text detection boxes corresponding to the same text instance.

8. The method according to claim 7, characterized in that The deduplication processing of the text detection boxes corresponding to the same text instance includes: Clustering the text detection frames according to their positions to obtain a plurality of text detection frame clusters; the text detection frames in the same text detection frame cluster correspond to the same text instance, and the text detection frames in different text detection frame clusters correspond to different text instances; Determine a text detection box cluster including multiple text detection boxes as a dense text detection box cluster; Constructing an instance subgraph corresponding to each dense text detection box cluster, wherein the nodes in the instance subgraph are text detection boxes included in the corresponding dense text detection box cluster, and the edges in the instance subgraph represent the set relationship between the text detection boxes in the corresponding dense text detection box cluster; According to the example sub-graph, the text detection frames included in the corresponding dense text detection frame cluster are merged into one text detection frame.

9. A text image processing device, characterized in that: include: An image acquisition module, used to acquire the text image to be processed; An object detection module is used to perform object detection on the text image based on a preset text type to obtain a plurality of text detection frames; The detection frame correction module is used to correct the text detection frame that does not conform to the dependency relationship between text types.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the method according to any one of claims 1 to 8 under the control of the computer program.