Method, apparatus, device, and storage medium for processing text images
By extracting and stitching the text area features in text images, the problem of difficulty in integrating and sorting discrete text information in the prior art is solved, and a more accurate output order of text recognition results and higher readability of text images are achieved.
Patent Information
- Application Number
- CN202110872093.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-30
AI Technical Summary
The existing OCR layout analysis scheme is difficult to effectively integrate and sort discrete text information, resulting in inaccurate output order of text recognition results.
By extracting the initial feature map of the text image, the text content features and spatial position features of the text area are determined, and the text area is spliced into area features. The sorting results of the text area are determined based on these features, and the sorted text recognition results are finally obtained.
It improves the accuracy of the output order of text recognition results, improves the readability of text images, and has high applicability.
Smart Images

Figure CN113822143B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence. Specifically, the present application relates to a method, device, equipment, and storage medium for processing text images. Background Art
[0002] OCR (Optical Character Recognition) layout analysis has always been a very important direction. With the development of technology, information has shown an explosive growth. When facing a large amount of OCR text, it is very limited and costly to obtain the content information therein by using manpower. Therefore, more and more OCR layout analysis methods have emerged. Existing OCR layout analysis solutions include using a text classifier based on CNN (or GCN (Graph Convolutional Network)) to classify text, tables, and pictures (or titles, authors, abstracts, etc.) in the text.
[0003] However, it has been found that most of the existing OCR layout analysis solutions classify the text content of OCR, and the integration and sorting of discrete text information are also very important in layout analysis. However, there is currently no solution to the problem of sorting discrete text. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, equipment, and storage medium for processing text images, which can effectively sort the text content in the text image and have high applicability.
[0005] On the one hand, embodiments of the present application provide a method for processing a text image, the method including:
[0006] Extracting an initial feature map of the text image to be processed;
[0007] According to the above initial feature map, determining the text content features and spatial position features of each text region in at least one text region included in the text image to be processed;
[0008] For each of the above text regions, splicing the text content features and spatial position features of the text region to obtain the region feature of the text region;
[0009] Based on the region features of each of the above text regions, determining the sorting result of each of the above text regions, where the sorting result represents the output order of the text recognition results of each of the above text regions;
[0010] Obtaining the text recognition results of each of the above text regions, and sorting the text recognition results of each of the above text regions based on the sorting result to obtain the text recognition result of the text image to be processed.
[0011] On the other hand, an embodiment of the present application provides a processing device for text images, and the device includes:
[0012] An initial feature map extraction module, configured to extract an initial feature map of a text image to be processed;
[0013] An initial feature map processing module, configured to determine text content features and spatial position features of each text region in at least one text region included in the text image to be processed according to the initial feature map;
[0014] A region feature determination module, configured to splice the text content features and spatial position features of each text region to obtain a region feature of the text region for each text region;
[0015] A sorting result determination module, configured to determine a sorting result of each text region based on the region features of each text region, where the sorting result represents an output order of text recognition results of each text region;
[0016] A text sorting module, configured to obtain text recognition results of each text region, and sort the text recognition results of each text region based on the region sorting result to obtain a text recognition result of the text image to be processed.
[0017] Wherein, the initial feature map processing module is configured to:
[0018] Determine a positional relationship between each feature point in the initial feature map based on the positions of each feature point in the initial feature map;
[0019] Perform feature extraction on the initial feature map based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map to obtain a target feature map;
[0020] Determine a classification result corresponding to each feature point in the target feature map according to the target feature map, where the classification result represents whether each feature point in the target feature map belongs to a text region or a background region;
[0021] Determine at least one text region in the target feature map according to the classification result corresponding to each feature point in the target feature map;
[0022] For each text region, determine text content features corresponding to the text region according to the feature values of each feature point corresponding to the text region in the target feature map, and determine spatial position features corresponding to the text region according to the positions of each feature point corresponding to the text region in the target feature map.
[0023] Optionally, the above-mentioned initial feature map processing module is used for:
[0024] Based on the positions of the feature points in the above-mentioned initial feature map, determine the distances between the feature points in the above-mentioned initial feature map;
[0025] Based on the feature values of the feature points in the above-mentioned initial feature map and the distances between the feature points in the above-mentioned initial feature map, construct a graph structure, perform feature extraction on the above-mentioned graph structure to obtain a target feature map, where each feature point in the above-mentioned initial feature map corresponds to a node in the above-mentioned graph structure, and there is an edge between the nodes corresponding to the feature points whose distances are less than or equal to a set value.
[0026] Optionally, the above-mentioned region feature determination module is used for:
[0027] Fuse the feature values of the same channel among the feature points corresponding to the above-mentioned text region in the above-mentioned target feature map to obtain the fused feature values corresponding to each channel;
[0028] Based on the fused feature values corresponding to the above-mentioned text region, determine the text content features corresponding to the above-mentioned text region.
[0029] Optionally, the above-mentioned sorting result determination module is used for:
[0030] Based on the feature sequence including the region features of the above-mentioned text regions, predict the probabilities corresponding to each of the above-mentioned region features at each time step, and based on the probabilities corresponding to each of the above-mentioned region features at each time step, determine the sorting results of the above-mentioned text regions;
[0031] Among them, the probability corresponding to a region feature at a time step represents the probability that the sorting of the text region corresponding to the region feature among the above-mentioned text regions corresponds to the above-mentioned time step.
[0032] Optionally, the above-mentioned sorting result determination module is used for:
[0033] Perform encoding processing on the region features of the above-mentioned text regions in the above-mentioned feature sequence to obtain the encoding result of the above-mentioned feature sequence;
[0034] For each time step, based on the above-mentioned encoding result and the historical output result corresponding to this time step, predict the probabilities corresponding to each of the above-mentioned region features at this time step;
[0035] Among them, the historical output result corresponding to the first time step is a preset feature. For each time step except the first time step, the historical output result corresponding to this time step includes the prediction result corresponding to the previous time step of this time step. The above prediction result is determined based on the regional features of the text region corresponding to the maximum probability among the probabilities corresponding to the previous time step.
[0036] Optionally, the above sorting result determination module is used for:
[0037] Perform encoding processing on the regional features of each of the above text regions in the above feature sequence to obtain the corresponding hidden state features of each of the above regional features and the encoding result of the above feature sequence;
[0038] For each time step, based on the above encoding result and the historical output result corresponding to this time step, determine the first feature corresponding to this time step; perform feature extraction on each of the above hidden state features to obtain the second feature corresponding to each of the above regional features, and obtain the probability corresponding to each of the above regional features at this time step according to the correlation between the second feature corresponding to each of the above regional features and the above first feature.
[0039] Optionally, the above text sorting module is used for:
[0040] For each of the above text regions, based on the text content features of the above text region, obtain the text recognition result of the above text region.
[0041] Optionally, the above processing device of the text image determines the text content features and spatial position features of each text region in at least one text region included in the above to-be-processed text image according to the above initial feature map, and for each of the above text regions, splicing the text content features and spatial position features of the above text region to obtain the regional features of the above text region is implemented by a graph processing model. The above graph processing model is trained by a model training module. The above model training module is used for:
[0042] Obtain a training sample set, the above training sample set includes at least one sample text image, and the text regions and background regions of the above sample text image are marked with a first sample label, and the above first sample label represents the true result of whether the corresponding region belongs to a text region or a background region;
[0043] Extract the sample initial features of each of the above sample text images;
[0044] For each of the above sample text images, input the sample initial features of the above sample text images into the initial graph processing model to obtain the predicted classification results of each feature point in the sample target feature map corresponding to the above sample text images, where the above predicted classification results represent the predicted results of whether each feature point in the above sample target feature map belongs to the text region or the background region;
[0045] Based on the first sample labels of each of the above sample text images, determine the true results of whether each feature point in the sample target feature map corresponding to each of the above sample text images belongs to the text region or the background region;
[0046] Determine the first training loss value based on the above true results and the above predicted results, and train the above initial graph processing model based on the above first training loss value and the above training sample set until the above first training loss value meets the first training end condition, and then determine the model at the end of training as the above graph processing model.
[0047] Optionally, the processing device of the above text images determines the sorting results of each of the above text regions based on the region features of each of the above text regions through a text sorting model; wherein, each text region of the above sample text images is labeled with a second sample label, and the above second sample label represents the true sorting results of each text region of the above sample text images, and the above text sorting model is trained by the above model training module, and the above model training module is used for:
[0048] Determine each sample feature sequence based on the above graph processing model, and each of the above sample feature sequences includes the region features of each text region of a sample text image;
[0049] For each of the above sample text images, input the sample feature sequence corresponding to the above sample text images into the initial text sorting model to obtain the predicted sorting results of the text regions of the above sample text images;
[0050] Determine the second training loss value based on the above true sorting results and the above predicted sorting results, and train the above initial text sorting model based on the above second training loss value and each of the above sample feature sequences until the above second training loss value meets the second training end condition, and then determine the model at the end of training as the above text sorting model.
[0051] On the other hand, an embodiment of the present application provides an electronic device, including a processor and a memory, and the processor and the memory are connected to each other;
[0052] The above memory is used to store a computer program;
[0053] The above processor is configured to execute the text image processing method provided by the embodiment of the present application when calling the above computer program.
[0054] On the other hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is executed by a processor to implement the method for processing a text image provided by the embodiment of the present application.
[0055] On the other hand, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for processing a text image provided by the embodiment of the present application.
[0056] The beneficial effects brought by the technical solution provided by the embodiment of the present application are as follows:
[0057] By determining the text content features and spatial position features of each text region in the text image to be processed, it is possible to effectively determine the output order of the text recognition results of each text region based on the positions and spatial features of the text regions corresponding to the discrete text information in the text image to be processed, improve the accuracy of the output order of the text recognition results, improve the readability of the text recognition results of the obtained text image, and have high applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments of the present application.
[0059] Figure 1 is a schematic flowchart of a method for processing a text image provided by an embodiment of the present application;
[0060] Figure 2 is a network architecture diagram of a feature extraction model provided by an embodiment of the present application;
[0061] Figure 3 is a flowchart of a method for determining text content features and spatial position features provided by an embodiment of the present application;
[0062] Figure 4 is a schematic network architecture diagram of a graph processing model provided by an embodiment of the present application;
[0063] Figure 5 is a schematic diagram of a scene for determining a text region provided by an embodiment of the present application;
[0064] Figure 6 is a schematic network architecture diagram of a text image processing method provided by an embodiment of the present application;
[0065] Figure 7 It is a schematic structural diagram of a processing device for text images provided by an embodiment of the present application;
[0066] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0067] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.
[0068] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0069] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making.
[0070] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0071] Optionally, when sorting the text recognition results of each text area in the text image to be processed in the embodiments of the present application, it specifically involves the natural language processing (NLP) technology in the field of artificial intelligence. Among them, natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language commonly used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graph.
[0072] The following will specifically describe the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0073] See Figure 1 , Figure 1 is a schematic flowchart of a method for processing a text image provided by an embodiment of the present application. As Figure 1 shown, the method for processing a text image provided by an embodiment of the present application may include the following steps:
[0074] Step S11: Extract the initial feature map of the text image to be processed.
[0075] In the embodiments of the present application, the text image to be processed can be any image including text information, such as website screenshots, web page snapshots, and literature pictures that can include text information, as well as web page images including discrete text information, etc., which are not limited herein.
[0076] Specifically, after obtaining the text image to be processed, while reducing the resolution of the text image to be processed and reducing the amount of data processing, the text image to be processed can be preprocessed based on a feature extraction tool, a feature extraction architecture constructed based on a neural network, etc., to obtain the initial feature map of the text image to be processed.
[0077] Among them, the feature extraction structure constructed based on a neural network may include a feature extraction structure constructed based on a convolutional neural network CNN (Convolutional Neural Networks, CNN), a recurrent neural network (Recurrent Neural Networks, RNN), a long short-term memory neural network (Long short-term memory, LSTM), etc., such as the VGG-16 architecture based on the CNN network, etc., which can be specifically determined based on the requirements of the actual application scenario and will not be limited here.
[0078] As an example, after obtaining the text image to be processed, the feature extraction model based on the VGG-16 architecture can be used to process the text image to be processed to extract the initial feature map of the text image to be processed. Among them, the VGG-16 architecture uses a 3*3 convolutional kernel and a 2*2 max pooling layer, and is specifically composed of 16 hidden layers, including 13 convolutional layers and 3 fully connected layers. Therefore, under the action of the above-mentioned multiple hidden layers, while effectively extracting the initial feature map of the text image to be processed, the initial feature map of the text image to be processed has good generalization performance.
[0079] To ensure that the initial feature map of the text image to be processed is in the form of a point set, therefore, in the feature extraction model based on VGG-16 in the embodiments of the present application, the fully connected layer in the VGG-16 architecture can be cancelled, and then a new feature extraction model can be constructed. As Figure 2 shown, Figure 2 is a network architecture diagram of the feature extraction model provided by the embodiments of the present application. The feature extraction model provided by the embodiments of the present application can be composed of multiple CBR network layers and a pooling layer (Max Pool). Among them, the CBR network layer is a comprehensive network layer composed of a convolutional network (Conv), a batch normalization network (Batch Norm), and an activation function (ReLu). Different CBRs have different data processing capabilities, and the pooling layer can be a max pooling layer.
[0080] As an example, for the high-resolution text image I_s to be processed, its resolution size can be adjusted to 1024×1024. Then, the data size of the 3-channel text image to be processed input into the feature extraction model based on the convolutional neural network is (1024, 1024, 3). After being processed by the feature extraction model based on the convolutional neural network, the data size of the obtained initial feature map I_cnn is (32, 32, 512). Then, the number of data channels is reduced to 8 through multi-layer one-dimensional convolution, that is, the final data size of the initial image feature I_cnn of the text image to be processed is (h, w, c)_(h = 32, w = 32, c = 8).
[0081] Among them, h represents the height of the initial image feature, w represents the width of the initial image feature, and c represents the number of channels of the initial image feature. It should be particularly noted that the feature extraction model provided in the embodiments of the present application is only an example, and can be specifically determined based on the requirements of the actual application scenario, and is not limited herein.
[0082] Step S12: Determine the text content features and spatial position features of each text region in at least one text region included in the text image to be processed according to the initial feature map.
[0083] In the embodiments of the present application, since the initial feature map of the text image to be processed can represent the features of various information in the text image to be processed, such as the features of text information and background information. Among them, the background information of the text image to be processed includes, but is not limited to, non-text information such as color background information and picture information, and can be specifically determined based on the requirements of the actual application scenario, and is not limited herein.
[0084] Based on this, after determining the initial feature map of the text image to be processed, the text content features and spatial position features of each text region in the text image to be processed can be determined based on the initial feature map of the text image to be processed.
[0085] For example, the text image to be processed includes text paragraph 1, text paragraph 2, and background information. After determining the initial feature map of the text image to be processed, the text content features and spatial position features of text paragraph 1 and text paragraph 2 in the text image to be processed can be determined based on the initial feature map.
[0086] Among them, for the specific implementation manner of determining the text content features and spatial position features of each text region in at least one text region included in the text image to be processed according to the initial map feature, reference can be made to Figure 3 . See Figure 3 , Figure 3 which is the flowchart of the method for determining text content features and spatial position features provided by the embodiments of the present application, and specifically includes the following steps:
[0087] Step S121: Determine the positional relationship between each feature point in the initial feature map based on the positions of each feature point in the initial feature map.
[0088] Specifically, the positional relationship between each feature point in the initial feature map can be represented by the distance between each feature point in the initial feature map. Among them, the distance between each feature point in the initial feature map is the distance between the relative positions of each feature point in the initial feature map.
[0089] Step S122: Perform feature extraction on the initial feature map based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map to obtain a target feature map.
[0090] In an embodiment of the present application, when the positional relationship between feature points in the initial feature map can be represented by the distances between feature points in the initial feature map, before performing feature extraction on the initial features, a graph structure can be constructed based on the feature values of feature points in the initial feature map and the positional relationship between feature points in the initial feature map.
[0091] Among them, each feature point in the initial feature map corresponds to a node in the graph structure, the feature value of each feature point in the initial feature map is based on the node feature value of the corresponding node in the graph structure, and there is an edge between nodes corresponding to feature points whose distances between feature points in the initial feature map are less than or equal to a set value.
[0092] In other words, for the edges between nodes in the graph structure, the distance between any feature point in the initial feature map and other feature points can be determined, and an edge is established between the feature points whose distances are less than or equal to the set value and this feature point, so that a graph structure composed of nodes and edges can be constructed.
[0093] Optionally, when determining the edges between nodes in the graph structure, it can also be determined based on the feature values of feature points in the initial feature map. For example, for two feature points in the initial feature map with the same feature value and / or the difference between feature values less than or equal to the set value, it can be determined that there is an edge between the two nodes corresponding to the two feature points in the graph structure.
[0094] Optionally, when determining the edges between nodes in the graph structure, it can also be determined based on the feature values of feature points in the initial graph feature and the distances between feature points in the initial graph feature. If there are two feature points in the initial graph feature whose difference between feature values is less than or equal to the set value and the distance between feature points is less than or equal to the set value, it can be determined that there is an edge between the two nodes corresponding to the two feature points in the graph structure.
[0095] In an embodiment of the present application, after constructing the graph structure based on the initial graph feature, the graph structure can be subjected to feature extraction based on a graph processing model to obtain a target feature map.
[0096] Among them, the above graph processing model includes but is not limited to a neural network model constructed based on a graph convolutional network (GCN) and related graph processing tools, such as a model architecture based on a deep graph convolutional network, a dense convolutional network, etc., which can be specifically determined according to the requirements of the actual application scenario and are not limited here. Refer to Figure 4 , Figure 4 is a schematic diagram of a network architecture of the graph processing model provided by an embodiment of the present application. Figure 4The network architecture shown includes multiple densely connected modules, a convolutional layer, a pooling layer, etc. connected in sequence. To reduce the problem of vanishing gradients caused by deepening the network depth, based on Figure 4 The network structure shown can directly connect all layers on the premise of maximizing information transmission between layers including the graph processing model. Each layer concatenates the inputs of all previous layers and then passes the output features to all subsequent layers.
[0097] It should be specifically noted that Figure 4 The network structure shown is only an example of the graph processing model in the embodiments of this application. The network architecture of the graph processing model can be specifically determined based on the requirements of the actual application scenario and is not limited here.
[0098] Furthermore, after constructing a graph structure based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map, the above graph structure can be processed based on the above graph processing model to extract features from the graph structure and obtain a target feature map.
[0099] As an example, after processing the text image to be processed based on the feature extraction model, the size of the data I_cnn of the obtained initial feature map is (h, w, c)_(h = 32, w = 32, c = 8). Through the graph processing model, based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map, a graph structure can be constructed and features can be extracted from the graph structure to obtain a target feature map.
[0100] Step S123: Determine the classification results corresponding to each feature point in the target feature map according to the target feature map.
[0101] In the embodiments of this application, each feature point in the target feature map can be classified based on the graph processing model. Based on the classification results, it is determined whether each feature point belongs to the local area or the background area, that is, based on the classification results, it is determined whether the information corresponding to each feature point in the text image to be processed belongs to the text area or the background area. The classification results of the feature points in the target feature map represent whether each feature point in the target feature map belongs to the text area or the background area
[0102] Among them, for each feature point in the target feature map, the feature point can be binary classified to obtain a classification result. If the classification result of each feature point in the target feature map is the first value (such as 1), or greater than the set probability value, it can be determined that the feature point represents that it belongs to the text area. If the classification result of each feature point in the target feature map is the second value (such as 0), or not greater than the set probability value, it can be determined that the feature point represents that it belongs to the background area.
[0103] For example, through the graph processing model, the initial feature map I_cnn with a data size of (h, w, c)_(h = 32, w = 32, c = 8) can be finally converted into 1024 feature points. Then, based on the positional relationship and eigenvalues of the 1024 feature points, a graph structure is constructed, thereby obtaining the classification results of the 1024 feature points.
[0104] Step S124: Determine at least one text region in the target feature map according to the classification results corresponding to the feature points in the target feature map.
[0105] In the embodiment of the present application, based on the classification results of the feature points in the target feature map, feature points with the same and adjacent classification results can be determined as a connected domain. Thus, a target connected domain whose classification result represents a text region is determined from each connected domain, and the target connected domain is determined as the text region in the target feature map.
[0106] As Figure 5 shown, Figure 5 is a schematic diagram of the scenario for determining a text region provided in the embodiment of the present application. The initial feature map of the text image to be processed is processed by a feature extraction model such as ResNet and DenseNet, increasing the number of network layers. Among them, each layer corresponds to an intermediate feature map, and the last layer corresponds to the target feature map. Further processing the target feature map to obtain the classification results of the feature points in the target feature map. For each feature point in the target feature map, if the classification result of the feature point represents that it belongs to a text region, it can be marked with a first identifier (such as a first preset value), and if the classification result of the feature point represents that it belongs to a background region, it can be marked with a second identifier (such as a second preset value). Then, based on the marking results of the feature points in the target feature map, feature points with the same and adjacent classification results are determined as a connected domain. If the feature points representing the text region are marked with the first preset value 1, the connected domain corresponding to the text region in the target feature map is as Figure 5 shown in.
[0107] Step S125: For each text region, determine the text content feature corresponding to the text region according to the eigenvalues of the feature points corresponding to the text region in the target feature map, and determine the spatial position feature corresponding to the text region according to the positions of the feature points corresponding to the text region in the target feature map.
[0108] In the embodiment of the present application, since the initial feature map includes multiple channels, the above-mentioned target feature map includes feature maps of multiple channels, and the target feature map includes a feature map corresponding to each channel. In this case, for each text region in the target feature map, the eigenvalues of the feature points corresponding to the text region in the target feature map can be determined.
[0109] Further, for each feature point of the text region in the target feature map, the feature values corresponding to the same channel of each feature point can be fused to obtain the fused feature value corresponding to the same channel of each feature point. For example, for each feature point of the text region in the target feature map, the average value of different feature values corresponding to the same channel of each feature point can be determined as the fused feature value corresponding to this channel of each feature point, and then the fused feature values corresponding to different channels of each feature point can be obtained.
[0110] Optionally, for each feature point of the text region in the target feature map, when fusing the feature values corresponding to the same channel of each feature point, it can also be performed by taking the maximum value, the minimum value, etc., which can be specifically determined based on the requirements of the actual application scenario and will not be limited here.
[0111] Further, for each text region of the target feature map, a fused feature value sequence can be further obtained based on the fused feature values corresponding to different channels of each feature point corresponding to the text region in the target feature map. Among them, the length of the fused feature value sequence is the same as the number of channels corresponding to the target feature map. Based on this, for each text region, the fused feature value sequence corresponding to the text region can be determined as the text content feature corresponding to the text region.
[0112] In the embodiment of the present application, after determining each text region in the target feature map, the spatial position feature corresponding to the text region can be determined based on the positions of the feature points corresponding to each text region in the target feature map.
[0113] Specifically, the position of the text region in the target feature map can be determined based on the positions of the feature points corresponding to the text region in the target feature map, such as coordinates. And according to the positions of the feature points corresponding to the text region in the target feature map, the height and width corresponding to the text region in the target feature map are determined, so as to determine the spatial position feature (x, y, w, h) of the text region based on the position, the corresponding height and width of the text region in the target feature map. Among them, x and y are used to determine the position of the text region in the target feature map, w represents the height of the text region in the target feature map, and h represents the height of the text region in the target feature map.
[0114] Step S13: For each text region, splice the text content feature and the spatial position feature of the text region to obtain the region feature of the text region.
[0115] In the embodiment of the present application, for each text region included in the target feature map, the text content feature and the spatial position feature corresponding to the text region can be fused to obtain the region feature of the text region.
[0116] Among them, when fusing the text content features and spatial position features corresponding to each text region, the text content features and spatial position features corresponding to each text region can be concatenated to obtain the region features corresponding to the text region.
[0117] As an example, the region features corresponding to each text region in the target feature map can be expressed as S i = CONCAT(s i , (x, y, w, h)) for i ∈ 1, 2…, n. Where i represents the index of the text region, and CONCAT(s i , (x, y, w, h)) represents the concatenation of the text content feature s i and the spatial position feature (x, y, w, h) of text region i, and n represents the number of text regions in the target feature map.
[0118] Step S14: Determine the sorting result of each text region based on the region features of each text region.
[0119] In the embodiment of the present application, after determining the region features of each text region in the text image to be processed based on the above implementation, a feature sequence including the region features of each text region can be determined, and the sorting result of each text region can be determined based on the feature sequence including the region features of each text region.
[0120] Specifically, the above feature sequence can be input into a text sorting model to determine the sorting result of each text region in the text image to be processed based on the text sorting model. Among them, the above text sorting model can be a text sorting model constructed based on a Pointer Network (PN), or a text sorting model constructed based on other neural network architectures, such as a Sequence2Sequence architecture, which can be specifically determined according to the requirements of the actual application scenario and will not be limited here.
[0121] Specifically, the above feature sequence can be input into a text sorting model to predict the probability corresponding to each region feature in the feature sequence at each time step, and then determine the sorting result of the text region based on the probability corresponding to each region feature at each time step. Among them, the probability corresponding to a region feature in the above feature sequence at a time step represents the probability that the text region corresponding to the region feature is sorted among the text regions corresponding to the time step.
[0122] As an example, if the above feature sequence includes region feature 1 corresponding to text region 1, region feature 2 corresponding to text region 2, and region feature 3 corresponding to text region 3. At this time, the feature sequence including region feature 1, region feature 2, and region feature 3 is input into the text sorting model. At the first time step, the probabilities of region feature 1, region feature 2, and region feature 3 corresponding to the first time step are determined. If the probability of region feature 2 corresponding to the first time step is the largest, then the sorting of the text region corresponding to region feature 2 is determined as the first sorting.
[0123] Further, at the second time step, the probabilities of region feature 1, region feature 2, and region feature 3 corresponding to the second time step are determined. Then, based on the probabilities of each region feature corresponding to the second time step, the sorting of the text region corresponding to the region feature with the largest probability (except region feature 2) is determined as the second sorting, and the sorting of the text region corresponding to the remaining region features is determined as the third sorting. The first sorting, the second sorting, and the third sorting are arranged in sequence to obtain the sorting result of each text region in the text image to be processed.
[0124] In the embodiment of the present application, when determining the probability of each region feature in the feature sequence corresponding to each time step, the region features of each text region in the feature sequence can be encoded to obtain the encoding result of the feature sequence. For each time step, based on the encoding result and the historical output result corresponding to the current time step, the probability of each region feature corresponding to this time step is determined.
[0125] Among them, the historical output result corresponding to the first time step is a preset feature. For each time step except the first time step, the historical output result corresponding to this time step includes the prediction result corresponding to the previous time step of this time step, and this prediction result is determined based on the region feature of the text region corresponding to the largest probability among the probabilities corresponding to the previous time step.
[0126] As an example, if the above feature sequence includes region feature 1 corresponding to text region 1, region feature 2 corresponding to text region 2, and region feature 3 corresponding to text region 3. At this time, the feature sequence including region feature 1, region feature 2, and region feature 3 is input into the text sorting model, and the region features in the feature sequence are encoded to obtain the encoding result of the feature sequence. At the first time step, based on the preset feature and the encoded feature, the probabilities of each region feature corresponding to the first time step can be determined. If the probability of region feature 2 corresponding to the first time step is the largest, then region feature 2 can be determined as the historical output result corresponding to the next time step.
[0127] Further, for the second time step, based on the encoded features and region feature 2, the probability corresponding to the second time step for each region feature can be determined. If, when aggregating other region features except region feature 2, the probability of region feature 1 corresponding to the second time step is the largest, then region feature 1 can be determined as the historical output result corresponding to the next time step. By analogy, the probabilities corresponding to each region feature at each time step can be obtained.
[0128] In the embodiments of the present application, when encoding the region features of each text region in the feature sequence, the hidden state features corresponding to each region feature and the encoding result of the feature sequence can be obtained during the encoding process.
[0129] For each time step, based on the encoding result and the historical output result corresponding to this time step, the first feature corresponding to this time step is determined, and the above first feature is the attention feature corresponding to this time step. Among them, the first feature (attention feature) corresponding to each time step can be expressed as
[0130] where, W 1 is the attention-related parameter when calculating the attention feature, and can be specifically determined based on the actual model architecture and the requirements of the actual application scenario, and is not limited here. Among them, Z G is the encoding result corresponding to each region feature in the feature sequence. Among them, t represents the time step. When t is 1, it represents the first time step, and the historical output result corresponding to the first time step is the preset feature v input ; at each other time step after the first time step (when t>1), the historical output result corresponding to this time step is the prediction result h t-1 corresponding to the previous time step.
[0131] For each time step, feature extraction can be performed on the hidden state features corresponding to each region feature to obtain the second feature corresponding to each region feature, and the above second feature can be expressed as k i =W 2 h i where, W 2 is the related parameter for calculating the second feature, and can be specifically determined based on the actual model architecture and the requirements of the actual application scenario, and is not limited here. Among them, h i represents the hidden state features corresponding to each region feature.
[0132] Further, for each time step, when determining the probability corresponding to each region feature at this time step based on the encoding result and the historical output result corresponding to this time step, the probability corresponding to each region feature at this time step can be determined according to the correlation between the second feature and the first feature corresponding to each region feature.
[0133] Specifically for each time step, based on the first feature and the second feature of each region feature corresponding to this time step, the correlation coefficient a of the first feature and the second feature corresponding to each region feature can be determined. i Among them, the correlation coefficient of the first feature and the second feature corresponding to each region feature can be expressed as
[0134] Among them, n represents the region features corresponding to all text regions, and n′ is the index of the region feature corresponding to the maximum probability determined at any time step before the current time step. For any region feature i in each time step, if the maximum probability of the historical output result corresponding to any time step between this time step includes the probability corresponding to region feature i, then the correlation coefficient a of region feature i corresponding to this time step i is the preset value A.
[0135] Conversely, the correlation coefficient a corresponding to region feature i i is where q is the first feature corresponding to this time step, k i is the second feature corresponding to this time step, and d is the dimension of the feature sequence.
[0136] Furthermore, for each time step, after determining the correlation coefficient of each region feature i corresponding to this time step, the probability of each region feature i corresponding to this time step can be determined based on the softmax function. Furthermore, the maximum probability is determined from the probabilities of each region feature corresponding to this time step, and the historical output result corresponding to the next time step is determined based on the maximum probability.
[0137] Step S15: Obtain the text recognition results of each text region, and sort the text recognition results of each text region based on the sorting result to obtain the text recognition result of the text image to be processed.
[0138] In the embodiment of the present application, the text recognition results of each text region can be obtained based on natural language processing technologies such as OCR technology, and then the text recognition results of each text region are sorted based on the sorting results of each text region. And based on the sorting results of each text region, the output order of the text recognition results of each text region is determined, and then the text recognition results of each text region are output based on this output order.
[0139] Among them, for each text region, the text recognition result of this text region can be determined based on the text content feature corresponding to this text region.
[0140] In the embodiment of the present application, Figure 1In step S12, according to the initial feature map, determining the text content features and spatial position features of each text region in at least one text region included in the text image to be processed, and in step S13, for each text region, splicing the text content features and spatial position features of the text region to obtain the region features of the text region can be implemented by a graph processing model.
[0141] Among them, the graph processing model is specifically determined in the following manner:
[0142] Obtain a training sample set, the training sample set includes at least one sample text image, and the text regions and background regions of the sample text image are labeled with a first sample label, and the first sample label represents the true result of whether the corresponding region belongs to the text region or the background region;
[0143] Extract the sample initial features of each sample text image;
[0144] For each sample text image, input the sample initial features of the sample text image into the initial graph processing model to obtain the predicted classification results of each feature point in the sample target feature map corresponding to the sample text image, and the predicted classification results represent the predicted results of whether each feature point in the sample target feature map belongs to the text region or the background region;
[0145] Based on the first sample label of each sample text image, determine the true result of whether each feature point in the sample target feature map corresponding to each sample text image belongs to the text region or the background region;
[0146] Determine the first training loss value based on the true result and the predicted result, and train the initial graph processing model based on the first training loss value and the training sample set until the first training loss value meets the first training end condition, and determine the model at the end of training as the graph processing model.
[0147] Among them, for the specific manner of determining the predicted classification results of each feature point in the sample target feature map corresponding to each sample text image, reference can be made to Figures 1 to 2 the implementation manner of determining the classification results corresponding to each feature point in the target feature map corresponding to the text image to be processed shown, which will not be elaborated here.
[0148] Among them, the above first training loss value can be specifically determined based on the cross-entropy loss function or other loss functions, and can be specifically determined according to the actual application scenario requirements, and no limitation is made here.
[0149] As an example, the above first training loss value can be based on to determine. Among them, r represents each feature point in the sample target feature map, y pd represents the predicted result, y gtrepresents the true result, and n represents the number of feature points in the sample feature map.
[0150] According to the first training loss value and each sample text image in the training sample set, the initial graph processing model can be iteratively trained based on the above implementation method, and relevant parameters in the initial graph processing model can be adjusted during each training process. When the first training loss value meets the training end condition, the model at the end of training can be determined as the final graph processing model. Among them, the above training end condition can be that the first training loss value reaches a convergence state, or that the first training loss value is lower than a preset threshold, etc., which can be specifically determined based on the requirements of the actual application scenario and is not limited here.
[0151] In the embodiments of the present application, Figure 1 In step S14, based on the regional features of each text region, determining the sorting result of each text region can be implemented through a text sorting model.
[0152] Among them, when the second sample label is marked on each text region of each sample text image in the training sample set, and the second sample label represents the true sorting result of each text region of the sample text image, the text sorting model is determined in the following manner:
[0153] Based on the graph processing model, each sample feature sequence is determined, and each sample feature sequence includes the regional features of each text region of a sample text image;
[0154] For each sample text image, the sample feature sequence corresponding to the sample text image is input into the initial text sorting model to obtain the predicted sorting result of the text regions of the sample text image;
[0155] Based on the true sorting result and the predicted sorting result, the second training loss value is determined, and the initial text sorting model is trained based on the second training loss value and each sample feature sequence until the second training loss value meets the second training end condition, and the model at the end of training is determined as the text sorting model.
[0156] Among them, for the specific manner of the predicted sorting result of the text regions of each sample text image, reference can be made to Figures 1 to 2 the implementation method of determining the sorting result of each text region in the to-be-processed text image shown, which will not be elaborated here.
[0157] Among them, the above second training loss value can be specifically determined based on the cross-entropy loss function or other loss functions, which can be specifically determined based on the requirements of the actual application scenario and is not limited here.
[0158] Determine the sample feature sequences according to the second training loss value and the graph processing model. The initial text sorting model can be iteratively trained based on the above implementation method, and the relevant parameters of the initial text sorting model can be adjusted during each training process. When the second training loss value meets the training end condition, the model at the end of training can be determined as the final text sorting model. Among them, the above training end condition can be that the second training loss value reaches a convergence state, or that the second training loss value is lower than a preset threshold, etc., which can be specifically determined based on the requirements of the actual application scenario and is not limited here.
[0159] Optionally, when extracting the sample initial features of each sample text image, it can also be obtained through the training of the feature extraction model. And before the initial graph processing model, the sample text images in the training sample set can be input into the initial feature extraction model to obtain the sample initial features of each training text image. And input each sample initial feature into the initial graph processing model to obtain the sample feature sequences of each sample text image, and train the initial text sorting model based on the sample feature sequences obtained from the initial graph processing model. After determining the first training loss value and the second training loss value respectively according to the above implementation method, determine the total training loss value based on the first training loss value and the second training loss value.
[0160] Further, according to the total training loss value and the training sample set, iteratively train the initial feature extraction model, the initial graph processing model, and the initial text sorting model, and adjust the relevant parameters of the initial feature extraction model, the initial text sorting model, and the initial graph processing model during each training process. When the total training loss value meets the training end condition, the models at the end of training can be determined as the final feature extraction model, graph processing model, and text sorting model. Among them, the above training end condition can be that the total training loss value reaches a convergence state, or that the total training loss value is lower than a preset threshold, etc., which can be specifically determined based on the requirements of the actual application scenario and is not limited here.
[0161] Based on the above implementation method, the graph processing model and the text sorting model can be determined. Furthermore, based on the feature extraction model, the graph processing model, and the text sorting model, determine the output order of the text recognition results of each text region in the to-be-processed text image summary. See Figure 6 , Figure 6 is a schematic diagram of the network architecture of the text image processing method provided by the embodiment of the present application. As Figure 6As shown, based on the feature extraction model, the initial feature map of the text image to be processed can be determined. Based on the graph processing model, according to the initial feature map, the text content features and spatial position features of each text region in at least one text region included in the text image to be processed can be determined; for each text region, the text content features and spatial position features of the text region are concatenated to obtain the region feature of the text region. Through the text sorting model, based on the feature sequence obtained by the graph processing model, the sorting result of each text region can be determined, thereby determining the output order of the text recognition results of each text region in the text image to be processed.
[0162] Among them, the sample text images in the above training sample set can be obtained by acquiring user historical access records, text image sampling, and big data, etc., or can be obtained from a database for storing text images, cloud storage, or blockchain. Specifically, it can be determined based on the actual application scenario requirements and is not limited here. Among them, the database can be regarded as an electronic filing cabinet in short - a place for storing electronic files, and can be used to store the sample training set in this application.
[0163] Among them, based on technologies such as data mining in big data, the text images can be mined to form the training sample set in this application.
[0164] Among them, blockchain is a new application mode of computer technologies such as distributed data storage, peer - to - peer transmission, consensus mechanism, and encryption algorithms. Blockchain is essentially a decentralized database and is a string of data blocks generated by using cryptographic methods. In this application, each data block in the blockchain can store the above - mentioned training sample set.
[0165] Among them, cloud storage is a new concept extended and developed from the concept of cloud computing. It refers to the function of clustering applications, grid technology, and distributed storage file systems, etc., which combines a large number of various types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together to store a large number of text images.
[0166] In the embodiment of this application, by determining the text content features and spatial position features of each text region in the text image to be processed, it is possible to effectively determine the output order of the text recognition results of each text region based on the positions and spatial features of the text regions corresponding to the discrete text information in the text image to be processed, improve the accuracy of the output order of the text recognition results, improve the readability of the text recognition results of the obtained text image, and have high applicability.
[0167] See Figure 7 , Figure 7It is a schematic structural diagram of a processing device for text images provided by an embodiment of the present application. The processing device for text images provided by the embodiment of the present application includes:
[0168] An initial feature map extraction module 71, configured to extract an initial feature map of a text image to be processed;
[0169] An initial feature map processing module 72, configured to determine text content features and spatial position features of each text region in at least one text region included in the text image to be processed according to the initial feature map;
[0170] A region feature determination module 73, configured to splice the text content features and spatial position features of each text region to obtain a region feature of the text region for each text region;
[0171] A sorting result determination module 74, configured to determine a sorting result of each text region based on the region features of each text region, where the sorting result represents an output order of text recognition results of each text region;
[0172] A text sorting module 75, configured to obtain text recognition results of each text region, and sort the text recognition results of each text region based on the region sorting result to obtain a text recognition result of the text image to be processed.
[0173] In some embodiments, the initial feature map processing module 72 is configured to:
[0174] Determine a positional relationship between each feature point in the initial feature map based on the positions of each feature point in the initial feature map;
[0175] Extract features from the initial feature map based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map to obtain a target feature map;
[0176] Determine a classification result corresponding to each feature point in the target feature map according to the target feature map, where the classification result represents whether each feature point in the target feature map belongs to a text region or a background region;
[0177] Determine at least one text region in the target feature map according to the classification result corresponding to each feature point in the target feature map;
[0178] For each text region, determine text content features corresponding to the text region according to the feature values of each feature point corresponding to the text region in the target feature map, and determine spatial position features corresponding to the text region according to the positions of each feature point corresponding to the text region in the target feature map.
[0179] In some embodiments, the above-mentioned initial feature map processing module 72 is configured to:
[0180] Based on the positions of the feature points in the above-mentioned initial feature map, determine the distances between the feature points in the above-mentioned initial feature map;
[0181] Based on the feature values of the feature points in the above-mentioned initial feature map and the distances between the feature points in the above-mentioned initial feature map, construct a graph structure, perform feature extraction on the above-mentioned graph structure to obtain a target feature map, wherein each feature point in the above-mentioned initial feature map corresponds to a node in the above-mentioned graph structure, and there is an edge between the nodes corresponding to the feature points whose distances are less than or equal to a set value.
[0182] In some embodiments, the above-mentioned region feature determination module 73 is configured to:
[0183] Fuse the feature values of the same channel among the feature points corresponding to the above-mentioned text region in the above-mentioned target feature map to obtain the fused feature values corresponding to each channel;
[0184] Based on the fused feature values corresponding to the above-mentioned text region, determine the text content features corresponding to the above-mentioned text region.
[0185] In some embodiments, the above-mentioned sorting result determination module 74 is configured to:
[0186] Based on the feature sequence including the region features of each of the above-mentioned text regions, predict the probability corresponding to each of the above-mentioned region features at each time step, and based on the probabilities corresponding to each of the above-mentioned region features at each time step, determine the sorting results of each of the above-mentioned text regions;
[0187] Wherein, the probability corresponding to a region feature at a time step represents the probability that the sorting of the text region corresponding to the region feature among the above-mentioned text regions corresponds to the time step.
[0188] In some embodiments, the above-mentioned sorting result determination module 74 is configured to:
[0189] Perform encoding processing on the region features of each of the above-mentioned text regions in the above-mentioned feature sequence to obtain the encoding result of the above-mentioned feature sequence;
[0190] For each time step, based on the above-mentioned encoding result and the historical output result corresponding to the time step, predict the probability corresponding to each of the above-mentioned region features at the time step;
[0191] Among them, the historical output result corresponding to the first time step is a preset feature. For each time step except the first time step, the historical output result corresponding to this time step includes the prediction result corresponding to the previous time step of this time step. The above prediction result is determined based on the regional features of the text region corresponding to the maximum probability among the probabilities corresponding to the previous time step.
[0192] In some embodiments, the above sorting result determination module 74 is configured to:
[0193] Perform encoding processing on the regional features of each of the above text regions in the above feature sequence to obtain the corresponding hidden state features of each of the above regional features and the encoding result of the above feature sequence;
[0194] For each time step, based on the above encoding result and the historical output result corresponding to this time step, determine the first feature corresponding to this time step; perform feature extraction on each of the above hidden state features to obtain the second feature corresponding to each of the above regional features, and obtain the probability corresponding to each of the above regional features at this time step according to the correlation between the second feature corresponding to each of the above regional features and the above first feature.
[0195] In some embodiments, the above text sorting module 75 is configured to:
[0196] For each of the above text regions, based on the text content features of the above text region, obtain the text recognition result of the above text region.
[0197] In some embodiments, the above text image processing device determines the text content features and spatial position features of each text region in at least one text region included in the above to-be-processed text image according to the above initial feature map, and for each of the above text regions, splicing the text content features and spatial position features of the above text region to obtain the regional features of the above text region is implemented by a graph processing model. The above graph processing model is trained by a model training module, and the above model training module is used for:
[0198] Obtain a training sample set, the above training sample set includes at least one sample text image, the text regions and background regions of the above sample text image are labeled with a first sample label, and the above first sample label represents the true result of whether the corresponding region belongs to a text region or a background region;
[0199] Extract the sample initial features of each of the above sample text images;
[0200] For each of the above sample text images, input the sample initial features of the above sample text images into the initial graph processing model to obtain the predicted classification results of each feature point in the sample target feature map corresponding to the above sample text images, where the predicted classification results represent the predicted results of whether each feature point in the sample target feature map belongs to the text region or the background region;
[0201] Based on the first sample labels of each of the above sample text images, determine the true results of whether each feature point in the sample target feature map corresponding to each of the above sample text images belongs to the text region or the background region;
[0202] Determine the first training loss value based on the above true results and the above predicted results, and train the above initial graph processing model based on the above first training loss value and the above training sample set until the first training loss value meets the first training end condition, and then determine the model at the end of training as the above graph processing model.
[0203] In some embodiments, the processing device of the above text images determines the sorting results of each of the above text regions based on the regional features of each of the above text regions through a text sorting model; where, each text region of the above sample text images is labeled with a second sample label, and the second sample label represents the true sorting results of each text region of the above sample text images, and the text sorting model is trained by a model training module, and the model training module is used for:
[0204] Determine each sample feature sequence based on the above graph processing model, and each of the above sample feature sequences includes the regional features of each text region of a sample text image;
[0205] For each of the above sample text images, input the sample feature sequence corresponding to the above sample text images into the initial text sorting model to obtain the predicted sorting results of the text regions of the above sample text images;
[0206] Determine the second training loss value based on the above true sorting results and the above predicted sorting results, and train the above initial text sorting model based on the above second training loss value and each of the above sample feature sequences until the second training loss value meets the second training end condition, and then determine the model at the end of training as the above text sorting model.
[0207] In specific implementation, the processing device of the text images provided in the embodiments of the present application can execute the processing method of the text images provided in the embodiments of the present application through each built-in functional module, and its implementation principle is similar and will not be elaborated here.
[0208] The processing device for text images provided in the embodiments of the present application may be a computer program (including program code) running on a computer device. For example, the device is an application software, and the device may be used to execute the corresponding steps in the method provided in the embodiments of the present application.
[0209] In some embodiments, the processing device for text images provided in the embodiments of the present application may be implemented in a combination of software and hardware. As an example, the processing device for text images provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the processing method for text images provided in the embodiments of the present invention. For example, the processor in the form of a hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0210] See Figure 8 , Figure 8 is a schematic structural diagram of the electronic device provided in the embodiments of the present application. As Figure 8 shown, the electronic device 1000 in this embodiment may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above-mentioned electronic device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may further be at least one storage device located far from the aforementioned processor 1001. As Figure 8 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0211] In Figure 8In the electronic device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for users to input; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement the website type determination method provided by the embodiments of the present application.
[0212] It should be understood that in some feasible embodiments, the above-mentioned processor 1001 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0213] In specific implementation, the above-mentioned electronic device 1000 can execute the implementation manners provided by each step in the text image processing method provided by the embodiments of the present application through its built-in various functional modules. For specific reference, please refer to the implementation manners provided by each of the above steps, which will not be elaborated here.
[0214] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, which is executed by a processor to implement the implementation manners provided by each step in the text image processing method provided by the embodiments of the present application. For specific reference, please refer to the implementation manners provided by each of the above steps, which will not be elaborated here.
[0215] The above computer-readable storage medium may be an internal storage unit of the device and / or electronic device provided in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. The above computer-readable storage medium may also include magnetic disks, optical disks, read-only memory (ROM), or random access memory (RAM), etc. Further, the computer-readable storage medium may include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0216] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the implementation manners provided in each step of the text image processing method provided in the embodiment of the present application.
[0217] The terms "first", "second", etc. in the claims, the description, and the drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or electronic device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or electronic devices. The mention of "embodiment" in this article means that a specific feature, structure, or characteristic described in combination with the embodiment may be included in at least one embodiment of the present application. The display of this phrase at various positions in the description does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments. The term "and / or" used in the description and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0218] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0219] The above-disclosed are only the preferred embodiments of this application, and thus cannot be used to limit the scope of rights of this application. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. A method for processing text images, characterized in that, Including: Extracting an initial feature map of the text image to be processed; Determining, according to the initial feature map, the text content features and spatial position features of each text region in at least one text region included in the text image to be processed; For each of the text regions, splicing the text content features and spatial position features of the text region to obtain the region feature of the text region; Based on a feature sequence including the region features of the text regions, predicting the probability corresponding to each region feature in the feature sequence at each time step, and determining the sorting result of each text region based on the probability corresponding to each region feature at each time step; wherein, the probability corresponding to a region feature at a time step represents the probability of the sorting of the text region corresponding to the region feature among the text regions corresponding to this time step, and the sorting result represents the output order of the text recognition results of the text regions; Obtaining the text recognition results of the text regions, and sorting the text recognition results of the text regions based on the sorting result to obtain the text recognition result of the text image to be processed.
2. The method according to claim 1, wherein The determining, according to the initial feature map, the text content features and spatial position features of each text region in at least one text region included in the text image to be processed includes: Based on the positions of the feature points in the initial feature map, determining the positional relationship between the feature points in the initial feature map; Performing feature extraction on the initial feature map based on the feature values of the feature points in the initial feature map and the positional relationship between the feature points in the initial feature map to obtain a target feature map; According to the target feature map, determining the classification result corresponding to each feature point in the target feature map, where the classification result represents whether each feature point in the target feature map belongs to a text region or a background region; According to the classification results corresponding to the feature points in the target feature map, determining at least one text region in the target feature map; For each of the text regions, determining the text content feature corresponding to the text region according to the feature values of the feature points corresponding to the text region in the target feature map, and determining the spatial position feature corresponding to the text region according to the positions of the feature points corresponding to the text region in the target feature map.
3. The method according to claim 2, wherein The determining, based on the positions of the feature points in the initial feature map, the positional relationship between the feature points in the initial feature map includes: Based on the positions of the feature points in the initial feature map, determining the distance between the feature points in the initial feature map; The performing feature extraction on the initial feature map based on the feature values of the feature points in the initial feature map and the positional relationship between the feature points in the initial feature map to obtain a target feature map includes: Construct a graph structure based on the feature values of each feature point in the initial feature map and the distances between each feature point in the initial feature map, and perform feature extraction on the graph structure to obtain a target feature map, where each feature point in the initial feature map corresponds to a node in the graph structure, and there is an edge between the nodes corresponding to the feature points whose distances are less than or equal to a set value.
4. The method according to claim 2, wherein The target feature map includes feature maps of multiple channels. For each text region, determining the text content feature corresponding to the text region according to the feature values of each feature point corresponding to the text region in the target feature map includes: Fuse the feature values of the same channel among the feature points corresponding to the text region in the target feature map to obtain the fused feature values corresponding to each channel; Determine the text content feature corresponding to the text region based on the fused feature values corresponding to the text region.
5. The method according to claim 1, wherein Predicting the probabilities corresponding to each of the region features in the feature sequence at each time step based on the feature sequence including the region features of each text region includes: Performing encoding processing on the region features of each text region in the feature sequence to obtain the encoding result of the feature sequence; For each time step, predict the probabilities corresponding to each of the region features at this time step based on the encoding result and the historical output result corresponding to this time step; Wherein, the historical output result corresponding to the first time step is a preset feature. For each time step except the first time step, the historical output result corresponding to this time step includes the prediction result corresponding to the previous time step of this time step, and the prediction result is determined based on the region feature of the text region corresponding to the maximum probability among the probabilities corresponding to the previous time step.
6. The method according to claim 5, characterized in that, Performing encoding processing on the region features of each text region in the feature sequence to obtain the encoding result of the feature sequence includes: Performing encoding processing on the region features of each text region in the feature sequence to obtain the corresponding hidden state features of each region feature and the encoding result of the feature sequence; For each time step, predicting the probabilities corresponding to each of the region features at this time step based on the encoding result and the historical output result corresponding to this time step includes: For each time step, determine the first feature corresponding to this time step based on the encoding result and the historical output result corresponding to this time step; perform feature extraction on each of the hidden state features to obtain the second features corresponding to each region feature, and obtain the probabilities corresponding to each region feature at this time step according to the correlation between the second features corresponding to each region feature and the first feature.
7. The method according to claim 1, characterized in that, Obtaining the text recognition results of each text region includes: For each text region, obtain the text recognition result of the text region based on the text content feature of the text region.
8. The method according to claim 1, wherein Determining the text content features and spatial location features of each text region in at least one text region included in the text image to be processed according to the initial feature map, and for each text region, splicing the text content features and spatial location features of the text region to obtain the region feature of the text region is implemented by a graph processing model; Among them, the graph processing model is determined in the following manner: Obtain a training sample set, which includes at least one sample text image, and the text regions and background regions of the sample text image are labeled with first sample labels, and the first sample labels represent the true results of whether the corresponding regions belong to text regions or background regions; Extract the sample initial features of each sample text image; For each sample text image, input the sample initial features of the sample text image into an initial graph processing model to obtain the predicted classification results of each feature point in the sample target feature map corresponding to the sample text image, and the predicted classification results represent the predicted results of whether each feature point in the sample target feature map belongs to a text region or a background region; Based on the first sample labels of each sample text image, determine the true results of whether each feature point in the sample target feature map corresponding to each sample text image belongs to a text region or a background region; Determine a first training loss value based on the true results and the predicted results, and train the initial graph processing model based on the first training loss value and the training sample set until the first training loss value meets the first training end condition, and determine the model at the end of training as the graph processing model.
9. The method according to claim 8, wherein Predicting the probabilities corresponding to each region feature in each time step of the feature sequence based on the feature sequence including the region features of each text region, and determining the sorting results of each text region based on the probabilities corresponding to each region feature in each time step is implemented by a text sorting model; Among them, each text region of the sample text image is labeled with a second sample label, and the second sample label represents the true sorting result of each text region of the sample text image, and the text sorting model is determined in the following manner: Determine each sample feature sequence based on the graph processing model, and each sample feature sequence includes the region features of each text region of a sample text image; For each sample text image, input the sample feature sequence corresponding to the sample text image into an initial text sorting model to obtain the predicted sorting result of the text region of the sample text image; Determine a second training loss value based on the true sorting result and the predicted sorting result, and train the initial text sorting model based on the second training loss value and each sample feature sequence until the second training loss value meets the second training end condition, and determine the model at the end of training as the text sorting model.
10. A processing device for text images, characterized in that, The device includes: An initial feature map extraction module, configured to extract an initial feature map of a text image to be processed; An initial feature map processing module, configured to determine text content features and spatial position features of each text region in at least one text region included in the text image to be processed according to the initial feature map; A region feature determination module, configured to splice the text content features and spatial position features of each text region to obtain the region feature of the text region for each text region; A sorting result determination module, configured to predict probabilities corresponding to each region feature in the feature sequence at each time step based on a feature sequence including region features of each text region, and determine a sorting result of each text region based on the probabilities corresponding to each region feature at each time step; wherein, the probability corresponding to a region feature at a time step represents the probability of the sorting of the text region corresponding to the region feature among each text region corresponding to this time step, and the sorting result represents the output order of text recognition results of each text region; A text sorting module, configured to obtain text recognition results of each text region, and sort the text recognition results of each text region based on the region sorting result to obtain a text recognition result of the text image to be processed.
11. The device according to claim 10, characterized in that, When determining text content features and spatial position features of each text region in at least one text region included in the text image to be processed according to the initial feature map, the initial feature map processing module is configured to: Determine the positional relationship between each feature point in the initial feature map based on the positions of each feature point in the initial feature map; Perform feature extraction on the initial feature map based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map to obtain a target feature map; Determine a classification result corresponding to each feature point in the target feature map according to the target feature map, where the classification result represents whether each feature point in the target feature map belongs to a text region or a background region; Determine at least one text region in the target feature map according to the classification results corresponding to each feature point in the target feature map; For each text region, determine the text content feature corresponding to the text region according to the feature values of each feature point corresponding to the text region in the target feature map, and determine the spatial position feature corresponding to the text region according to the positions of each feature point corresponding to the text region in the target feature map.
12. The device according to claim 11, wherein When determining the positional relationship between each feature point in the initial feature map based on the positions of each feature point in the initial feature map, the initial feature map processing module is configured to: Determine the distance between each feature point in the initial feature map based on the positions of each feature point in the initial feature map; The performing feature extraction on the initial feature map based on the feature values of each feature point in the initial feature map and the positional relationship between each feature point in the initial feature map to obtain a target feature map includes: Construct a graph structure based on the feature values of each feature point in the initial feature map and the distances between the feature points in the initial feature map, and perform feature extraction on the graph structure to obtain a target feature map, where each feature point in the initial feature map corresponds to a node in the graph structure, and there is an edge between the nodes corresponding to the feature points whose distances are less than or equal to a set value.
13. The device according to claim 11, characterized in that, The target feature map includes feature maps of multiple channels. For each of the text regions, when the region feature determination module determines the text content feature corresponding to the text region according to the feature values of the feature points corresponding to the text region in the target feature map, it is used for: Fuse the feature values of the same channel among the feature points corresponding to the text region in the target feature map to obtain the fused feature values corresponding to each channel; Determine the text content feature corresponding to the text region based on the fused feature values corresponding to the text region.
14. The device according to claim 10, characterized in that, When the sorting result determination module predicts the probability corresponding to each of the region features in the feature sequence at each time step based on the feature sequence including the region features of each of the text regions, it is used for: Perform encoding processing on the region features of each of the text regions in the feature sequence to obtain the encoding result of the feature sequence; For each time step, predict the probability corresponding to each of the region features at this time step based on the encoding result and the historical output result corresponding to this time step; Among them, the historical output result corresponding to the first time step is a preset feature. For each time step except the first time step, the historical output result corresponding to this time step includes the prediction result corresponding to the previous time step of this time step, and the prediction result is determined based on the region feature of the text region corresponding to the maximum probability among the probabilities corresponding to the previous time step.
15. The device according to claim 14, characterized in that, When the sorting result determination module performs encoding processing on the region features of each of the text regions in the feature sequence to obtain the encoding result of the feature sequence, it is used for: Perform encoding processing on the region features of each of the text regions in the feature sequence to obtain the corresponding hidden state features of each of the region features and the encoding result of the feature sequence; For each time step, the prediction of the probability corresponding to each of the region features at this time step based on the encoding result and the historical output result corresponding to this time step includes: For each time step, determine the first feature corresponding to this time step based on the encoding result and the historical output result corresponding to this time step; perform feature extraction on each of the hidden state features to obtain the second features corresponding to each of the region features, and obtain the probability corresponding to each of the region features at this time step according to the correlation between the second features corresponding to each of the region features and the first feature.
16. The device according to claim 10, characterized in that, When the text sorting module obtains the text recognition results of each of the text regions, it is used for: For each of the text regions, obtain the text recognition result of the text region based on the text content feature of the text region.
17. The device according to claim 10, characterized in that, The determination of the text content features and spatial position features of each text region in at least one text region included in the text image to be processed, and for each of the text regions, the splicing of the text content features and spatial position features of the text region to obtain the region feature of the text region is implemented by a graph processing model; Among them, the graph processing model is trained by the model training module in the following manner: Obtain a training sample set, which includes at least one sample text image, and the text regions and background regions of the sample text image are labeled with first sample labels, and the first sample labels represent the true results of whether the corresponding regions belong to text regions or background regions; Extract the sample initial features of each of the sample text images; For each of the sample text images, input the sample initial features of the sample text image into an initial graph processing model to obtain the predicted classification results of each feature point in the sample target feature map corresponding to the sample text image, and the predicted classification results represent the predicted results of whether each feature point in the sample target feature map belongs to a text region or a background region; Based on the first sample labels of each of the sample text images, determine the true results of whether each feature point in the sample target feature map corresponding to each of the sample text images belongs to a text region or a background region; Determine a first training loss value based on the true results and the predicted results, and train the initial graph processing model based on the first training loss value and the training sample set until the first training loss value meets the first training end condition, and determine the model at the end of training as the graph processing model.
18. The device according to claim 17, wherein The prediction of the probabilities corresponding to each of the region features in each time step in the feature sequence based on the feature sequence including the region features of each of the text regions, and the determination of the sorting results of each of the text regions based on the probabilities corresponding to each of the region features in each time step are implemented by a text sorting model; Among them, each text region of the sample text image is labeled with a second sample label, and the second sample label represents the true sorting result of each text region of the sample text image, and the text sorting model is trained by the model training module in the following manner: Determine each sample feature sequence based on the graph processing model, and each of the sample feature sequences includes the region features of each text region of a sample text image; For each of the sample text images, input the sample feature sequence corresponding to the sample text image into an initial text sorting model to obtain the predicted sorting result of the text region of the sample text image; Determine a second training loss value based on the true sorting result and the predicted sorting result, and train the initial text sorting model based on the second training loss value and each of the sample feature sequences until the second training loss value meets the second training end condition, and determine the model at the end of training as the text sorting model.
19. An electronic device, characterized in that, It includes a processor and a memory, and the processor and the memory are connected to each other; The memory is used to store a computer program; The processor is configured to execute the method according to any one of claims 1 to 9 when the computer program is called.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 9.
21. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the electronic device executes the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method and device, equipment and medium
CN112287763A
Image text content recognition method and device, equipment and storage medium
CN112784692A