A method, device and electronic device for text block recognition

Through the combination of feature extraction network and attention network, multiple lines of text in text block images are directly identified, which solves the problem of inaccurate line division in the prior art and improves the accuracy and efficiency of text block recognition.

CN114283432BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110931940.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-13
Publication Date
2025-06-27
Estimated Expiration
2041-08-13

AI Technical Summary

Technical Problem

When identifying text in text block images, the prior art encounters distortion and up and down adhesion between multiple lines of text, resulting in inaccurate line division and reducing the accuracy of text block recognition.

Method used

The feature extraction network is used to extract the features of text block images through the feature extraction module and the classification module, and the codec of the attention network directly recognizes the text in multiple lines of text to improve the recognition accuracy.

Benefits of technology

By extracting the features of the entire text block and identifying it using an attention network, the accuracy and efficiency of text block recognition can be effectively improved, especially when processing multi-line text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283432B_ABST
    Figure CN114283432B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a method, apparatus, and electronic device for text block recognition. The method includes: obtaining a text block image, where the text block image includes one or more lines of text; extracting features of the text block image through a feature extraction network to obtain a first feature map; and recognizing the text in the first feature map through an encoder-decoder to obtain a recognition result of the one or more lines of text, where the encoder-decoder includes an attention network. Embodiments of the present invention can improve the accuracy of text block recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of natural language processing, and in particular, to a method, apparatus, and electronic device for text block recognition. Background Art

[0002] Based on various requirements, it is necessary to recognize the text in a text block image. Currently, a common method for recognizing the text in a text block image is as follows: first, divide the text block by rows, then recognize the text in each row, and then splice the recognized text in each row to obtain the final recognition result. However, in the case where there are distortions, vertical adhesions, etc. between multiple lines of text included in the text block image, during the division, it is possible to divide the text of this row into that row, resulting in inaccurate divided rows, so that the text in each row cannot be accurately recognized, reducing the accuracy of text block recognition. Summary of the Invention

[0003] Embodiments of the present invention disclose a method, apparatus, and electronic device for text block recognition, which are used to improve the accuracy of text block recognition.

[0004] In a first aspect, a method for text block recognition is disclosed, which is characterized by including:

[0005] Obtain a text block image, where the text block image includes one or more lines of text;

[0006] Extract the features of the text block image through a feature extraction network to obtain a first feature map;

[0007] Recognize the text in the first feature map through an encoder-decoder to obtain the recognition result of the one or more lines of text, where the encoder-decoder includes an attention network.

[0008] As a possible implementation, the feature extraction network includes a feature extraction module and a classification module. The step of extracting the features of the text block image through the feature extraction network to obtain a first feature map includes:

[0009] Extract the features of the text block image through the feature extraction module to obtain N feature maps, where the sizes of the N feature maps are different, and N is an integer greater than or equal to 4;

[0010] Select one feature map from the N feature maps through the classification module to obtain a first feature map.

[0011] As a possible implementation, the feature extraction module includes N convolutional neural network (CNN) blocks, and the classification module includes N classification units, and the N CNN blocks correspond to the N classification units one by one;

[0012] The extraction of the features of the text block image by the feature extraction module to obtain N feature maps includes:

[0013] Input the text block image into the first CNN block to obtain the first feature map;

[0014] Use the (i + 1)-th CNN block to perform dimensionality reduction on the i-th feature map to obtain the (i + 1)-th feature map, where i = 1, 2, …, N - 1;

[0015] The selection of one feature map from the N feature maps by the classification module to obtain the first feature map includes:

[0016] When the size of the second feature map is the size of the target feature map corresponding to the first classification unit, determine the second feature map as the first feature map, where the first classification unit is any one of the N classification units, and the second feature map is the feature map corresponding to the first classification unit among the first feature map and the (i + 1)-th feature map.

[0017] As a possible implementation, the N classification units correspond one-to-one with N threshold ranges, and there is no intersection between any two of the N threshold ranges. The j-th classification unit includes a pooling layer, a fully connected layer (FC), and a classification layer. The pooling layer is used to convert the j-th feature map into a feature map of a first size. The FC layer is used to perform dimensionality reduction on the converted j-th feature map. The classification layer is used to determine whether the number of lines of text included in the dimensionality-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit. When it is determined that the number of lines of text included in the dimensionality-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit, determine the size of the j-th feature map as the size of the target feature map corresponding to the j-th classification unit, and determine the j-th feature map as the first feature map, where j = 1, 2, …, N.

[0018] As a possible implementation, the size of the (i + 1)-th feature map is smaller than the size of the i-th feature map.

[0019] As a possible implementation, the method further includes:

[0020] Convert the text block image into an image of a second size;

[0021] The extraction of the features of the text block image by the feature extraction network to obtain the first feature map includes:

[0022] Extract the features of the converted text block image through the feature extraction network to obtain the first feature map.

[0023] As a possible implementation, the method further includes:

[0024] Inputting the text block image into a text block detection network to obtain multiple text block image segments;

[0025] The extracting, by a feature extraction network, of the features of the text block image to obtain a first feature map includes:

[0026] Inputting the multiple text block image segments into a feature extraction network to obtain multiple feature maps;

[0027] The recognizing, by an encoder-decoder, of the text in the first feature map to obtain a recognition result includes:

[0028] Inputting the multiple feature maps into an encoder-decoder to obtain a recognition result.

[0029] A second aspect discloses a text block recognition device, including:

[0030] An acquisition unit, configured to acquire a text block image, where the text block image includes one or more lines of text;

[0031] An extraction unit, configured to extract the features of the text block image through a feature extraction network to obtain a first feature map;

[0032] A recognition unit, configured to recognize the text in the first feature map through an encoder-decoder to obtain a recognition result of the one or more lines of text, where the encoder-decoder includes an attention network.

[0033] As a possible implementation, the feature extraction network includes a feature extraction module and a classification module, and the extraction unit is specifically configured to:

[0034] Extract the features of the text block image through the feature extraction module to obtain N feature maps, where the sizes of the N feature maps are different, and N is an integer greater than or equal to 4;

[0035] Select one feature map from the N feature maps through the classification module to obtain a first feature map.

[0036] As a possible implementation, the feature extraction module includes N CNN blocks, the classification module includes N classification units, and the N CNN blocks correspond to the N classification units one by one;

[0037] The extraction unit extracts the features of the text block image through the feature extraction module to obtain N feature maps, including:

[0038] Inputting the text block image into the first CNN block to obtain a first feature map;

[0039] The (i + 1)-th CNN block is used to reduce the dimension of the i-th feature map to obtain the (i + 1)-th feature map, where i = 1, 2, …, N - 1;

[0040] The step of selecting one feature map from the N feature maps by the classification module to obtain the first feature map includes:

[0041] When the size of the second feature map is the size of the target feature map corresponding to the first classification unit, the second feature map is determined as the first feature map, where the first classification unit is any one of the N classification units, and the second feature map is the feature map corresponding to the first classification unit among the first feature map and the (i + 1)-th feature map.

[0042] As a possible implementation manner, the N classification units correspond one-to-one to N threshold ranges, and there is no intersection between any two of the N threshold ranges. The j-th classification unit includes a pooling layer, an FC, and a classification layer. The pooling layer is used to convert the j-th feature map into a feature map of a first size, the FC is used to reduce the dimension of the converted j-th feature map, and the classification layer is used to determine whether the number of lines of text included in the dimension-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit. When it is determined that the number of lines of text included in the dimension-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit, the size of the j-th feature map is determined as the size of the target feature map corresponding to the j-th classification unit, and the j-th feature map is determined as the first feature map, where j = 1, 2, …, N.

[0043] As a possible implementation manner, the size of the (i + 1)-th feature map is smaller than the size of the i-th feature map.

[0044] As a possible implementation manner, the device further includes:

[0045] A conversion unit, configured to convert the text block image into an image of a second size;

[0046] The extraction unit is specifically configured to extract the features of the converted text block image through a feature extraction network to obtain a first feature map.

[0047] As a possible implementation manner, the device further includes:

[0048] An input unit, configured to input the text block image into a text block detection network to obtain a plurality of text block image segments;

[0049] The extraction unit is specifically configured to input the plurality of text block image segments into a feature extraction network to obtain a plurality of feature maps;

[0050] The recognition unit is specifically configured to input the multiple feature maps into an encoder-decoder to obtain a recognition result.

[0051] A third aspect discloses an electronic device, which may include a processor and a memory. The memory is used to store a set of computer program codes. When the processor is used to call the computer program codes stored in the memory, the processor is caused to execute the method disclosed in the first aspect or any possible implementation manner of the first aspect.

[0052] A fourth aspect discloses an electronic device, which may include a processor, a memory, and a transceiver. The memory is used to store a set of computer program codes. The transceiver is used to receive information from other electronic devices outside the electronic device and output information to other electronic devices outside the electronic device. When the processor is used to call the computer program codes stored in the memory, the processor is caused to execute the method disclosed in the first aspect or any possible implementation manner of the first aspect.

[0053] A fifth aspect discloses a computer-readable storage medium, on which computer programs or computer instructions are stored. When the computer programs or computer instructions run, the method disclosed in the first aspect or any possible implementation manner of the first aspect is implemented.

[0054] A sixth aspect discloses a computer program product, which causes a computer to execute the method disclosed in the first aspect or any possible implementation manner of the first aspect when running on the computer.

[0055] In an embodiment of the present invention, a text block image including one or more lines of text can be obtained. The features of the text block image can be extracted through a feature extraction network to obtain a first feature map. The text in the first feature map can be recognized through an encoder-decoder including an attention network to obtain the recognition result of the above one or more lines of text. The features of the entire text block can be extracted first, and then the encoder-decoder can be directly used to recognize the features of the entire text block. Since the attention network can see the features at any position in the two-dimensional feature map, the encoder-decoder including the attention network can directly recognize multiple lines of text. Therefore, the accuracy of text block recognition can be improved. In addition, since multiple lines of text can be directly recognized without first detecting each line one by one and then recognizing, the efficiency of text block recognition can be improved. Description of the Drawings

[0056] Figure 1 is a schematic flowchart of a text block recognition method disclosed in an embodiment of the present invention;

[0057] Figure 2 It is a schematic structural diagram of a CNN block disclosed in an embodiment of the present invention;

[0058] Figure 3 It is a schematic structural diagram of a feature extraction network disclosed in an embodiment of the present invention;

[0059] Figure 4 It is a schematic structural diagram of a text block recognition model disclosed in an embodiment of the present invention;

[0060] Figure 5 It is a schematic diagram of using the recognition result of a text block recognition model disclosed in an embodiment of the present invention;

[0061] Figure 6 It is a schematic structural diagram of a text block recognition device disclosed in an embodiment of the present invention;

[0062] Figure 7 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present invention. Detailed implementation manners

[0063] An embodiment of the present invention discloses a text block recognition method, device and electronic device, which are used to improve the recognition accuracy. The following will be described in detail respectively.

[0064] To better understand the embodiments of the present invention, the related technologies will be described first. Artificial intelligence (AI) is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.

[0065] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence include speech technology.

[0066] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0067] The present invention realizes the recognition of characters in a text block through NLP technology. The specific implementation is described through the following embodiments.

[0068] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a text block recognition method disclosed in an embodiment of the present invention. Among them, the text block recognition method can be applied to an electronic device capable of image processing, and the electronic device can be provided with a graphics processing unit (GPU). The text block recognition method can also be applied to an application installed on the electronic device. As Figure 1 shown, the text block recognition method can include the following steps.

[0069] 101. Obtain a text block image including one or more lines of text.

[0070] In the case where it is necessary to recognize the characters in the text block, a text block image can be obtained. The text block image can be an image including characters taken by a camera, a PDF document including characters, an electronic scanned document including characters, or a document or picture including characters obtained by other means. The text block image can include one line of text or multiple lines of text. The characters can be characters in different languages and can include numbers, characters, etc. One line of text can be understood as a line of mathematical formulas, a line of characters, a formula + characters, or a line of other text, which is not limited here.

[0071] The text block image can be obtained from the text block images stored locally, from the server, or from a dedicated database. The text block image can also be obtained through an inserted storage device. For example, when the electronic device is a laptop, a desktop computer, etc., the text block image can be obtained from a USB flash drive, a portable hard drive, etc. inserted into the laptop, the desktop computer, etc. The text block image can also be obtained by intercepting displayed documents, pictures, etc. For example, when the user needs to input some text in a PDF document into a Word document, but the user cannot directly copy the required content from the PDF document, the user can capture this part of the content by taking a screenshot to obtain a text block image.

[0072] 102. Extract the features of the text block image through a feature extraction network to obtain a first feature map.

[0073] After obtaining a text block image including one or more lines of text, the features of the text block image can be extracted through a feature extraction network to obtain a first feature map, that is, the text block image can be input into the feature extraction network, and the output of the feature extraction network is the first feature map. When the text block image is an integral image, the first feature map is a single feature map.

[0074] The feature extraction network can include a feature extraction module and a classification module. First, the features of the text block image can be extracted through the feature extraction module to obtain N feature maps, that is, the text block image can be input into the feature extraction module, and the feature extraction module will output N feature maps. The sizes of these N feature maps are different, that is, the sizes of any two of the N feature maps are different. However, the number of lines and the content of the text included in the N feature maps are the same, and are both the number of lines and the content of the text included in the text block image. N is an integer greater than or equal to 4. Since the deeper the network, the higher the level of abstract features that the network can learn, and the higher the level of abstract features, the more accurate the extracted abstract features. Therefore, in order to improve the accuracy of text block recognition, the number of CNN blocks included in the feature extraction module can be 4, or more than 4.

[0075] Afterwards, one feature map can be selected from the N feature maps through the classification module to obtain the first feature map. That is, these N feature maps can be input into the classification module, and the output of the classification module is the first feature map. The first feature map is one of the N feature maps. It can be seen that although there are N feature maps input into the classification module, the feature map output by the classification module is one, and this one feature map is one of the N feature maps. Since the recognition effect of the codec is the best when the size of the text area corresponding to each unit in the feature map input into the codec is fixed, that is, each unit corresponds to a fixed number of lines of text. Therefore, in order to improve the recognition accuracy of the codec, the larger the number of lines included in the text block image, the larger the size of the corresponding first feature map, so that the text block with more lines uses a larger feature map, and the feature block with fewer lines uses a smaller feature map, making the size of the text area corresponding to each unit in the feature map basically the same, and the best recognition effect can be achieved.

[0076] The feature extraction module may include N CNN blocks, and the classification module may include N classification units. The N CNN blocks and the N classification units are in one-to-one correspondence, that is, each of the N CNN blocks is connected to one of the N classification units, and the classification units connected by different CNN blocks are different. The N CNN blocks are connected in sequence. It can be understood that the output of the first CNN block is connected to the input of the second CNN block, the output of the second CNN block is connected to the input of the third CNN block,..., and the output of the (N - 1)th CNN block is connected to the input of the Nth CNN block.

[0077] The text block image can be first input into the first CNN block to obtain the first feature map, that is, the output of the first CNN block is the first feature map. Afterwards, the (i + 1)th CNN block can be used to perform dimensionality reduction processing on the ith feature map to obtain the (i + 1)th feature map. That is, the ith feature map is input into the (i + 1)th CNN block, and the output of the (i + 1)th CNN block is the (i + 1)th feature map, that is, the output of the previous CCN block is used as the input of the next CNN block. It can be seen that by inputting the text block image into the feature extraction module, the first feature map and the (i + 1)th feature map can be obtained, a total of N feature maps, that is, each CNN block will output one feature map. i = 1, 2,..., N - 1.

[0078] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a CNN block disclosed in an embodiment of the present invention. As Figure 2As shown, the CNN block may include two 3×3 convolution layers, two batch normalization (BN) layers, two activation layers, and one pooling layer. The padding of these two convolution layers can be 1 to ensure that the size of the feature map remains unchanged before and after convolution. The convolution layer is used to extract features. The BN layer is used to normalize the output result of the convolution layer to prevent the output of the CNN block from being too large. The activation layer can be a rectified linear unit (ReLU) layer, which is used to transform the linear into the non-linear. The first 3×3 convolution layer is connected to the first BN layer, the first BN layer is connected to the first activation layer, the first activation layer is connected to the second convolution layer, the second convolution layer is connected to the second BN layer, the second BN layer is connected to the second activation layer, and finally connected to a 2×2 pooling layer. The pooling layer is used for downsampling, that is, for dimensionality reduction, and can reduce the length and width of the feature map to 1 / 2 of the original.

[0079] Therefore, the size of the feature map output by the second CNN block is 1 / 2 of the size of the feature map output by the first CNN block, that is, the length of the feature map output by the second CNN block is 1 / 2 of the length of the feature map output by the first CNN block, and the width of the feature map output by the second CNN block is 1 / 2 of the width of the feature map output by the first CNN block, that is, both the length and width are reduced to 1 / 2 of the original. Similarly, the size of the feature map output by the third CNN block is 1 / 2 of the size of the feature map output by the second CNN block, and the size of the feature map output by the fourth CNN block is 1 / 2 of the size of the feature map output by the third CNN block.

[0080] It should be understood that Figure 2 This is only an exemplary illustration of the structure of the CNN block, but does not limit the structure of the CNN block. For example, the CNN block may include three 3×3 convolution layers, three batch normalization layers, three activation layers, and one pooling layer.

[0081] After receiving the second feature map input by the corresponding CNN block, the first classification unit can determine whether the size of the second feature map is the target feature map size corresponding to the first classification unit. When it is determined that the size of the second feature map is the target feature map size corresponding to the first classification unit, the second feature map can be determined as the first feature map. The first classification unit is any one of the N classification units, the second feature map is the feature map corresponding to the first classification unit among the first feature map and the (i + 1)-th feature map, that is, the feature map output by the CNN block corresponding to the first classification unit. It can be seen that the outputs of the N CNN blocks will be used as the inputs of the N classification units respectively, and the input of the j-th classification unit is the output of the j-th CNN block. The structures and processing processes of the N classification units are the same, but the sizes of the feature maps input to the N classification units are different.

[0082] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a feature extraction network disclosed in an embodiment of the present invention. As Figure 3 shown, N is 4, that is, the feature extraction module includes 4 CNN blocks, and the classification module includes 4 classification units. Figure 3 The "c" after the numbers in

[0083] refers to the channel. It can be seen that the 4 CNN blocks are connected in sequence, and each output position of the 4 CNN blocks is connected to a classification unit. As the CNN blocks become deeper, the number of channels of the CNN blocks will also increase continuously. For example, the number of channels of the first CNN block is 64, the number of channels of the second CNN block is 128, the number of channels of the third CNN block is 256, and the number of channels of the fourth CNN block is 256. It can be seen that the number of channels of the next CNN block can be greater than or equal to the number of channels of the previous CNN block. Figure 3As shown, the j-th classification unit may include a pooling layer, an FC, and a classification layer. That is, each of the N classification units includes a pooling layer, an FC, and a classification layer. The pooling layer is used to convert the j-th feature map into a feature map of a first size, that is, to convert the input feature map into a first size. The first sizes corresponding to the N classification units are the same. That is, the pooling layers in the N classification units will convert the N feature maps output by the N CNN blocks into feature maps of the same size. The FC is used to reduce the dimension of the converted j-th feature map. That is, the FC performs dimensionality reduction processing on the feature map output by the pooling layer to reduce it to a feature map with 1024 channels. The classification layer is used to determine whether the number of lines of text included in the reduced-dimensional j-th feature map is within the threshold range corresponding to the j-th classification unit. That is, to determine whether the number of lines of text included in the feature map output by the FC is within the threshold range corresponding to the j-th classification unit. When it is determined that the number of lines of text included in the feature map output by the FC is within the threshold range corresponding to the j-th classification unit, it can be determined that the size of the j-th feature map is the target feature map size corresponding to the j-th classification unit. That is, it can be determined that the size of the j-th feature map is the size of the required feature map. The j-th feature map can be classified into the (YES) category. That is, the classification category of the j-th feature map is the YES category. After that, the j-th feature map can be determined as the first feature map, and the first feature map can be output to the codec. When it is determined that the number of lines of text included in the feature map output by the FC is outside the threshold range corresponding to the j-th classification unit, it can be determined that the size of the j-th feature map is not the target feature map size corresponding to the j-th classification unit. That is, it can be determined that the size of the j-th feature map is not the size of the required feature map. The j-th feature map can be classified into the (NO) category. That is, the corresponding classification category of the j-th feature map is (NO). It can be seen that the classification layer is a binary classifier. When the size of the j-th feature map is the target feature map size corresponding to the j-th classification unit, the j-th feature map can be classified into the YES category; when the size of the j-th feature map is not the target feature map size corresponding to the j-th classification unit, the j-th feature map can be classified into the NO category. The classification results of the j-th feature map are different, and the processing results are different. j = 1, 2,..., N.

[0084] It can be seen that the structures and processing processes of the N classification units are the same, but the sizes of the feature maps input to the N classification units and the corresponding target feature map sizes are different. That is, the corresponding threshold ranges are different. That is, the basis or criterion for classification is different. For example, when N is 4, the threshold range corresponding to the first classification unit can be greater than 6, the threshold range corresponding to the second classification unit can be greater than 4 and less than or equal to 6, the threshold range corresponding to the third classification unit can be greater than 2 and less than or equal to 4, and the threshold range corresponding to the fourth classification unit can be less than or equal to 2.

[0085] It can be seen that although each of the N classification units has an input feature map, that is, there is a corresponding feature map input, since the number of rows included in these feature maps is the same, but the target feature map sizes corresponding to each of the N classification units, that is, the threshold ranges, that is, the classification criteria, are different, the number of rows of text included in the text block image will only fall within one of the N threshold ranges corresponding to the N classification units. Therefore, for the same text block image, only one of the N classification units will have a classification result of the YES category, and the classification results of the remaining N - 1 classification units will be the NO category.

[0086] The pooling layer in the classification unit can be an adaptive average pooling layer, which can convert the size of the feature map to 1*1. The pooling layer can also be other pooling layers. The layering layer can also convert the feature map to a feature map of other sizes, which is not limited here. Then the output of the pooling layer can be input into the FC, and afterwards the channel dimension can be changed to 1024. In addition, there can also be an activation function between the FC and the classification layer, which is used to activate the feature map output by the FC. For example, it can be activated through a rectified linear unit (ReLU) function. The classification layer can be an FC with a channel dimension of 2, which performs binary classification on the activated feature map, that is, classifies whether it is a feature map of the target size. The classification layer can also be other classifiers with binary classification functions. When the number of rows of text included in the feature map input into the classification unit is within the corresponding threshold range, it indicates that the feature map input into the classification unit is a feature map of the target feature map size, and the feature map input into the classification unit can be sent to the codec. When the number of rows of text included in the feature map input into the classification unit is outside the corresponding threshold range, it indicates that the feature map input into the classification unit is not a feature map of the target feature map size, and the result may not be sent to the codec.

[0087] The feature extraction network is a trained network. The training data can include various types of text block images obtained through various methods. These text block images can include marked lines, that is, the positions of each line of text are marked, as well as the annotated number of rows.

[0088] 103. Recognize the text in the first feature map through the codec to obtain a recognition result.

[0089] After extracting the features of the text block image through the feature extraction network to obtain the first feature map, the text in the first feature map can be recognized through the codec to obtain the recognition result of one or more lines of text, that is, the first feature map can be input into the codec, and the output of the codec is one or more lines of text included in the recognized text block image. The codec can include an attention network.

[0090] The codec may include an encoder and a decoder. The first feature map may sequentially pass through the encoder and the decoder. The encoder may include a layer of bidirectional long short term memory (LSTM) network. Assuming the first feature map is denoted as V, the features of each line of text in the first feature map input to the encoder can obtain new features for each line of text Obtained from V The calculation formula can be expressed as follows:

[0091]

[0092] Where, is the initial state, and w represents the line number of the current line of text in the text block. is the encoding result of the encoder for the w-th line, is the encoding result of the encoder for the (w - 1)-th line, V w is the feature of the w-th line in the first feature map. It can be seen that the encoding result of the encoder for the next line is related to the encoding result of the encoder for the previous line and the feature of the current line input. The recurrent neural network (RNN) in the above formula can be understood as an LSTM network.

[0093] The decoder may include a layer of bidirectional LSTM network and an attention network. The decoder can predict the next character y t based on the previous prediction results (y1, y2,..., y ) and the input feature t+1 , and its calculation formula can be expressed as follows:

[0094]

[0095] Where, represents the result of predicting y t based on y1, y2,..., y and t+1 . is the output matrix of the encoder, that is, the first feature map, which is also the input of the decoder. W is a linear transformation matrix. softmax(·) represents the softmax function. o t can be expressed as follows:

[0096] o t = tanh(W[h t ; c t )

[0097] Where, h t is the hidden state of the decoder, and the calculation formula of h t can be expressed as follows:

[0098] h t = RNN(h t-1 ; o t-1 )

[0099] c t is the context feature matrix, which can be calculated based on the feature matrices output by the attention network and the encoder c t The calculation formula can be expressed as follows:

[0100]

[0101] i is the position of the context feature matrix, is the feature matrix at time t, α t is the attention weight matrix. It can be seen that the above formula means: the value at position i in the context feature matrix is the weighted sum of the feature matrix at time t and the corresponding attention weight matrix at position i. The calculation formula of α t can be expressed as follows:

[0102]

[0103] W h is the weight matrix of h i-1 W v is 's weight matrix.

[0104] The encoder-decoder is a trained encoder-decoder. The training data used for training the encoder-decoder can include real data and constructed data. The real data can be obtained from the collected text block images, and the real data can include the labeled text content and the number of lines. In addition, the real data can be augmented. For example, the data can be rotated and radiated. For another example, the brightness of the data can be changed. For another example, the contrast of the data can be changed. For another example, noise can be added to the data. The constructed data can be used to render text block images and their corresponding label information according to the font library file, and the label information can include the labeled text content and the number of lines. The training of the decoder can be used as the training of a conditional language model.

[0105] At Figure 1In the described method, an image of a text block including one or more lines of text can be obtained. The features of the text block image can be extracted through a feature extraction network to obtain a first feature map. The text in the first feature map can be recognized through an encoder-decoder including an attention network to obtain the recognition result of the above one or more lines of text. The features of the entire text block can be extracted first, and then the encoder-decoder can be directly used to recognize the features of the entire text block. Since the attention network can see the features at any position in the two-dimensional feature map, the encoder-decoder including the attention network can directly recognize multiple lines of text. Therefore, the accuracy of text block recognition can be improved. In addition, since multiple lines of text can be directly recognized without detecting each line first and then recognizing, the recognition efficiency can be improved.

[0106] Optionally, before performing step 101 in the above text block recognition method, the image of the text block can be converted into an image of a second size first, that is, the size of the image is converted into a fixed size. Then, the features of the converted text block image can be extracted through a feature extraction network to obtain a first feature map. Since batch training cannot be performed when the sizes of different training images are different, during the training of the feature extraction network, in order to achieve batch training, the training text block images can be converted into text block images of the second size first and then trained. Correspondingly, after the feature extraction network is trained well, during the text block recognition process, before the text block image is input into the feature extraction network, the text block image also needs to be converted into a text block image of the second size first and then input into the feature extraction network. It can be seen that regardless of the size of the obtained text block image, the size of the text block image input into the feature extraction network is fixed. The text block image can be converted into an image of the second size through a size conversion module.

[0107] Optionally, before performing step 101 in the above text block recognition method, the text block image can be input into a text block detection network first to obtain multiple text block image segments, that is, the text block detection network can be used to divide the text block image first, and a text block image can be divided into multiple text block image segments. A text block image segment can include a paragraph of text, that is, the text block image can be divided by paragraph. A text block image segment can also include a formula, that is, the text block image can be divided by formula. A text block image segment can also include a fixed number of lines of text, that is, the text block image can be divided by a fixed number of lines. For example, each text box image segment can include 5 lines of text. The text block image can also be divided in other ways. The division method of the text block image can also include two or more of the above, which is not limited here.

[0108] Afterwards, multiple text block image segments can be input into the feature extraction network to obtain multiple feature maps. The multiple feature maps correspond one-to-one with the multiple text block image segments, that is, one text block image segment can obtain one feature map. The way the feature extraction network processes multiple text block image segments is similar to the way it processes one text block image. The feature extraction network can process each text block image segment in the multiple text block image segments separately, rather than processing them in a jumbled manner.

[0109] Multiple text block image segments can be input into the feature extraction module first. Each of the N CNN blocks included in the feature extraction module will output multiple feature maps. The multiple feature maps output by each CNN block correspond one-to-one with the multiple text block image segments, that is, each CNN block outputs one feature map for each text block image segment in the multiple text block image segments.

[0110] The multiple feature maps output by each CNN block will be input into the corresponding classification unit. The classification unit can process each of the feature maps output by the corresponding CNN block separately. Finally, the classification module will output multiple feature maps. The multiple feature maps output by the classification module correspond one-to-one with the multiple text block image segments. The multiple feature maps output by the classification module can be output by the same classification unit or by multiple classification units. For example, in the case where the number of lines of text included in the multiple text block image segments is within the same threshold range among the N threshold ranges corresponding to the N classification units, the multiple feature maps output by the classification module can be output by the same classification unit. Another example is that in the case where the number of lines of text included in the multiple text block image segments is within different threshold ranges among the N threshold ranges corresponding to the N classification units, the multiple feature maps output by the classification module can be output by multiple classification units. It should be understood that the feature maps output by the classification module corresponding to the text block image segments with the number of lines of text within the same threshold range are output by the same classification unit.

[0111] For example, the text block detection network can divide a text block image into 4 text block image segments, and each of the N CNN blocks will output 4 feature maps. The 4 feature maps output by one CNN block are the feature maps corresponding to the 4 text block image segments. Each classification unit will input the 4 feature maps output by one CNN block. For each of the 4 input feature maps, each classification unit can determine whether the size of the feature map is the target feature map size corresponding to this classification unit. When it is determined that the size of the feature map is the target feature map size corresponding to this classification unit, it indicates that the size of the feature map is the required feature map size, and the feature map can be classified into the YES category. After that, the feature map can be output to the codec. When it is determined that the size of the feature map is not the target feature map size corresponding to this classification unit, it indicates that the size of the feature map is not the required feature map size, and the feature map can be classified into the NO category.

[0112] Before inputting the text block image segments into the feature extraction network, multiple text block image segments can be first converted into text block image segments of a second size, and then the converted multiple text block image segments can be input into the feature extraction network. After that, the multiple feature maps output by the feature extraction network can be input into the codec to obtain the recognition result.

[0113] When the number of lines of the text included in the text block image is large, the speed of the text block recognition model for recognizing the text block image is slow. Therefore, a text block image can be cut into multiple text block image segments through the text block detection network and then recognized. In this way, the number of lines of the text included in the text block image segments can be reduced, and thus the text block recognition efficiency can be improved.

[0114] In one case, the text block recognition model can include a feature extraction network and a codec. In another case, the text block recognition model can include a size conversion module, a feature extraction network and a codec. In yet another case, the text block recognition model can include a text block detection network, a feature extraction network and a codec. In yet another case, the text block recognition model can include a text block detection network, a size conversion module, a feature extraction network and a codec. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a text block recognition model disclosed in an embodiment of the present invention. As Figure 4 shown, the text block recognition model can include a feature extraction network and a codec. It should be understood that Figure 4 the text block recognition model shown is a schematic illustration of the structure of the text block recognition model, and does not limit the structure of the text block recognition model. For example, the text block recognition model can also include a text block detection network. For another example, the text block recognition model can also include a size conversion module.

[0115] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the recognition result using the text block recognition model disclosed in an embodiment of the present invention. As Figure 5 shown, the left part is the input text block image, and the right part is the recognition result output by the text block recognition model. It can be seen that the text block recognition model can accurately recognize the characters that are stuck together.

[0116] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a text block recognition device disclosed in an embodiment of the present invention. Among them, the text block recognition device can be an electronic device capable of performing image processing, and the electronic device can be provided with a GPU. The text block recognition device can also be an application installed on the electronic device. As Figure 6 shown, the text block recognition device may include:

[0117] An acquisition unit 601, configured to acquire a text block image, where the text block image includes one or more lines of text;

[0118] An extraction unit 602, configured to extract the features of the text block image through a feature extraction network to obtain a first feature map;

[0119] A recognition unit 603, configured to recognize the text in the first feature map through an encoder-decoder to obtain the recognition result of the above one or more lines of text, and the encoder-decoder includes an attention network.

[0120] In one embodiment, the feature extraction network may include a feature extraction module and a classification module. The extraction unit 602 is specifically configured to:

[0121] Extract the features of the text block image through the feature extraction module to obtain N feature maps, where the sizes of the N feature maps are different, and N is an integer greater than or equal to 4;

[0122] Select one feature map from the above N feature maps through the classification module to obtain the first feature map.

[0123] In one embodiment, the feature extraction module may include N CNN blocks, and the classification module may include N classification units, and the N CNN blocks correspond to the N classification units one by one;

[0124] The extraction unit 602 extracting the features of the text block image through the feature extraction module to obtain N feature maps may include:

[0125] Input the text block image into the first CNN block to obtain the first feature map;

[0126] Use the (i + 1)-th CNN block to perform dimensionality reduction processing on the i-th feature map to obtain the (i + 1)-th feature map, where i = 1, 2,..., N - 1;

[0127] The extraction unit 602 selects one feature map from the N feature maps through the classification module, and obtaining the first feature map includes:

[0128] When the size of the second feature map is the size of the target feature map corresponding to the first classification unit, the second feature map is determined as the first feature map, the first classification unit is any one of the N classification units, and the second feature map is the feature map corresponding to the first classification unit among the first feature map and the (i + 1)-th feature map.

[0129] In one embodiment, the N classification units correspond one-to-one with N threshold ranges, and there is no intersection between any two of the N threshold ranges. The j-th classification unit includes a pooling layer, an FC, and a classification layer. The pooling layer is used to convert the j-th feature map into a feature map of a first size, the FC is used to perform dimensionality reduction processing on the converted j-th feature map, and the classification layer is used to determine whether the number of lines of text included in the dimension-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit. When it is determined that the number of lines of text included in the dimension-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit, the size of the j-th feature map is determined as the size of the target feature map corresponding to the j-th classification unit, and the j-th feature map is determined as the first feature map, where j = 1, 2,..., N.

[0130] In one embodiment, the size of the (i + 1)-th feature map is smaller than the size of the i-th feature map.

[0131] In one embodiment, the text block recognition device may further include:

[0132] A conversion unit 604, configured to convert the text block image into an image of a second size;

[0133] The extraction unit 602 is specifically configured to extract the features of the converted text block image through a feature extraction network to obtain a first feature map.

[0134] In one embodiment, the text block recognition device may further include:

[0135] An input unit 605, configured to input the text block image into a text block detection network to obtain a plurality of text block image segments;

[0136] The extraction unit 602 is specifically configured to input the plurality of text block image segments into a feature extraction network to obtain a plurality of feature maps;

[0137] The recognition unit 603 is specifically configured to input the plurality of feature maps into an encoder-decoder to obtain a recognition result.

[0138] For a detailed description of the above-mentioned acquisition unit 601, extraction unit 602, recognition unit 603, conversion unit 604, and input unit 605, reference can be directly made to the method embodiments described above Figure 1 and will not be elaborated here.

[0139] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device disclosed in an embodiment of the present invention. As Figure 7 shown, the electronic device may include a processor 701, a memory 702, and a connection line 703. The memory 702 may exist independently, and the connection line 703 is connected to the processor 701. The memory 702 may also be integrated with the processor 701. The connection line 703 may include a path for transmitting information between the above components. The processor 701 includes a GPU. Among them, computer program instructions are stored in the memory 702, and the processor 701 is used to call the computer program instructions stored in the memory 702 to perform the following operations:

[0140] Obtain a text block image, where the text block image includes one or more lines of text;

[0141] Extract the features of the text block image through a feature extraction network to obtain a first feature map;

[0142] Identify the text in the first feature map through an encoder-decoder to obtain the recognition result of the above one or more lines of text. The encoder-decoder includes an attention network.

[0143] In one embodiment, the feature extraction network may include a feature extraction module and a classification module. The processor 701 extracts the features of the text block image through the feature extraction network to obtain a first feature map, including:

[0144] Extract the features of the text block image through the feature extraction module to obtain N feature maps, where the sizes of the N feature maps are different, and N is an integer greater than or equal to 4;

[0145] Select one feature map from the N feature maps through the classification module to obtain the first feature map.

[0146] In one embodiment, the feature extraction module may include N CNN blocks, and the classification module may include N classification units. The N CNN blocks and the N classification units are in one-to-one correspondence;

[0147] The processor 701 extracts the features of the text block image through the feature extraction module to obtain N feature maps, which may include:

[0148] Input the text block image into the first CNN block to obtain the first feature map;

[0149] The (i + 1)-th CNN block is used to perform dimensionality reduction on the i-th feature map to obtain the (i + 1)-th feature map, where i = 1, 2, …, N - 1;

[0150] The processor 701 selects one feature map from the N feature maps through the classification module. Obtaining the first feature map may include:

[0151] When the size of the second feature map is the size of the target feature map corresponding to the first classification unit, the second feature map is determined as the first feature map. The first classification unit is any one of the N classification units, and the second feature map is the feature map corresponding to the first classification unit among the first feature map and the (i + 1)-th feature map.

[0152] In one embodiment, the N classification units correspond one-to-one with N threshold ranges, and there is no intersection between any two threshold ranges among the N threshold ranges. The j-th classification unit includes a pooling layer, an FC, and a classification layer. The pooling layer is used to convert the j-th feature map into a feature map of a first size. The FC is used to perform dimensionality reduction on the converted j-th feature map. The classification layer is used to determine whether the number of lines of text included in the dimensionality-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit. When it is determined that the number of lines of text included in the dimensionality-reduced j-th feature map is within the threshold range corresponding to the j-th classification unit, the size of the j-th feature map is determined as the size of the target feature map corresponding to the j-th classification unit, and the j-th feature map is determined as the first feature map, where j = 1, 2, …, N.

[0153] In one embodiment, the size of the (i + 1)-th feature map is smaller than the size of the i-th feature map.

[0154] In one embodiment, the processor 701 is further configured to call the computer program instructions stored in the memory 702 to perform the following operations:

[0155] Convert the text block image into an image of a second size;

[0156] The processor 701 extracts the features of the text block image through the feature extraction network. Obtaining the first feature map may include:

[0157] Extract the features of the converted text block image through the feature extraction network to obtain the first feature map.

[0158] In one embodiment, the processor 701 is further configured to call the computer program instructions stored in the memory 702 to perform the following operations:

[0159] Input the text block image into the text block detection network to obtain multiple text block image segments;

[0160] The processor 701 extracts the features of the text block image through the feature extraction network, and the obtained first feature map may include:

[0161] Input multiple text block image segments into the feature extraction network to obtain multiple feature maps;

[0162] The processor 701 identifies the text in the first feature map through the codec, and the obtained recognition result includes:

[0163] Input multiple feature maps into the codec to obtain the recognition result.

[0164] In addition, the electronic device may further include a transceiver 704. The transceiver 704 is used to output information to other electronic devices outside the electronic device and receive information sent by other electronic devices outside the electronic device.

[0165] An embodiment of the present invention also discloses a computer-readable storage medium, on which a computer program or computer instructions are stored. When the computer program or computer instructions are executed, the methods in the above method embodiments are executed.

[0166] An embodiment of the present invention also discloses a computer program product containing a computer program or computer instructions. When the computer program or computer instructions are executed, the methods in the above method embodiments are executed.

[0167] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for identifying text blocks, characterized in that, Including: Obtain a text block image, where the text block image includes one or more lines of text; Extract the features of the text block image through a feature extraction network to obtain a first feature map; Identify the text in the first feature map through an encoder-decoder to obtain the recognition result of the one or more lines of text, where the encoder-decoder includes an attention network; Wherein, the feature extraction network includes a feature extraction module and a classification module, and the extracting the features of the text block image through the feature extraction network to obtain a first feature map includes: Extract the features of the text block image through the feature extraction module to obtain N feature maps, where the sizes of the N feature maps are different, and N is an integer greater than or equal to 4; Select one feature map from the N feature maps through the classification module to obtain a first feature map; Wherein, the classification module includes a plurality of classification units. The j-th classification unit is used to judge whether the number of lines of text included in the j-th feature map after dimensionality reduction is within the threshold range corresponding to the j-th classification unit through the classification layer of the j-th classification unit. When the judgment result is yes, determine the size of the j-th feature map as the target feature map size corresponding to the j-th classification unit, and classify the j-th feature map into the yes category, so as to obtain the first feature map from the feature maps judged to be in the yes category. j is a positive integer less than or equal to N.

2. The method according to claim 1, characterized in that The feature extraction module includes N convolutional neural network (CNN) blocks, and the classification module includes N classification units, and the N CNN blocks correspond to the N classification units one by one; The extracting the features of the text block image through the feature extraction module to obtain N feature maps includes: Input the text block image into the first CNN block to obtain a first feature map; Use the (i + 1)-th CNN block to perform dimensionality reduction processing on the i-th feature map to obtain the (i + 1)-th feature map, where i = 1, 2, …, N - 1; The selecting one feature map from the N feature maps through the classification module to obtain a first feature map includes: In the case where the size of the second feature map is the target feature map size corresponding to the first classification unit, determine the second feature map as the first feature map, where the first classification unit is any one of the N classification units, and the second feature map is the feature map corresponding to the first classification unit among the first feature map and the (i + 1)-th feature map.

3. The method according to claim 2, characterized in that, The N classification units correspond one-to-one with N threshold ranges, and there is no intersection between any two of the N threshold ranges. The j-th classification unit includes a pooling layer, a fully connected layer FC, and a classification layer. The pooling layer is used to convert the j-th feature map into a feature map of a first size. The FC is used to perform dimensionality reduction processing on the converted j-th feature map. The classification layer is used to determine whether the number of lines of text included in the j-th feature map after dimensionality reduction is within the threshold range corresponding to the j-th classification unit. When it is determined that the number of lines of text included in the j-th feature map after dimensionality reduction is within the threshold range corresponding to the j-th classification unit, the size of the j-th feature map is determined as the target feature map size corresponding to the j-th classification unit, and the j-th feature map is determined as the first feature map, where j = 1, 2,..., N.

4. The method according to claim 2 or 3, characterized in that, The size of the (i + 1)-th feature map is smaller than the size of the i-th feature map.

5. The method according to claim 1, wherein The method further includes: Converting the text block image into an image of a second size; The obtaining the first feature map by extracting the features of the text block image through the feature extraction network includes: Extracting the features of the converted text block image through the feature extraction network to obtain the first feature map.

6. The method according to claim 1, wherein The method further includes: Inputting the text block image into a text block detection network to obtain a plurality of text block image segments; The obtaining the first feature map by extracting the features of the text block image through the feature extraction network includes: Inputting the plurality of text block image segments into the feature extraction network to obtain a plurality of feature maps, and the plurality of feature maps correspond one-to-one with the plurality of text block segment images; The obtaining the recognition result by identifying the text in the first feature map through the codec includes: Inputting the plurality of feature maps into the codec to obtain the recognition result.

7. A text block recognition device, characterized in that, Including: An acquisition unit for acquiring a text block image, where the text block image includes one or more lines of text; An extraction unit for extracting the features of the text block image through the feature extraction network to obtain the first feature map; A recognition unit for identifying the text in the first feature map through the codec to obtain the recognition result of the one or more lines of text, and the codec includes an attention network; Wherein, the feature extraction network includes a feature extraction module and a classification module. The extraction unit is used to extract the features of the text block image through the feature extraction module to obtain N feature maps, the sizes of the N feature maps are different, and N is an integer greater than or equal to 4; and is used to select one feature map from the N feature maps through the classification module to obtain the first feature map; Among them, the classification module includes a plurality of classification units. The j-th classification unit is used to determine whether the number of lines of the text included in the j-th feature map after dimensionality reduction is within the threshold range corresponding to the j-th classification unit through the classification layer of the j-th classification unit. When the judgment result is yes, it is determined that the size of the j-th feature map is the target feature map size corresponding to the j-th classification unit, and the j-th feature map is classified into the yes category, so as to obtain the first feature map from the feature maps judged to be in the yes category, where j is a positive integer less than or equal to N.

8. An electronic device, characterized in that, It includes a processor and a memory. The memory is used to store a set of computer program codes, and the processor is used to call the computer program codes stored in the memory to implement the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, A computer program or computer instructions are stored in the computer-readable storage medium. When the computer program or computer instructions are run, the method according to any one of claims 1-6 is implemented.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Multidirectional scene character recognition method and system based on multivariate attention mechanism

    CN112215223A