Text recognition methods, electronic devices and computer storage media

By using an attention-based machine learning model and multi-granularity label parsing technology, the problem of low accuracy of the Transformer model in low-quality image text recognition is solved, achieving more efficient and accurate text recognition.

CN115393864BActive Publication Date: 2026-04-03ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing Transformer models have low accuracy in recognizing text in low-quality images, and the lack of semantic information leads to inaccurate recognition.

Method used

A machine learning model based on an attention mechanism is used to extract features from image patch sequences, and semantic information is implicitly injected into the model through multi-granularity label parsing (character granularity, sub-word granularity, and whole word granularity) for text recognition.

Benefits of technology

It improves the accuracy and efficiency of text recognition, and obtains more semantic information through multi-granularity tag parsing, thereby enhancing text recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393864B_ABST
    Figure CN115393864B_ABST
Patent Text Reader

Abstract

This application provides a text recognition method, an electronic device, and a computer storage medium. The text recognition method includes: obtaining an image block sequence corresponding to a text image to be recognized; extracting image features from the image block sequence using a machine learning model based on an attention mechanism to obtain corresponding image features; performing multi-granularity label parsing based on the image features, wherein the multi-granularity label parsing includes: character-level label parsing, sub-word-level label parsing, and whole-word-level label parsing; and performing text recognition on the text image based on the multi-granularity label parsing results. This application embodiment improves the performance of text recognition to a higher level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to a text recognition method, electronic device, and computer storage medium. Background Technology

[0002] Text recognition is a technique that detects and identifies text in images to obtain their corresponding textual information. With the widespread use of the Transformer model in the field of natural language processing, more and more technicians are trying to replace the convolutional neural network model with it to apply it to the field of image recognition, specifically for text recognition of images containing text.

[0003] However, since images containing text often lack semantic information, applying the Transformer model to images containing text for text recognition often results in inaccurate recognition, especially for low-quality images, where the recognition accuracy is low. Summary of the Invention

[0004] In view of this, embodiments of this application provide a text recognition scheme to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of the embodiments of this application, a text recognition method is provided, comprising: obtaining an image block sequence corresponding to a text image to be recognized; extracting image features from the image block sequence using a machine learning model based on an attention mechanism to obtain corresponding image features; performing multi-granularity label parsing based on the image features, wherein the multi-granularity label parsing includes: character-level label parsing, sub-word-level label parsing, and whole-word-level label parsing; and performing text recognition on the text image based on the multi-granularity label parsing results.

[0006] According to a second aspect of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, which causes the processor to perform an operation corresponding to the method described in the first aspect.

[0007] According to a third aspect of the embodiments of this application, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0008] According to a fourth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to perform an operation corresponding to the method described in the first aspect.

[0009] According to the solution provided in the embodiments of this application, by performing multi-granularity labeling and parsing on image features, semantic information can be implicitly injected into the model that processes text images. This allows the model to simultaneously combine image features and semantic information for text recognition, improving recognition efficiency and accuracy. Furthermore, this semantic information is represented at multiple granularities—character granularity, sub-word granularity, and whole-word granularity—enabling the acquisition of image features and semantic information from multiple granularities, thereby elevating the performance of text recognition to a higher level. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0011] Figure 1 A schematic diagram of an exemplary system to which the embodiments of this application are applicable;

[0012] Figure 2A This is a flowchart illustrating the steps of a text recognition method according to an embodiment of this application.

[0013] Figure 2B for Figure 2A A schematic diagram of the structure of a text recognition model in the embodiment shown;

[0014] Figure 2C For use Figure 2B The diagram shows the multi-granularity prediction results obtained by the text recognition model.

[0015] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0016] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0017] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.

[0018] Figure 1 An exemplary system applicable to embodiments of this application is shown. For example... Figure 1 As shown, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106. Figure 1 The example in the text shows multiple user devices.

[0019] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, in some embodiments, the cloud server 102 can perform text recognition on text images. As an optional example, in some embodiments, the cloud server 102 can perform multi-granularity label parsing based on the image features corresponding to the text image to integrate text semantic information into the image features for text recognition. As an optional example, in some embodiments, the cloud server 102 can set up a text recognition model and perform text recognition on text images through the text recognition model. As another example, in some embodiments, the cloud server 102 can receive a text recognition request sent by the user device 106 and perform text recognition on the text image requested by the request. Further, in some embodiments, the cloud server 102 can also send the text recognition result back to the user device 106.

[0020] In some embodiments, communication network 104 may be any suitable combination of one or more wired and / or wireless networks. For example, communication network 104 may include any one or more of the following: the Internet, intranet, wide area network (WAN), local area network (LAN), wireless network, digital subscriber line (DSL) network, frame relay network, asynchronous transfer mode (ATM) network, virtual private network (VPN), and / or any other suitable communication network. User equipment 106 may be connected to communication network 104 via one or more communication links (e.g., communication link 112), and communication network 104 may be linked to cloud server 102 via one or more communication links (e.g., communication link 114). Communication links may be any communication link suitable for transmitting data between user equipment 106 and cloud server 102, such as network links, dial-up links, wireless links, hardwired links, any other suitable communication links, or any suitable combination of such links.

[0021] User device 106 may include any one or more user devices suitable for interaction, such as interacting with a user or cloud server 102. In some embodiments, user device 106 may send a text recognition request to cloud server 102, enabling cloud server 102 to acquire the requested text image for text recognition. In some embodiments, the text recognition request sent by user device 106 to cloud server 102 carries information about the corresponding text image or the address for acquiring the corresponding text image, enabling cloud server 102 to acquire the text image. In some embodiments, user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include mobile devices, tablet computers, laptop computers, desktop computers, wearable computers, game consoles, media players, vehicle entertainment systems, and / or any other suitable type of user device.

[0022] Based on the above system, the text recognition scheme of this application will be described below through embodiments.

[0023] Reference Figure 2A The diagram illustrates a flowchart of the steps of a text recognition method according to an embodiment of this application.

[0024] The text recognition method in this embodiment includes the following steps:

[0025] Step S202: Obtain the image block sequence corresponding to the text image to be recognized.

[0026] In this application embodiment, the text image includes plain text images and images that simultaneously contain text image elements and non-text image elements, both of which are applicable to the solutions in this application embodiment.

[0027] Furthermore, in this embodiment, to facilitate image feature extraction by the subsequent attention-based machine learning model, the text image is first segmented into a series of non-overlapping continuous image patches, i.e., an image patch sequence. These image patches, along with their positional information, serve as input to the subsequent attention-based machine learning model.

[0028] For example, the text image to be recognized can be divided into blocks using convolution, each block can be flattened into a sequence, and then positional encoding and CLS token (classification label) can be added to the image block sequence and input into a subsequent attention-based machine learning model, such as an encoder with a Transformer structure.

[0029] Step S204: Extract image features from the image patch sequence using an attention-based machine learning model to obtain the corresponding image features.

[0030] Unlike conventional image feature extraction using convolutional neural networks, this embodiment employs an attention-based machine learning model to extract features from image patch sequences, taking advantage of the excellent performance of the attention mechanism in feature extraction. For example, this machine learning model can utilize the encoder structure from the Transformer class.

[0031] In order to use the Transformer encoder for image feature extraction without changing its structure, as mentioned earlier, after processing the image into a sequence of image blocks, positional encoding and CLS tokens are added so that the Transformer encoder can process it directly.

[0032] By using attention-based machine learning models, such as the Transformer encoder, to extract image features from image patch sequences, the corresponding image features can be obtained and represented as corresponding feature tokens.

[0033] Step S206: Perform multi-granularity label parsing based on image features.

[0034] The multi-granularity tag parsing includes: character-level tag parsing, sub-word-level tag parsing, and whole-word-level tag parsing.

[0035] Tokenization, also known as tokenization, generates character substrings based on feature representations, each with relatively complete semantics. In this embodiment, tokenization introduces textual semantic information into image processing, enabling text recognition of text images to be performed simultaneously based on image features and semantic information, thus achieving more accurate recognition results.

[0036] Furthermore, in this embodiment of the application, the tag parsing is implemented at three granularities: character granularity, sub-word granularity, and whole-word granularity.

[0037] Among them, the character granularity is character-level. By performing label parsing on the image features at the character granularity, the most basic characters can be parsed and predicted, such as 'a', 'b', 'c' in English, or '你', '我', '他' in Chinese, etc. The whole-word granularity is wordpiece-level. By performing label parsing on the image features at the whole-word granularity, natural language units can be parsed and predicted, such as 'Transformer', 'coffee' in English, or '这个餐馆做的菜很好', '这件衣服真漂亮' in Chinese, etc. The subword granularity is subword-level, which is a granularity between the character granularity and the whole-word granularity. It considers word segmentation based on common combinations. By performing label parsing on the image features at the subword granularity, subwords between characters and whole words can be parsed and predicted. For example, 'Transformer' in English will be parsed and predicted as 'Transform', 'er'; 'coffee' will be parsed and predicted as 'co', 'ff', 'ee', etc.; or, '这个餐馆做的菜很好' in Chinese will be parsed and predicted as '这个', '餐馆', '做的菜', '很好'; '这件衣服真漂亮' will be parsed and predicted as '这件', '衣服', '真', '漂亮', etc.

[0038] In a feasible manner, the label parsing of the subword granularity in the embodiments of this application can be implemented as the label parsing of the byte pair encoding (BPE) granularity. The label parsing of the BPE granularity will perform subword segmentation according to the frequency of the most frequently combined characters, so that the predicted subwords are more reasonable. In fact, the label parsing of the wordpiece-level can also be regarded as a variant of BPE, which is a coarser granularity encoding format based on subwords.

[0039] The above multi-granularity label parsing can be implemented by means of attention processing + classification. For example, the label parsing of the character granularity based on image features can be implemented as follows: According to the single character granularity, use the spatial attention function to select the image features related to the i-th character in the image features, where i is the number of text characters in the text image, such as 'coffee' has 6 characters; then, aggregate the selected image features to generate a vector corresponding to the i-th character; and then, classify and identify through a classifier to obtain the character text corresponding to the i-th character, such as 'c', etc.

[0040] Sub-word granularity labeling and parsing based on image features can be implemented as follows: Spatial attention calculation and classification processing of image features are performed using the BPE algorithm, and the processing result is used as the sub-word granularity labeling and parsing result. The BPE algorithm counts the frequency of each consecutive byte pair and selects the highest frequency pair to merge into a new highest frequency byte pair. Based on this, the spatial attention function obtained after training can select image features related to the j-th sub-word according to the sub-word granularity, and aggregate them to generate a vector corresponding to the j-th sub-word; then, a classifier is used for classification and recognition to obtain the character text corresponding to the j-th sub-word. Here, j is the number of text sub-words in the text image. For example, "coffee" includes 2 sub-words, and the character text corresponding to the j-th sub-word might be "co" or "ffee".

[0041] The whole-word granularity labeling and parsing based on image features can be implemented as follows: The WordPiece algorithm is used to perform spatial attention calculation and classification on image features, and the result is used as the whole-word granularity labeling and parsing result. The WordPiece algorithm can be seen as a variant of the BPE algorithm, the difference being that the WordPiece algorithm generates new subwords based on probability rather than the next most frequent byte pair. Based on this, the spatial attention function obtained after training can select image features related to the m-th whole word from the image features according to the whole-word granularity, and aggregate them to generate a vector corresponding to the m-th whole word; then, a classifier is used for classification and recognition to obtain the character text corresponding to the m-th whole word. Here, m is the number of whole words in the text image. For example, "I like coffee" includes 3 whole words, and the character text corresponding to the m-th subword may be "I", "like", or "coffee".

[0042] Through the above multi-granularity label parsing, on the one hand, more and richer image features and semantic information can be obtained through multiple levels of parsing and prediction, thereby improving the text recognition performance for text images; on the other hand, by using spatial attention mechanism and classification processing, accurate parsing and prediction at different levels can be performed, thereby improving the accuracy and efficiency of parsing and prediction.

[0043] Step S208: Perform text recognition on the text image based on the multi-granularity label parsing results.

[0044] After obtaining the multi-granularity label parsing results, text recognition can be performed based on these results. This includes: obtaining text prediction results of multiple granularities based on the multi-granularity label parsing results; fusing the text prediction results of multiple granularities; and performing text recognition on the text image based on the fusion result. It should be noted that obtaining the text prediction results requires not only the label parsing results but also a pre-set vocabulary. Different granularity label parsing results correspond to different vocabularys. For example, a character-granularity vocabulary includes characters such as a, b, c, and d, as well as other symbols such as #, @, and *. In this embodiment, the size and specific implementation of the vocabulary (characters, sub-words, or whole words) in different granularity vocabularys are not limited. However, for ease of explanation, the following description uses a vocabulary containing 256 elements as an example.

[0045] Among them, fusing text prediction results at multiple granularities and performing text recognition on text images based on the fusion result can be achieved as follows:

[0046] Method 1: Calculate the mean of the probability distributions indicated by the text prediction results at each of the multiple granularities to obtain the mean of the probability distributions corresponding to the text prediction results at each granularity; determine the text prediction result corresponding to the maximum mean of the multiple probability distributions corresponding to the multiple granularities as the target text prediction result; perform text recognition on the text image based on the target text prediction result.

[0047] Continuing with the example of the text "coffee" in a text image, based on the 256 characters in the vocabulary, each character in "coffee" has a corresponding probability, forming a probability distribution such as [0, 0, 0.9, 0, 0.05, 0, 0, ..., 0.05, 0, 0 ...]. This probability distribution indicates that the probability of "c" being the character 'c' is 0.9, the probability of it being the character 'e' is 0.02, the probability of it being the character 'o' is 0.05, and the probability of it being any other character is 0. Similarly, each other character also has a similar probability distribution. Therefore, at the character granularity, the mean of the probability distributions for all characters corresponding to "coffee" can be calculated to obtain the mean of the probability distribution at that character granularity, denoted as P1.

[0048] Similarly, at the sub-word granularity level, "coffee" also has a corresponding probability distribution of multiple sub-words based on the sub-word granularity vocabulary. Therefore, at the sub-word granularity level, the mean of the probability distribution of all sub-words corresponding to "coffee" can be obtained, denoted as P2.

[0049] At the whole-word granularity level, "coffee" also has a corresponding probability distribution based on the whole-word granularity vocabulary. Therefore, at the whole-word granularity level, the mean of the probability distribution corresponding to "coffee" can be obtained as the mean of the probability distribution at that whole-word granularity level, denoted as P3.

[0050] Assuming that in the example above, P3>P2>P1, then the text prediction result corresponding to P3 is determined as the target text prediction result. Based on this result, whole word recognition is performed, and the character corresponding to the "coffee" image portion in the text image is identified as "coffee".

[0051] By using the mean of the probability distribution, the prediction results of text prediction at each granularity can be represented in a relatively balanced way, and more objective and accurate prediction results can be obtained.

[0052] Method 2: Multiply the probability distributions indicated by the text prediction results of each granularity to obtain the probability product results corresponding to the text prediction results of each granularity; determine the text prediction result corresponding to the maximum product among the multiple probability product results corresponding to multiple granularities as the target text prediction result; perform text recognition on the text image based on the target text prediction result.

[0053] Continuing with the example of the text "coffee" in a text image, based on the 256 characters in the vocabulary, each character in "coffee" has a corresponding probability, forming a probability distribution such as [0, 0, 0.9, 0, 0.05, 0, 0, ..., 0.05, 0, 0...]. This probability distribution indicates that the probability of "c" being the character 'c' is 0.9, the probability of it being the character 'e' is 0.02, the probability of it being the character 'o' is 0.05, and the probability of it being any other character is 0. Similarly, each other character also has a similar probability distribution. Therefore, at the character granularity, multiplying the probability distributions of all characters corresponding to "coffee" yields the probability distribution product at that character granularity, denoted as M1.

[0054] Similarly, at the sub-word granularity level, "coffee" also has a corresponding probability distribution of multiple sub-words based on the sub-word granularity vocabulary. Therefore, at the sub-word granularity level, multiplying the probability distributions of all the sub-words corresponding to "coffee" yields the probability distribution product at that sub-word granularity level, denoted as M2.

[0055] At the whole-word granularity level, "coffee" also has a corresponding probability distribution based on the whole-word granularity vocabulary. Therefore, at the whole-word granularity level, multiplying the probability distributions corresponding to "coffee" yields the probability distribution product at that whole-word granularity level, denoted as M3.

[0056] Assuming that in the example above, M1>M2>M3, the text prediction result corresponding to M1 is determined as the target text prediction result. Based on this result, character recognition is performed, and the characters corresponding to the "coffee" image portion in the text image are identified as c, o, f, f, e, e. Based on this, they are combined to form the word "coffee".

[0057] By using the product of probability distributions, the more accurate text prediction results at that granular level are highlighted, allowing for faster and more efficient acquisition of accurate prediction results.

[0058] Method 3: Calculate the mean of the probability distributions indicated by the text prediction results at each of the multiple granularities to obtain the mean of the probability distributions corresponding to the text prediction results at each granularity; determine the text prediction result corresponding to the maximum mean of the multiple probability distributions corresponding to the multiple granularities as the first text prediction result; furthermore, multiply the probability distributions indicated by the text prediction results at each of the multiple granularities to obtain the probability product result corresponding to the text prediction results at each granularity; determine the text prediction result corresponding to the maximum product of the multiple probability product results corresponding to the multiple granularities as the second text prediction result; determine the target text prediction result based on the confidence scores corresponding to the first text prediction result and the second text prediction result; and perform text recognition on the text image based on the target text prediction result.

[0059] In this method, based on the results obtained from Method 1 and Method 2, the prediction results of both methods are comprehensively considered, and the prediction result corresponding to the superior method is selected. The text prediction result corresponding to Method 1 is the first text prediction result in this method, and the text prediction result corresponding to Method 2 is the second text prediction result. Furthermore, the confidence levels of both methods are determined. Based on the respective confidence levels, the text prediction result of the method with the higher confidence level is selected as the target text prediction result, and text recognition is then performed on the text image based on this result. The specific methods for obtaining the confidence levels corresponding to the first and second text prediction results can be obtained using conventional methods and will not be detailed here.

[0060] In this way, better target text prediction results can be selected, providing an accurate basis for subsequent text recognition.

[0061] This embodiment utilizes multi-granularity labeling and parsing of image features to implicitly inject semantic information into the model processing text images. This allows the model to simultaneously combine image features and semantic information for text recognition, improving recognition efficiency and accuracy. Furthermore, this semantic information is represented at multiple granularities—character, sub-word, and whole-word—enabling the acquisition of both image features and semantic information at various levels, thereby elevating text recognition performance to a higher level.

[0062] In practical applications, the above-mentioned text recognition method can also be implemented through a text recognition model. One feasible implementation of a text recognition model may include: a linear projection layer, an attention-based encoder, an adaptive addressing and aggregation layer, and a fusion output layer.

[0063] in:

[0064] A linear projection layer is used to project a sequence of image patches from a text image into a vector of a preset dimension.

[0065] An attention-based encoder, such as the Transformer encoder, is used to implement the functions of the aforementioned attention-based machine learning model, that is, to perform attention calculations on the vectors of the preset dimensions in order to extract and output the corresponding image features from the vectors of the preset dimensions.

[0066] An adaptive addressing and aggregation layer is used to perform multi-granularity label parsing based on image features to obtain corresponding label parsing results at multiple granularities; based on the label parsing results at multiple granularities, corresponding text prediction results at multiple granularities are obtained.

[0067] The fusion output layer is used to determine and output the target text prediction result based on the text prediction results at multiple granularities, so as to obtain the text recognition result of the text image based on the target text prediction result.

[0068] The aforementioned adaptive addressing and aggregation layers may include: character-level adaptive addressing and aggregation layers, sub-word-level adaptive addressing and aggregation layers, and whole-word-level adaptive addressing and aggregation layers.

[0069] An exemplary implementation of the above text recognition model is as follows: Figure 2B As shown in the figure, a W*H RGB image is divided into a series of non-overlapping image patches through a patch operation. These image patches form a sequence of image patches, which is illustrated in the figure as P*P Patches. Here, P*P represents the resolution of each image patch.

[0070] The sequence of image patches is linearly projected into a D-dimensional image patch vector through a linear projection layer. Here, D is set by those skilled in the art according to actual needs.

[0071] After obtaining the D-dimensional image patch vector, the text recognition model adds a learnable [CLS] token to the head of the vector, and adds the position information of the image patch to the [CLS] token and the image patch vector corresponding to each image patch, thus forming a Position+Patch embedding with the [CLS] token added, which is then input into the attention-based encoder. Figure 2B In this diagram, the vector is shown as the vector between the linear projection layer and the Transformer encoder. In this part of the diagram, 0, 1, ... 10 represent position vectors, "*" represents the [CLS] token, and the hollow ellipse formed by combining each position vector 1, 2, ... 10 represents the image patch vector.

[0072] The vector processed above will be output into the attention-based encoder. Figure 2B The diagram shows a Transformer encoder, which extracts image features based on this vector to obtain the corresponding image features.

[0073] Next, the image features are input into the Adaptive Addressing and Aggregation layer, which will be referred to as the A3 layer for ease of explanation. Multiple granularities of label parsing are performed in this A3 layer. Because multi-granularity label parsing is required, the A3 layer is implemented through three independent modules: a character-level A3 module, a BPE-level A3 module, and a word-level A3 module.

[0074] In traditional methods, when performing text recognition on text images, the Transformer encoder directly uses the first 27 tokens of the 256 output sequences as the final output, while the other tokens are not fully utilized, and much useful information is discarded. To fully utilize the information in the Transformer encoder's output sequences for text sequence prediction, the solution in this application uses multiple granularity A3 modules to integrate all the output tokens of the Transformer encoder into a sequence of a preset length (exemplarily, this preset length can be 27, i.e., assuming the longest character is 27). Assume that the element output by each granularity A3 module after attention processing is y. i Let i represent the i-th element, z be the output of the Transformer encoder, and A be the aggregation function. Then the transformation formula used by module A3 is: y i =A i (z L ).

[0075] In one feasible approach, y i =A i (z L ) = softmax(α i (z L )) T (z L U) T

[0076] Where, α i (·) represents a group convolution with a 1x1 kernel (also called a grouped convolution); z L represents the token sequence output by the Transformer encoder; U represents a learnable linear mapping matrix.

[0077] Based on this, for a certain A3 module, the vector Y output after the above calculation is represented as:

[0078] Y = [y1y2; ...; y T ] = [A1(z L A2(z) L );...;A T (z L )]

[0079] Where T is the preset text length, as shown above, and in this embodiment it is 27.

[0080] Next, the text sequence is predicted by classifying the vector Y using the classifiers corresponding to the A3 modules at each granularity.

[0081] In one example, the classifier is represented as G = YWT Where W represents the linear mapping matrix. However, as mentioned earlier, different granularities of the A3 module correspond to different classifiers. In this embodiment, G is used uniformly and no further distinction is made. However, those skilled in the art should understand that Y and W are different in the classifiers corresponding to different granularities, therefore, the classifier G is also different.

[0082] After processing by the A3 module, text prediction results at multiple granularities can be obtained. An example of a multi-granularity prediction result is shown below. Figure 2C As shown in the diagram, for character-level granularity, it can break down words into the finest individual characters. For BPE granularity, it breaks down words into common character substrings; for example, "methodist" is broken down into "method" and "ist"; "university" is broken down into "un" and "iversity"; and "41km" is broken down into "41" and "km". For whole-word granularity, it predicts an entire word.

[0083] Specifically Figure 2B In the example shown, the character-level prediction result for "coffee" is a single character, the BPE-level prediction result is "co" and "ffee", and the whole-word-level prediction result is "coffee".

[0084] Furthermore, the multi-granularity prediction results output by the A3 module will be fused and output through the fusion output layer. It can be fused and output using any one of the fusion output methods described above, namely, method one, method two, and method three, which will not be elaborated here.

[0085] As can be seen from the above, Figure 2B In the text recognition model example shown, image processing and semantic processing of the text image share the same backbone (linear projection layer and attention-based encoder), without a separate semantic processing module. Furthermore, semantic information is implicitly injected into the text recognition model through the A3 module. Moreover, the A3 module can automatically aggregate and select all tokens output by the Transformer encoder through a spatial attention mechanism. In addition, the A3 module includes modules of multiple granularities, predicting at the character, sub-word, and whole-word granularities respectively, thereby obtaining image features and semantic information from more granularities to achieve more accurate and efficient text recognition.

[0086] Reference Figure 3 This document illustrates a schematic diagram of an electronic device according to an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.

[0087] like Figure 3As shown, the electronic device may include: a processor 302, a communications interface 304, a memory 306, and a communications bus 308.

[0088] in:

[0089] The processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 308.

[0090] Communication interface 304 is used to communicate with other electronic devices or servers.

[0091] The processor 302 is used to execute program 310, which can specifically execute the relevant steps in the above-described text recognition method embodiment.

[0092] Specifically, program 310 may include program code that includes computer operation instructions.

[0093] Processor 302 may be a CPU, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The smart device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.

[0094] Memory 306 is used to store program 310. Memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0095] Specifically, program 310 can be used to cause processor 302 to perform the operation corresponding to the text recognition method described in any of the foregoing multiple method embodiments.

[0096] The specific implementation of each step in procedure 310 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0097] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the text recognition methods in the above-described multiple method embodiments.

[0098] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0099] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0100] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0101] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. A text recognition method, comprising: Obtain the image patch sequence corresponding to the text image to be identified; Image features are extracted from the image patch sequence using an attention-based machine learning model to obtain the corresponding image features; Multi-granularity tagging parsing is performed based on the image features, wherein the multi-granularity tagging parsing includes: character-level tagging parsing, sub-word-level tagging parsing, and whole-word-level tagging parsing, wherein the sub-word granularity is a granularity between the character granularity and the whole-word granularity; Based on the multi-granularity label parsing results, text recognition is performed on the text image; The text recognition process based on the multi-granularity label parsing results includes: obtaining text prediction results for multiple granularities based on the multi-granularity label parsing results; obtaining fusion statistics corresponding to the text prediction results for each granularity; and performing text recognition on the text image based on the fusion statistics.

2. The method according to claim 1, wherein, The sub-word granularity tag parsing is the byte-to-encoding granularity tag parsing.

3. The method according to claim 1 or 2, wherein, Based on the image features, sub-word granular tagging and parsing are performed, including: The image features are spatial attention calculated and classified based on the BPE algorithm, and the processing result is used as the tag parsing result at the sub-word granularity.

4. The method according to claim 1 or 2, wherein, Based on the image features, whole-word granular tagging and parsing are performed, including: The image features are spatial attention calculated and classified based on the WordPiece algorithm, and the processing result is used as the labeling and parsing result at the whole word level.

5. The method according to claim 1 or 2, wherein, The fusion statistic is the result of the probability distribution mean and / or probability product; Obtain the fusion statistics corresponding to the text prediction results at each granularity, and perform text recognition on the text image based on the fusion statistics, including: Obtain the mean and / or product of the probability distributions corresponding to the text prediction results at each granularity; Based on the mean of the probability distribution and / or the result of the probability product, the text prediction results of the multiple granularities are fused, and the text image is recognized according to the fusion result.

6. The method according to claim 5, wherein, Based on the mean of the probability distribution, the text prediction results of the multiple granularities are fused, and text recognition is performed on the text image according to the fusion result, including: The mean of the probability distribution indicated by the text prediction results of each of the multiple granularities is calculated to obtain the mean of the probability distribution corresponding to the text prediction results of each granularity. The text prediction result corresponding to the maximum mean of the multiple probability distributions corresponding to multiple granularities is determined as the target text prediction result; Based on the target text prediction result, text recognition is performed on the text image.

7. The method according to claim 5, wherein, Based on the probability product result, the text prediction results of the multiple granularities are fused, and text recognition is performed on the text image according to the fusion result, including: The probability product results corresponding to the text prediction results at each of the multiple granularities are obtained by multiplying the probability distributions indicated by the text prediction results at each granularity. The text prediction result corresponding to the maximum product among the multiple probability products corresponding to multiple granularities is determined as the target text prediction result; Based on the target text prediction result, text recognition is performed on the text image.

8. The method according to claim 5, wherein, Based on the mean of the probability distribution and the result of the probability product, the text prediction results of the multiple granularities are fused, and text recognition is performed on the text image according to the fusion result, including: The mean of the probability distribution indicated by the text prediction results of each of the multiple granularities is calculated to obtain the mean of the probability distribution corresponding to the text prediction results of each granularity; the text prediction result corresponding to the maximum mean of the multiple probability distributions corresponding to the multiple granularities is determined as the first text prediction result. Furthermore, the probability distributions indicated by the text prediction results of each of the multiple granularities are multiplied to obtain the probability product results corresponding to the text prediction results of each granularity; the text prediction result corresponding to the largest product among the multiple probability product results corresponding to the multiple granularities is determined as the second text prediction result. The target text prediction result is determined based on the confidence level corresponding to the first text prediction result and the confidence level corresponding to the second text prediction result. Based on the target text prediction result, text recognition is performed on the text image.

9. The method according to claim 1 or 2, wherein, The text recognition method is executed through a text recognition model, and the attention-based machine learning model is an attention-based encoder. The text recognition model includes: a linear projection layer, the encoder, an adaptive addressing and aggregation layer, and a fusion output layer; in: The linear projection layer is used to project the image patch sequence into a vector of a preset dimension; The encoder is used to perform attention calculation on the vector in order to extract and output the corresponding image features from the vector; The adaptive addressing and aggregation layer is used to perform multi-granularity label parsing based on the image features to obtain corresponding label parsing results at multiple granularities; and to obtain corresponding text prediction results at multiple granularities based on the label parsing results at multiple granularities. The fusion output layer is used to determine and output the target text prediction result based on the text prediction results of the multiple granularities, so as to obtain the text recognition result of the text image based on the target text prediction result.

10. The method according to claim 9, wherein, The adaptive addressing and aggregation layer includes: a character-level adaptive addressing and aggregation layer, a sub-word-level adaptive addressing and aggregation layer, and a whole-word-level adaptive addressing and aggregation layer.

11. An electronic device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform an operation corresponding to the method as described in any one of claims 1-10.

12. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-10.

13. A computer program product comprising computer instructions that instruct a computing device to perform an operation corresponding to any one of the methods described in claims 1-10.

Citation Information

Patent Citations

  • Speech recognition method and device, storage medium and electronic equipment

    CN113990293A

  • Text recognition method and device, readable medium and electronic equipment

    CN114611509A